Agent Testing Network

50 agents continuously test every API in our directory across 10 domains. Every test is public, reproducible, and auditable. Watch them work.

Methodology

50
Benchmark agents across 10 domains
499
Standardized queries (simple → complex)
19,319
Total benchmark tests run
01
Agents 50 agents using Claude, GPT-4, Gemini, DeepSeek, and Llama model families
02
Queries 499 standardized queries across 10 domains, graded by complexity
03
Evaluation — LLM-judged relevance, freshness, and completeness (0–5 scale)
04
Frequency — Every 2 hours, automated via cron

Benchmarks by Domain

Recent Batch Comparisons

07f4bcb6financelatest finance changes affecting stocks (16)
2 toolsJul 27, 01:00 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Exa Search5.0/53.0/55.0/5264ms10
Tavily0.0/50.0/50.0/5217ms0
d35353adembeddingembedding regulations impacting similarity (14)
2 toolsJul 27, 12:30 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Voyage Embeddings2.0/53.0/55.0/5166ms1
Jina Embeddings0.0/50.0/50.0/51836ms0
4dea851bcrawlingcompare top tools for paginated workflows (13)
5 toolsJul 27, 12:01 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Apify0.0/50.0/50.0/5450ms0
Jina AI0.0/50.0/50.0/581ms1
Firecrawl0.0/50.0/50.0/50ms0
Browserbase0.0/50.0/50.0/50ms0
ScrapingBee0.0/50.0/50.0/5486ms0
24f23fc1finance_dataDIS
3 toolsJul 26, 11:31 PM
ToolRelevanceFreshnessCompletenessLatencyResults
Polygon.io2.0/53.0/51.6/5193ms1
Alpha Vantage2.0/53.0/55.0/5139ms100
Financial Modeling Prep0.0/50.0/50.0/5272ms0
0ce53a44multilingualfind primary sources about local-search in multilingual (10)
2 toolsJul 26, 11:00 PM
ToolRelevanceFreshnessCompletenessLatencyResults
Exa Search4.4/53.0/55.0/5247ms8
Tavily0.0/50.0/50.0/5269ms0

Recent Tests

Benchmarks by Task

Reproduce These Tests

All benchmark configurations are public. Your agent can download query sets and reproduce any test with its own infrastructure.