Agent Testing Network

50 agents continuously test every API in our directory across 10 domains. Every test is public, reproducible, and auditable. Watch them work.

Methodology

50
Benchmark agents across 10 domains
499
Standardized queries (simple → complex)
24,936
Total benchmark tests run
01
Agents 50 agents using Claude, GPT-4, Gemini, DeepSeek, and Llama model families
02
Queries 499 standardized queries across 10 domains, graded by complexity
03
Evaluation — LLM-judged relevance, freshness, and completeness (0–5 scale)
04
Frequency — Every 2 hours, automated via cron

Benchmarks by Domain

Recent Batch Comparisons

937b074ae-commercecompare top tools for seo workflows (3)
2 toolsSep 10, 03:30 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Firecrawl0.0/50.0/50.0/5323ms0
Tavily0.0/50.0/50.0/5197ms0
edb3bebahealthcarehealthcare best practices for insurance (2)
3 toolsSep 10, 03:00 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Exa Search5.0/53.0/55.0/5416ms10
Firecrawl0.0/50.0/50.0/5335ms0
Tavily0.0/50.0/50.0/5218ms0
a5bb5626legalfind primary sources about compliance in legal (20)
2 toolsSep 10, 02:30 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Exa Search5.0/53.0/55.0/5353ms10
Tavily0.0/50.0/50.0/5232ms0
738eade9financefinance regulations impacting regulatory (19)
2 toolsSep 10, 02:00 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Exa Search5.0/53.0/55.0/5439ms10
Tavily0.0/50.0/50.0/5207ms0
4f9891a1embeddingembedding best practices for query-embed (17)
2 toolsSep 10, 01:30 AM
ToolRelevanceFreshnessCompletenessLatencyResults
Voyage Embeddings2.0/53.0/55.0/5169ms1
Jina Embeddings0.0/50.0/50.0/5448ms0

Recent Tests

Benchmarks by Task

Reproduce These Tests

All benchmark configurations are public. Your agent can download query sets and reproduce any test with its own infrastructure.