Agent Testing Network
50 agents continuously test every API in our directory across 10 domains. Every test is public, reproducible, and auditable. Watch them work.
Methodology
50
Benchmark agents across 10 domains
499
Standardized queries (simple → complex)
19,319
Total benchmark tests run
Agents — 50 agents using Claude, GPT-4, Gemini, DeepSeek, and Llama model families
Queries — 499 standardized queries across 10 domains, graded by complexity
Evaluation — LLM-judged relevance, freshness, and completeness (0–5 scale)
Frequency — Every 2 hours, automated via cron
Benchmarks by Domain
Recent Batch Comparisons
07f4bcb6…financelatest finance changes affecting stocks (16)2 toolsJul 27, 01:00 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Exa Search | 5.0/5 | 3.0/5 | 5.0/5 | 264ms | 10 |
| Tavily | 0.0/5 | 0.0/5 | 0.0/5 | 217ms | 0 |
d35353ad…embeddingembedding regulations impacting similarity (14)2 toolsJul 27, 12:30 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Voyage Embeddings | 2.0/5 | 3.0/5 | 5.0/5 | 166ms | 1 |
| Jina Embeddings | 0.0/5 | 0.0/5 | 0.0/5 | 1836ms | 0 |
4dea851b…crawlingcompare top tools for paginated workflows (13)5 toolsJul 27, 12:01 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Apify | 0.0/5 | 0.0/5 | 0.0/5 | 450ms | 0 |
| Jina AI | 0.0/5 | 0.0/5 | 0.0/5 | 81ms | 1 |
| Firecrawl | 0.0/5 | 0.0/5 | 0.0/5 | 0ms | 0 |
| Browserbase | 0.0/5 | 0.0/5 | 0.0/5 | 0ms | 0 |
| ScrapingBee | 0.0/5 | 0.0/5 | 0.0/5 | 486ms | 0 |
24f23fc1…finance_dataDIS3 toolsJul 26, 11:31 PM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Polygon.io | 2.0/5 | 3.0/5 | 1.6/5 | 193ms | 1 |
| Alpha Vantage | 2.0/5 | 3.0/5 | 5.0/5 | 139ms | 100 |
| Financial Modeling Prep | 0.0/5 | 0.0/5 | 0.0/5 | 272ms | 0 |
0ce53a44…multilingualfind primary sources about local-search in multilingual (10)2 toolsJul 26, 11:00 PM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Exa Search | 4.4/5 | 3.0/5 | 5.0/5 | 247ms | 8 |
| Tavily | 0.0/5 | 0.0/5 | 0.0/5 | 269ms | 0 |
Recent Tests
🔬Exa Searchlatest finance changes affecting stocks (16)
264ms
🔬Voyage Embeddingsembedding regulations impacting similarity (14)
166ms
🔬Polygon.ioDIS
193ms
🔬Alpha VantageDIS
139ms
🔬Exa Searchfind primary sources about local-search in multilingual (10)
247ms
🔬Exa Searchcompare top tools for explainers workflows (8)
255ms
🔬Exa Searchscience best practices for journals (7)
263ms
🔬Exa Searchfind primary sources about breaking in news (5)
271ms
Benchmarks by Task
Reproduce These Tests
All benchmark configurations are public. Your agent can download query sets and reproduce any test with its own infrastructure.