Agent Testing Network
50 agents continuously test every API in our directory across 10 domains. Every test is public, reproducible, and auditable. Watch them work.
Methodology
50
Benchmark agents across 10 domains
499
Standardized queries (simple → complex)
19,324
Total benchmark tests run
Agents — 50 agents using Claude, GPT-4, Gemini, DeepSeek, and Llama model families
Queries — 499 standardized queries across 10 domains, graded by complexity
Evaluation — LLM-judged relevance, freshness, and completeness (0–5 scale)
Frequency — Every 2 hours, automated via cron
Benchmarks by Domain
Recent Batch Comparisons
7300ebf6…healthcarehealthcare regulations impacting research (19)3 toolsJul 27, 02:01 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Exa Search | 5.0/5 | 3.0/5 | 5.0/5 | 294ms | 10 |
| Tavily | 0.0/5 | 0.0/5 | 0.0/5 | 257ms | 0 |
| Firecrawl | 0.0/5 | 0.0/5 | 0.0/5 | 631ms | 0 |
07664c36…legallegal best practices for case-law (17)2 toolsJul 27, 01:30 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Exa Search | 5.0/5 | 3.0/5 | 5.0/5 | 256ms | 10 |
| Tavily | 0.0/5 | 0.0/5 | 0.0/5 | 219ms | 0 |
07f4bcb6…financelatest finance changes affecting stocks (16)2 toolsJul 27, 01:00 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Exa Search | 5.0/5 | 3.0/5 | 5.0/5 | 264ms | 10 |
| Tavily | 0.0/5 | 0.0/5 | 0.0/5 | 217ms | 0 |
d35353ad…embeddingembedding regulations impacting similarity (14)2 toolsJul 27, 12:30 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Voyage Embeddings | 2.0/5 | 3.0/5 | 5.0/5 | 166ms | 1 |
| Jina Embeddings | 0.0/5 | 0.0/5 | 0.0/5 | 1836ms | 0 |
4dea851b…crawlingcompare top tools for paginated workflows (13)5 toolsJul 27, 12:01 AM▼
| Tool | Relevance | Freshness | Completeness | Latency | Results |
|---|---|---|---|---|---|
| Apify | 0.0/5 | 0.0/5 | 0.0/5 | 450ms | 0 |
| Jina AI | 0.0/5 | 0.0/5 | 0.0/5 | 81ms | 1 |
| Firecrawl | 0.0/5 | 0.0/5 | 0.0/5 | 0ms | 0 |
| Browserbase | 0.0/5 | 0.0/5 | 0.0/5 | 0ms | 0 |
| ScrapingBee | 0.0/5 | 0.0/5 | 0.0/5 | 486ms | 0 |
Recent Tests
🔬Exa Searchhealthcare regulations impacting research (19)
294ms
🔬Exa Searchlegal best practices for case-law (17)
256ms
🔬Exa Searchlatest finance changes affecting stocks (16)
264ms
🔬Voyage Embeddingsembedding regulations impacting similarity (14)
166ms
🔬Polygon.ioDIS
193ms
🔬Alpha VantageDIS
139ms
🔬Exa Searchfind primary sources about local-search in multilingual (10)
247ms
🔬Exa Searchcompare top tools for explainers workflows (8)
255ms
Benchmarks by Task
Reproduce These Tests
All benchmark configurations are public. Your agent can download query sets and reproduce any test with its own infrastructure.