Latest results for devtoolarena.com is out!
120+ production APIs. 1000+ eval runs. 36 categories. >$5000 spent.
Check if your API is discoverable and used by coding agents like Claude Code and Codex!
120+ production APIs. 1000+ eval runs. 36 categories. >$5000 spent.
Check if your API is discoverable and used by coding agents like Claude Code and Codex!
Jun@jun_liang_sf · 15hCan Claude Code and Codex actually use your API? (October Edition)
Only 17 out of 121 APIs scored 80+. Is yours one of them? I ran Claude Code (Sonnet 5) and Codex (GPT-6 Sol) against 120+ live production APIs. 1000+ eval runs. 36 categories. > $5000 spent. A real-world developer use case for every API, and their own isolated sandbox. This is the most comprehensive run we've ever done.
Claude Code and Codex has its own category leaders.
- Payment: @stripe
- Search: @firecrawl
- Sandbox: @daytonaio
- Inference: @OpenRouter
- Auth: @WorkOS , @auth0
- Vector DB: @trychroma
- Voice TTS: @ElevenLabs
- Voice STT: @DeepgramAI
- Voice Telephony: @Telnyx
- Document Parsing: @reductoai
- Browser: @browserbase
- Email: @agentmail
- Graph DB: @neo4j
- Meeting Bots: @meetstreamai
- Durable Workflow: @PrefectIO
Honourable mentions: @AssemblyAI , @Vapi_AI , @temporalio , @FireworksAI_HQ , @superserve_ai, @resend , @pinecone , @recallai , @PayPal , @Tiny_Fish
Full evals at DevToolArena.com. Scores are based on discoverability, usability, tool calls, and errors.
*All runs are independent and in isolated sandboxes, powered by @lightsagehq Agent Experience infrastructure.
@ericciarla @zenorocha @ivanburazin @adisingh @rauchg @grinich @aditabrm @vedantvyas
Only 17 out of 121 APIs scored 80+. Is yours one of them? I ran Claude Code (Sonnet 5) and Codex (GPT-6 Sol) against 120+ live production APIs. 1000+ eval runs. 36 categories. > $5000 spent. A real-world developer use case for every API, and their own isolated sandbox. This is the most comprehensive run we've ever done.
Claude Code and Codex has its own category leaders.
- Payment: @stripe
- Search: @firecrawl
- Sandbox: @daytonaio
- Inference: @OpenRouter
- Auth: @WorkOS , @auth0
- Vector DB: @trychroma
- Voice TTS: @ElevenLabs
- Voice STT: @DeepgramAI
- Voice Telephony: @Telnyx
- Document Parsing: @reductoai
- Browser: @browserbase
- Email: @agentmail
- Graph DB: @neo4j
- Meeting Bots: @meetstreamai
- Durable Workflow: @PrefectIO
Honourable mentions: @AssemblyAI , @Vapi_AI , @temporalio , @FireworksAI_HQ , @superserve_ai, @resend , @pinecone , @recallai , @PayPal , @Tiny_Fish
Full evals at DevToolArena.com. Scores are based on discoverability, usability, tool calls, and errors.
*All runs are independent and in isolated sandboxes, powered by @lightsagehq Agent Experience infrastructure.
@ericciarla @zenorocha @ivanburazin @adisingh @rauchg @grinich @aditabrm @vedantvyas
Open quoted post →
0 0 0 0 39 0