Agent platforms on Merge can now add 9 You.com tools: search, page contents, cited answers, research, and finance research. Users bring their own API key.
Claude Haiku 5.5 from Anthropic (@AnthropicAI) is now available on Merge Gateway.
Anthropic calls it their fastest model yet, and early customers are seeing up to 2.5x faster inference per agent turn while costing 75% less than Haiku 4.5.
Two frontier models a few points apart on the Intelligence Index can differ 7x in the tokens they burn per task.
GPT-6 Astra (@OpenAI) scores 52.7 and finishes an average Artificial Analysis Intelligence Index task in 27K output tokens. Claude Sonnet 5.5 scores 56.0 and uses 197K. GPT-6.1 Sol comes in at 38K, while Muse Spark 1.3, Gemini 4 Argon and MiMo-V2.6-Pro all land between 60K and 64K.
Output tokens are what you pay for. Sending simple requests to a leaner model is one of the easiest ways to cut inference spend, and Merge Gateway routes them for you.
.@FastCompany named @Merge_api one of the 7 next big things in foundational AI for 2026.
A year ago we launched Agent Handler on a bet that agents are only as useful as the systems they can safely reach.
Being recognized among the breakthroughs shaping how we work and live is an incredible milestone for our team and the customers building with us. Thank you, Fast Company!
Merge has been named one of @FastCompany's 7 Next Big Things in Foundational AI for 2026!
Agents are only as good as the data they can reach. Merge is the layer that connects them to tools like @Salesforce, @Workday, and @GitHub, while keeping access secure.
Switching models is scary when the only real test is in production.
Gateway Evals grades a new model on your agent's real tasks, then shadows it on live traffic. You see cost, latency, and every answer before a customer does.
Gemini 4 Argon (@GoogleDeepMind) makes up an answer about half as often as the next best frontier model.
On AA-Omniscience, when Argon doesn't know something it guesses wrong 15% of the time instead of saying so. Qwen3.8 Max is next at 29%, with Grok 4.7 at 29% and Muse Spark 1.3 at 32%. GPT-6 Astra has the highest accuracy of the six at 61%, but still guesses on 45% of the questions it misses.
Whether a wrong answer or no answer costs you more depends on the task. With Merge Gateway you can send each request to the model that fits it.
The biggest generational jump since June came from a model that costs $0.18 per million output tokens.
DeepSeek-V4-Flash-0731 (@Deepseek_AI) scored 26.8 points higher on Toolathlon than the release it replaced. GLM-5.3 gained 24.8 on the same board. Claude Opus 5.5 added 20.7 on SWE-bench Pro, and Claude Sonnet 5.5 added 18.1 there six days later.
When an integration stops syncing, the error rarely says how to fix it.
AI-powered issue resolution is now live in Merge Unified. Open any issue and Merge explains what went wrong, drawing on years of supporting integrations in production.
If the fix is on your customer's side, you get an explanation ready to send them. If it's on yours, Merge walks you through the fix in the dashboard.
Sending every agent task to the most capable model gets expensive fast.
@EragonAI now routes each task to the right model through one Merge Gateway endpoint. Customers get the same quality, while Eragon saves hundreds of thousands of dollars in inference.
The best open-weight model for code is third overall, at a 23rd of the price of the model above it.
MiMo-V2.6-Pro scores 60.9 on SciCode against Claude Opus 5.5's 66.9, and runs $0.87 per million output tokens against $20. Kimi K3 and GLM-5.3 follow within two points of it.
Six points for $19 per million tokens is a real tradeoff, and it's the kind a Build-Your-Own-Router policy on Merge Gateway lets you set per request rather than once.