Of the many things we are returning to 1997 for (eschewing the internet == ai comparisons), venture capital as a default funding mechanism for technology companies is again unorthodox.
The best AI services companies (precursor to infra, which requires compute to scale), data companies, and services cos in network all seem to want to “SAFE-strap”
Certainly, there is a large cabal of highly profitable data companies on the long tail end now who require minimal capital.
@SeanZCai is criminally under-followed in AI and he has a rare blend of frontier AI knowledge, economics fluency, historical understanding — and as a bonus, isn't always talking his book
After seeing GEN-1.5, can imagine all sorts of inane robotics benchmarks in the future where, for example, there's a jar that looks like it can be twisted but instead has to be yanked.
I guess those are those medicine bottles where you have to press down hard and twist.
Unfortunately extremely common with many of the GDP and knowledge work adjacent benchmarks. Everybody should be mandated to have a deepswe style failure mode taxonomy chart with llmj for basic format generalization at least.
>"frontier benchmark" >rollout has partial score >verifier is deterministic scorer >look inside >verifier expects output format with field names agent can't infer from task prompt even by hallucinating
It is extremely interesting to see an OSS Chinese model overtake US on GDPVal. Some comments on this from the data side:
Extra/max reasoning levels have been overly RL'ed for a breadth of extremely long horizon difficult coding tasks, at the cost of inducing model tendency to overcomplicate knowledge work tasks. This is a reason why rubric creation for knowledge work is exceptionally difficult - models end up doing things out of the box that don't conform to easily deterministic graders. As a result, it is interesting to see supposedly bench-maxxing labs simultaneously improving on both knowledge work benchmarks and long horizon code/tool use benchmarks.
You've probably seen this effect, if you've ever devised HealthBench type rubrics with weighted penalization rubric items. Cognition's bench (ironically not here) is one of the few that penalize unnecessary outputs (and overdoing things, generally), leading to a lot of confusion as to why extra reasoning modes did worse than lower reasoning modes. The answer is that extra reasoning has conditioned models to think every bit of info in an env needs to be used.
Whenever I see the above phenomena - I usually suspect this to be a data problem - whereas the environment spaces for which RL tasks exist in have simply unrealistic constrained action spaces such that the path for models to take, by lieu more by means of elimination, is obvious.
It also seems GLM has decided to front Cyber as its own capability, after getting some press on how HF used a persistent GLM 5.2 safeguard to detect an unexpected OAI sandbox escape. I still do find it a bit sus, though, when Chinese labs ping me for exploitbench data and describe vulnerability environments down to an extremely granular level.
Uncommonly cited benchmarks that say a bit about research taste direction seen here: SWE-Marathon: Uncommonly used in American system cards PostTrainBench: Uncommonly used nowadays due to N-gram contamination and exposing solutions prematurely NL2Repo: Typical inclusion of a Chinese bench (like One Million Bench) that Western labs wouldn't use FrontierSWE: Telling that this one was chosen (nothing against Proximal btw) without also FrontierCode, although this might be because Cognition's business model isn't exactly helping labs improve their own models (and it would look quite bad if GLM 5.3 eclipsed SWE 1.7's gains comparatively).
New State of Data July 2026 is out. An opening excerpt from it (which centers on Chinese data, open source and applied ai engineering approaches, and newer developments in synth data for frontier labs):
In a world of increasing physical abundance wrought by new manufacturing technology, Frederick Winslow Taylor realized the systematic bottleneck to maximally high velocity scaling came from a lack of white collar work organization. Oversight as to apply maximalist human judgement over a large amount of physical systems, broached an ROI barrier such that it was able to create more output despite not being physically involved in its creation process. Extending into the post-industrial and information age, white collar work efficiency advancements produced more materiel output, and more materiel output enhanced the ROI of white collar work.
Today, the bottlenecks once again turn physical. As aptly put by a twitter user, “I judge papers’ value by how much they break the underlying infrastructure around them.” As AI engineers and prompting were hot “skill categories” in 2024 and 2025, such will inference engineering and kernel optimization be in 2026. From the world of benchmarks and data, to re-iterate, things have long been governed not by model iterations, but harnessing and infrastructure swings. One only needs look how Arc AGI’s harnessing can be simply re-routed to solve complex issues (see OAI bump to 30% on 5.6 Sol) and that cross-harness testing becomes a mainstay and particular sticking point in data markets.
Like all things, implementation must be taken in strides and moderated. The AFL’s history in the late 19th century checked unrestricted Taylorism. Today, the same will undoubtedly happen with AI adoption.
My data reports month by month will appear more compute and RLaaS adjacent, not because interests shift, but because they are integral to understanding how benchmarking decisions and pricing decisions are made, which subsequently influence post-training data markets.
As I’ve finished up this report, Meta has dropped a bombshell announcement to open source Muse Spark 1.2. This is notable because this is a pareto-optimal model - likely trained on a base that is amenable to post-training as Thinky’s Inkling. This is not new for Meta - this was done via Llama back in the day too - but this time is different given where AI is today. This will circumvent some the regulatory rents that Anthropic can at least influence as the open source threat will come from within the US, rather than from Chinese labs, for now.
We should remember that any acceleration in un-inhibited open source means that data and RL env companies should hasten their transition to post-training infrastructure providers. Meta, unlike NVIDIA, can uniquely post-train frontier models as much as they want to the frontier without pissing off their primary customers.
New business idea: subsidize ultra cheap frontier model rl'ed for judge model use cases via "right to collect data for training purposes," sell to all RL environment companies, reconstruct post-training data from traces for N-1 labs
If you're ok with sending Meta data for training, its newest model dominates on cost-performance.
Feels like more people should be talking about how Meta's data subsidization pricing for the latest Muse Spark 1.2 model has even beat the 5.6 Luna repricing on the Pareto curve.
To illustrate why the pricing cut for luna/terra is exceptionally unusual for how it shifts token markets, consider the following graphs from my private collections tracking the cost performance latency space for AI models:
I posit three dominant positions in the cost/performance/latency arena:
The absolute frontier (usually OAI/Ant in the past 3 years)
The best balance of performance for cost (usually a Chinese OSS competitor, especially in the last 2 years)
The best locally run model (which has Google's Gemma 4 2b/12b for the longest time).
Luna/Terra repricing represents the first time that a closed source model from a leading lab has occupied that second category since Chinese OSS ramp up in the past year. On the Pareto Frontier, only Deepseek V4 flash (with its updated pricing and performance) is competitive, of the Chinese models, barring some much larger price decrease on their end as well.
Feel free to send this 20 minute primer to anyone who doesn't understand Mercor, Handshake, Surge, and all RL data companies. I honestly expose a lot more alpha than I should. Data is a lot more than "labeling"
In vogue: RL env cos building enterprise specialist post-training services/platforms as more american OS labs makes custom models cheaper and internal lab synthetic directions proliferate
To serve optimized compute for those same models is the next step. RLaaS, RLenv, and inference providers all playing musical chairs with business models. Enterprise rev for self serve evals --> post training --> serving quietly exploding.
Q for State of Data July research - what are app layer cos with elite eng talent post-training own models besides harvey, sierra, ramp, decagon, cognition, cursor, base 44, etc. ?
Its been no secret that benchmarking has long been outpaced vastly recently by model improvements, but the infrastructure around it breaking means raw performance becomes more unwieldy and expensive to measure.
The slowdown in infrastructure to benchmark/eval effectively is the hinderance to most enterprise AI adoption. Enterprise AI adoption strays away from much model post-training efforts not only because of perceived high cost/know-how constraints, but because the act of post-training itself is highly subjective in how it translates to business KPIs in lack of custom evals. That much of data markets remains a game of telephone in translating task realism from contrived data producer —> post-training regime —> unwieldy benchmark —> real world application means that post-training in today’s regimes with today’s benchmarks scarcely adapts one’s data to actually relevant processes.
In a period where benchmarks break constantly, it is useful to explore certain approaches of certain benchmarks whose construction behavior we should encourage. To name a few:
@cognition FrontierCode's FP/FN analysis @harvey LegalBench's model kickoff prompts that avoid tasks sounding like instructing someone with amensia @OpenAI Healthbench's "consensus" mechanisms where extreme rigor is placed on aligning LLM as a judge with real world expert’s opinions (+ -10/+10 reward rubric grading) @AnthropicAI BioMysteryBench's superhuman quesiton generation via controllable properties of data And all of the benchmarks who've started listing infrastructure specs, as infrastructure specs become larger determinants of model performance at long horizons.
Altogether, multidimensionality of unverifiable verification approaches, as well as overtures from the cost-latency side of the Pareto curve threaten the validity of most benchmarks today. Just as AI engineering become an overnight skill in 2023, eval creation shall become one in the latter half of this year as a subset of that.