Alfredo Andere
@AlfredoAndere
benchmarks (dot) bio v1 has many improvements. It now includes an overall leaderboard (1105 tasks), counts refusals as failures, and makes patches to handle misalignments. It has also been re-tested back to Sonnet 4.6 (feb 2026).
Kenny Workman@kenbwork · Sep 23Biology benchmarks must evolve with agent capabilities and behavior. We updated benchmarks.bio after observing agents investigate benchmark identities, browse unrelated material and supply fabricated API contact details.
The update includes:
- Restricted or disabled
Open quoted post →The update includes:
- Restricted or disabled
0 20