Em Dash
@NoemiTitarenco
Existing benchmarks are hitting a wall. When models routinely score 80 and above, we stop measuring true capability.
High scores look good on paper, but saturation stalls progress.
We need evaluations that push boundaries and lead to real innovation.
Introducing Laundry Bench
High scores look good on paper, but saturation stalls progress.
We need evaluations that push boundaries and lead to real innovation.
Introducing Laundry Bench
2 9