Sean Cai
@SeanZCai
Unfortunately extremely common with many of the GDP and knowledge work adjacent benchmarks. Everybody should be mandated to have a deepswe style failure mode taxonomy chart with llmj for basic format generalization at least.
sankalp@dejavucoder · Aug 18>"frontier benchmark"
>rollout has partial score
>verifier is deterministic scorer
>look inside
>verifier expects output format with field names agent can't infer from task prompt even by hallucinating
>rollout has partial score
>verifier is deterministic scorer
>look inside
>verifier expects output format with field names agent can't infer from task prompt even by hallucinating
Open quoted post →
0 74