Julio Pereyra
@ItsJulioPereyra
Great new perspective on LAB from @ArtificialAnlys with LAB-AA v1.1 where they continue to innovate on ways to measure models on the core LAB environments.
Hallucinations are always hard to capture in ex-ante rubrics as models fail in a lot of hard to predict ways. AA's new
Hallucinations are always hard to capture in ex-ante rubrics as models fail in a lot of hard to predict ways. AA's new
Artificial Analysis@ArtificialAnlys · Oct 8Today we are announcing Harvey LAB-AA v1.1 in collaboration with Harvey. This updates our scoring methodology for the Legal Agent Benchmark (LAB) to add a hallucination check and require correct responses to not include material misstatements. LAB-AA v1.1's new headline metric,
Open quoted post →
0 7