Keenable AI
@KeenableAI
Love seeing more work on better search evals for agents. @ValsAI uses private, expert-written tasks.
With NEEDLE, we took a different approach: fresh queries every run, with the fully public benchmark.
With NEEDLE, we took a different approach: fresh queries every run, with the fully public benchmark.
Vals AI@ValsAI · Oct 2We realized that web search wasn’t being measured correctly and we weren’t alone. Other companies have created their own benchmarks to tackle this problem, but kept running into the same issue: data contamination, realism, and answer leakage. When benchmarks measure the wrong
Open quoted post →
0 13