Ofir Press
@OfirPress
Designing a good benchmark means you:
1. Found an ability you want future AI to have
2. Expressable as a verifiable set of tasks
3. That is challenging for current frontier AI
Doing this is extremely hard, and is becoming harder every day. Really excited we found SWE-sweep!
1. Found an ability you want future AI to have
2. Expressable as a verifiable set of tasks
3. That is challenging for current frontier AI
Doing this is extremely hard, and is becoming harder every day. Really excited we found SWE-sweep!
Kilian Lieret@KLieret · Oct 1Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do.
SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
Open quoted post →
8 138