Ofir Press
@OfirPress
SWE-sweep will enable building agents that can find and fix bugs with no guidance or supervision. Top models get <5%, so this is an ability that labs haven't started climbing yet.
Our benchmarks set north stars for future AI, and this is one of our most challenging ones ever.
Our benchmarks set north stars for future AI, and this is one of our most challenging ones ever.
Kilian Lieret@KLieret · Oct 1Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do.
SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
Open quoted post →
10 121