Turing
@turingcom
Introducing Behavior2Code from Turing Frontier Research Lab: a more reliable benchmark than ProgramBench for evaluating AI agents on black-box program reconstruction and software reverse engineering tasks.
Best model: 19/60 solved after three attempts. 35 remain unsolved across
Best model: 19/60 solved after three attempts. 35 remain unsolved across
turing.comNew Behavior2Code benchmark: Evaluate how well agents reverse engineer software
0 6