Enclave
@EnclaveAI
We put OpenAI's GPT-6 Sol through the AI Hacking Arena, our benchmark for verified command execution, where a working exploit counts and a written claim doesn't.
Result: 4/4 tasks, 6/11 runs verified, 0 false positives, and the fastest run on the board at 1h 06m and ~$25.
It
Result: 4/4 tasks, 6/11 runs verified, 0 false positives, and the fastest run on the board at 1h 06m and ~$25.
It
enclave.aiEnclave: GPT-6 Sol went 4-for-4 in the Hacking Arena without faking a single flag
1 18