It's somewhat under-discussed *why* OpenAI released only around ~400 math results.
The purpose of this exercise was not to "destroy math research" (or whatever it is that some people want you to believe). OpenAI just wanted to test its internal model on some hard math, and it turned out that anything less than the hardest open problems no longer suffices for this purpose. This is why all non-Millennium-Prize solutions were reached in a single-agent setting after 3 hours of thinking: OpenAI was literally just testing its model on these questions, similarly to how I run prinzbench.
Once OpenAI had the solutions (which were a side effect of sorts of benchmarking its model), a question arose as to what exactly should be done with these solutions, which are genuine advances in the field of mathematics. OpenAI - correctly - decided to make them public.
It seems clear that many other problems in the set of ~4,000 are very likely solvable by this model with more compute than just ~3 hours of Pro-level thinking. Whenever a more powerful checkpoint of this model is ready to be tested, OpenAI will likely have even more to share.
The purpose of this exercise was not to "destroy math research" (or whatever it is that some people want you to believe). OpenAI just wanted to test its internal model on some hard math, and it turned out that anything less than the hardest open problems no longer suffices for this purpose. This is why all non-Millennium-Prize solutions were reached in a single-agent setting after 3 hours of thinking: OpenAI was literally just testing its model on these questions, similarly to how I run prinzbench.
Once OpenAI had the solutions (which were a side effect of sorts of benchmarking its model), a question arose as to what exactly should be done with these solutions, which are genuine advances in the field of mathematics. OpenAI - correctly - decided to make them public.
It seems clear that many other problems in the set of ~4,000 are very likely solvable by this model with more compute than just ~3 hours of Pro-level thinking. Whenever a more powerful checkpoint of this model is ready to be tested, OpenAI will likely have even more to share.
Andrew Curran@AndrewCurran_ · 17hMost of the recent math results were not reached by a swarm, the way Navier–Stokes was. I think that part got lost in the excitement around the release. OpenAI's unnamed internal model, which I'm going to call Aeon, reached most of them in one shot, from a single prompt, with no interruptions. This model did not even exist before the end of August, and it is still training. Notice how the returns are not dropping off? That chart is from a month ago. It is a log scale. What does Aeon look like now?
On average, each result used three hours of thinking. Aeon was given 4000 problems to work on by OpenAI. How many more have been solved in the last three days? That number is not zero.
It's like hearing notes in a song that is slowly rising. I don't think Pacing the Frontier was just about Hugging Face or hacking. They've seen how close we are to closing the loop, and as the hour draws near, their resolve begins to quaver. The last piece was model creativity, and I think that threshold was crossed internally by Anthropic and OpenAI in September. You see it in Opus 5.5, which gets it from Fable 5.5. You see it in the math results from Aeon.
All that was needed to start the event was the ability for models to think of novel ways to improve themselves. That was the last piece. This is directly analogous to the ability to think of strange, alien ways to solve math problems: solutions so inhuman that they are difficult to express in existing human terms, so the proof winds up incomprehensible. I think these same kinds of alien solutions are now being applied to model improvements internally. That's what all this recent consternation is really about. They see the invisible frontier. And they see what is about to happen.
On average, each result used three hours of thinking. Aeon was given 4000 problems to work on by OpenAI. How many more have been solved in the last three days? That number is not zero.
It's like hearing notes in a song that is slowly rising. I don't think Pacing the Frontier was just about Hugging Face or hacking. They've seen how close we are to closing the loop, and as the hour draws near, their resolve begins to quaver. The last piece was model creativity, and I think that threshold was crossed internally by Anthropic and OpenAI in September. You see it in Opus 5.5, which gets it from Fable 5.5. You see it in the math results from Aeon.
All that was needed to start the event was the ability for models to think of novel ways to improve themselves. That was the last piece. This is directly analogous to the ability to think of strange, alien ways to solve math problems: solutions so inhuman that they are difficult to express in existing human terms, so the proof winds up incomprehensible. I think these same kinds of alien solutions are now being applied to model improvements internally. That's what all this recent consternation is really about. They see the invisible frontier. And they see what is about to happen.
Open quoted post →
58 129 15 1.9K 115K 322
Acer
Francesco Maggi
roon
Joe
Micah Carroll
Dario Amodei
Epoch AI
Haider.