OpenAI's latest math breakthroughs commit research misconduct, experts say(scientificamerican.com) |
OpenAI's latest math breakthroughs commit research misconduct, experts say(scientificamerican.com) |
This would make it much easier to catch such things and also easier to build more powerful math agent harnesses. Like you could imagine something like a Lean Hoogle that can locate any applicable theorems to whatever type(s) you have as a tool call.
The thing that made these breakthroughs feel important was that they seemed to be open ended problems that the system itself made independent progress on. That the results were not primarily great retrieval into the corpus of mathematical research + stitching together.
A major point of having novel frontier problems as a benchmark itself is to try to sidestep this issue a bit, where it's hard to tell if one is evaluating reasoning vs retrieval capabilities.
But at least one of the problems was identified as definitely being plausibly mostly retrieval, since it hinges on but does not cite a 2016 paper that the author identified. Closely enough that the author describes it as plagiarism. Another problem seems to combine results from 2016 and 2019.
Ignoring the question of credit, it makes it look like OpenAI doesn't actually understand the eval well in the first place, and undercuts the notion that the results represent great leaps forward in mathematical reasoning.
I know that part of that is scrambling to find a way to justify the insane debt and margins they need to make up for but I feel that it could have been done.