https://chatgpt.com/codex/cloud/settings/data#settings/DataC...
https://claude.ai/new#settings/data-privacy-controls
I just realized I've been happily "improving the model for everyone"...
Visit this website https://privacy.openai.com/policies/en/ , click "Make a Privacy Request", choose "Do not train on my content", and complete the form. That submits a formal objection to training on your data, as required by GDPR/your local legislation.
They would do it, the say “ah sorry chaps, impossible to extract it from the dataset by now, anyway we anonymized it so can’t tell what’s what, and we can’t risk losing to China. Oh look - did you see Superman fly outside?”.
This looks incorrect. OpenAI have flat out said that there's no need of doing this and both ways are equivalent
> We respect our users' choice whether to use their data to “improve our models for everyone” regardless of where they express that choice. Users can opt out in the in-app settings or indeed also in our privacy portal. They do not need to opt out in both places, and we will make this clearer in our Help Center.
Humanity is better having solved this issue. But which humans were credited and benefited is the issue.
Or is it just all PFM? (Pure, Fanciful Magic)
How did it become part of their training data if it wasn't public? /confused
If it is oversight, then that's ok and OpenAI can volunteer to make him the lead author which they did. But he's pissed that OpenAI is not letting Levent as well, who had access to internal Anthropic models. So this guy thinks
1. oh my bad i forgot to turn off the consent thing in settings page
2. also i'll collaborate with a literal Anthropic employee who has access to their internal models
3. i'll also reject OpenAI's deal to be the lead author because i want an employee of the competitor to be a part of it
I don't get the mindset.
companies famously always honor those settings!
https://www.theguardian.com/technology/2023/sep/14/google-lo...
https://techhq.com/news/amazon-and-microsoft-both-fined-mill...
Maybe they were fine with contributing training data when their threat model didn't include the case of "OpenAI gets wind of our research and attempts to front-run us"?
oAI has made clear they did not specifically pull in any user data to context for this run.
“We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra.
The latter was only used for writeups and auditing our arguments. For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing).
This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.”
The OpenAI research post states they began training GPT-6 internally on August 28th, and that user chats are used to train models.
“We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .”
For an incredibly niche topic like this, I believe it’s extremely likely that Buckmaster/Levant’s work would influence the direction of OpenAI’s agents’ work even as a de-identified drop in the overall bucket of training data.
So what might have been the incentive for OpenAI to do all this shenanigans? It might have to do with getting its models certified for AGI and getting out of lockin with Microsoft - https://deadneurons.substack.com/p/the-quiet-unwinding-of-mi...
“We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
“After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
Buckmaster reached the key result on August 15, far after the training cutoff.
But I disagree that which humans were credited is the heart of the issue in this particular controversy. The question is what do you need to bring to the table for a result like this. A pre-release frontier model trained on the open literature and $15 million of inference? Or all that plus a year of the experts finding the path to the solution for the model to run with?
I think it makes a huge difference in terms of what we think the future of mathematical research will be like, and whether we should still encourage students to go into this field, which was the original topic of this thread
Option B is legally binding and permanent.
The one in ChatGPT settings only applies to ChatGPT. The one on the OpenAI privacy portal applies to all OpenAI products, present and future.
Here's my answer: If you think we're getting to AGI in the next 9 months, then you believe, with all your heart, that these problems will fall soon. However, there's an ocean to boil in terms of what you could point your limited clusters at. In the meantime, the market is desperate for any sign your company might be first to AGI. Therefore, news of tractability with current models might focus an organization intensely - internally they have a huge leg up on the public, and therefore it's minimal compute to check - and if they are successful, they get approximately $50 million of free PR, likely adding 10-20% to their valuation.
Likewise someone like Tristan is fighting for his (metaphorical) life right now, hoping to preserve his claims of primacy and have a shot at some of that prize money, despite being only partway to a full solution for N-S.
I don't think we see any behavior at all that isn't simple to understand and well described by the setup here, but tell me what you see differently.
In the process they have shafted real-world hardworking mathematicians, which is to say the least, despicable. Note that stories are now coming out from other mathematicians who have also been shafted in a similar manner. Also there are cases where they have had mathematicians accept their "deal" (like the one they offered Buckmaster that he refused) and have OpenAI name linked to their work.
Regarding "their proof", their claim is only solving Navier-Stokes partially for when a smooth force is applied and not a fully general solution (which is maybe impossible). The proof is still being verified and we don't know whether it is just an approximation/hallucination or not.
Their most blatant lie is that they "gave only the problem statement" to their system which then went ahead and solved it. This is almost an impossibility. Problems like these need to identify a specific lead/approach and some work to be done on that path before you can even know whether that approach is promising and worth pursuing. This problem has resisted all attempts at solution for over two centuries. This is where the Buckmaster/Levent's work's importance comes in. They identified a promising approach based on other mathematicians work and have been using both OpenAI and Anthropic's models to make progress and had reached a promising milestone which they published. But they inadvertently gave away their approach to the model's training data set which OpenAI capitalized on by throwing a large amount of compute at the problem to get to the finish first.
This is straightforward stealing of other people's work and building upon it to claim it as your own which can and should be sued. All Scientists/Mathematicians/Researchers who feel OpenAI has done them dirty should band together and file suit.
This whole thing could easily have been avoided if they had worked with the researchers so everybody's concerns/needs are met.
By engaging in this sort of backstabbing, OpenAI has effectively killed the "Goose that laid the Golden Eggs". viz. Researchers were giving away their hard-earned highly specialized knowledge freely to the models in the hope that it will help them get quicker to the result. But now everybody is going to lockdown their research findings and will stop sharing it with the models to the overall detriment of advancement of Science.
What’s a problem is that they haven’t outright denied it. That could be caution and them doing their due diligence first, it could be that it was intentional and they didn’t expect to get caught, or it could be because they have no way of knowing themselves.