Codex for open source(openai.com) |
Codex for open source(openai.com) |
i imagine the usage from maintainers of high quality projects are excellent training data. much better than average joe
I mean seriously, you already ripped off all the worlds open source code. Be more generous and don't demand anything else back. Six months is so little too.
Whenever companies do things like this, it's both, or at least trying hard to be. To the extent that it's perceived by developers (that is, potential OpenAI customers) as helping OSS, it's effective marketing. This perception may or may not correspond to reality.
Correction: only in part
Yes, Amazon is the only company named, but would anyone be surprised if OpenAI was one of the other five companies? It's hard to imagine a company that would materially benefit more from this event.
The evidence is circumstantial, of course, but can you blame people for making a connection?
[1] https://www.axios.com/2026/06/13/anthropic-amazon-white-hous...
I've forked tensorzero after they archived the repo and will be updating and fixing issues going forward.
https://github.com/agentify-sh/gateway
this is my 2nd attempt
I am using my idle codex usage but would benefit from more inference
Companies a thousandth their size are giving free or at-cost access for OSS projects.
If you have more than 100 stars, you can get $50 in starter credit.
Ideally organizations, more so than people, provide the bulk of future donations.
As for this program, ehh... Sceptical in general of any frontier program that ends at some time.
Once you're embedded, and all that...
(no internal knowledge, this is based on my experience with explainshell.com, thanks OAI!)
If they were really serious about supporting OSS, they'd offer it for free perpetually (well, with periodic checks to ensure the maintainers are still affiliated with the project). Anything less just makes it look like a marketing stunt.
And also, dumb how Github-centric this is, same as Anthropic's signup form. Most of my OSS contributions aren't on Github. Guess that means the projects I've worked on don't matter.
It's not like the Linux kernel isn't real. It's just that the kind of people who write Linux kernel patches and get them accepted are, in the eyes of an average open source developer, somewhere between "majestic magical creatures" and "madmen".
[1] https://www.jetbrains.com/store/?section=students&billing=ye...
I’m building EasyInvoicePDF - a free and open-source invoice generator. (900+ GitHub stars, 2k monthly users on average, 10k total invoices downloaded)
Its arguably even more self-serving than the drug dealer tactic because of the feedback loop involved (if you use it to maintain your open source project, OpenAI will surely use that new code [along with all the existing code in your project] to train future models).
So it would be like if the drug dealer gave you the first taste for free and also the drug caused you to shit out more drugs and the drug dealer harvested your shit to sell to both future you plus other people.
Isn't the thing open source and governed by its own license?
especially in larger projects where maintainership duties are heavily delegated, the last thing i want is some tool that can only be used by me, because suddenly i can no longer share the workload that tool targets with people who aren't "technically" maintainers.
IMO this is an insult if anything
It should be maintained by humans, relying on widely available hardware and software, requiring little of both.
Not saying that using LLMs as a convenience is forbidden or anything, but the direction is problematic.
(Also, this sounds like a cheap alternative to actually funding FOSS work.)
We got it yesterday, maybe they just started rolling it out and hence op posted this.
my approach to open source development with AI now is to include all of the agent sessions used in development in the repository, which makes this data freely available for training for both proprietary and open weights models, but that is just my own approach. every open source developer ultimately has to make their own judgement on the best way to integrate AI in accordance with their values.
I see it as a chance. Many OS projects themselves offer LLM readable websites, their docs.
This way the project at least not only gets ingested but receives referential treatment.
Some sort of collaboration. Ingested it will be, anyway.
absolutely. AI is the same as any other software, and open source has to integrate, adapt, and lead to make sure that open source values continue to propagate.
my personal approach is to focus on developing with open weights models, so that my work is optimized for them, and leads to their development. proprietary labs are free to copy, but they have a structural cost disadvantage. my objective is that open weights models remain competitive on capability but lead on capability/cost.
I'm sure they train their models on open source software, so how do I know that LLM generated code doesn't reproduce substantial chunks of, for example, GPL licensed code? If indeed there are GPL violations, what are AI companies doing to police themselves?
I wonder if open source licenses will start to include "not to be used for LLM training" clauses.
As if the LLM trainers would care. They've ignored every single license and copyright policy out there because "fair transformative use". It's undergoing litigation in various jurisdictions, and the chaotic side of me really wants to see what happens if a UK or California decide that training an LLM on pirated copyrighted material is not fair use, and the rights holders have to be compensated.
How dare they only give me this much free stuff! I want that much free stuff!
(Actually I don't, I want their stuff as Free Software and I mean everything, training data, pipelines and all)
https://www.oss.fund/explore/?pillar=operational-support&cat...
Also applied for Codex Open Source - within 2 days I got confirmation and using it since then. Great Job OpenAI! Shame on you Antrophic for not even sending out a refusal message with the cause.
It's like how after using Apple hardware for years I couldn't put up with most Windows laptops -- either they were HiDPI ultrabooks with no performance or they were sloppy gamer machines with no class.
Learning JetBrains gets you hooked.
In my experience most SaaS apps do not filter this out and allow re-sign ups with sub-addresses.
Gmail has an additional behavior that dot character is ignored in local component of the address . multiple@gmail.com, mult.iple@gmail.com mult.ip.le@gmail.com all route to the same inbox as well.
[1] https://datatracker.ietf.org/doc/html/rfc5233 [2] Less common in work hosted ESPs but almost universally default enabled in public ESPs for consumers.
Two or three mails have been misplaced in a decade.
It would be feasible to change something like that without breaking security now.
Google can hardly start allowing/routing a new account for first.last@gmail.com when you were getting it for years even though your account is firstlast@gmail.com and sensitive communication like say from your bank would routed there.