Oracle bans AI-generated code from OpenJDK(app.dealroom.co) |
Oracle bans AI-generated code from OpenJDK(app.dealroom.co) |
I'm shocked.
I'm not only shocked: I also see there a delicious irony in the countless of LLMish sloppy-pasta code that's now running and going to be run on man-made JVMs.
Don’t worry, I think most LLMs, when starting a greenfield codebase, don’t reach for Java, so it may be a smaller amount of code affected than you feared.
That Oracle's policies over its internal code base are different is irrelevant.
Probably won’t happen but clearly Oracle sees a potential for legal issues with LLM output.
Now slowly, the Empire strikes back - not just Oracle, but more and more resist the tyranny of AI skynet slop.
I am upset that these corporations drove up the RAM prices still. They need to compensate the rest of mankind for this - after all the chip market is a de-facto monopoly. They should all be sued into nothingness, then new laws must enforce healthy and fair competition, without unfair players driving up the prices willy-nilly style. Absolute AI mafia here.
And if I run out of tokens, well, that's OK too. K3 running on my own box will finish the job. I'll just have to wait another week, that's all.
Regardless, your timeframe on Oracle's death seems way too soon.
As we move forward it will be easier than ever to just maintain and keep your fork of software with the changes you want or need. No more approval, bureaucracy, or arguing. Just tell the AI agent want you want changed and you have it.
This will be used for huge things too. Like maybe you want a specific fork of Java that only supports for each iterators, goodby linters, hello compile time error.
Software as a list of requirements and that's it. The local LLM appliance everybody has taking in a document specifying hardware, interfaces, and requirements and spitting out software changeable locally via conversation with its users.
In the same way you have a cookbook with recipes to make dinner instead of ordering out.
I don't imagine this is practical or desirable for all situations. Good software is built from being battle tested by many users in many environments. Even with the advancements in AI tools, I don't imagine they'll become omnipotent anytime soon.
> No more approval, bureaucracy, or arguing
For software that can kill people or substantively affect someone’s life in a negative way, the bureaucracy is there for good reason. I don't think anyone should want someone at Phillips to vibe code the control software for an X-Ray machine or an employee at CrowdStrike vibe coding the next update before pushing it out to millions of machines.
We are forced to endure low-quality software because there is little or no accountability. I can only imagine what you propose would make an already poor situation worse.
Then you'll have the same problem everyone who forks a piece of software ends up having, sooner or later: as the original evolves, keeping your fork up to date with the upstream changes becomes harder and harder. The bigger and more invasive the changes are, the harder synchronizing with newer releases become.
Can't picture a functioning world where every piece of software is custom and requires factorial amount of AI comparisons and reviews to patch the API to communicate. In fact, it's impossible! There's not enough compute to handle a factorial explosion.
I really doubt SaaS going anywhere.
Even SaaS isn't safe. I don't even have to describe your product to my system, I just have to give it a harness with access to the interface and have it replicate it locally. Frankly you can probably already prompt for that.
The only thing holding this future back right now are pricing problems and code generation quality. Both of those barriers are constantly being knocked down. We might never arrive at that future, but it's definitely a higher probability than solving AGI's scaling issues, and would arrive much sooner for technical users.
Because SQLite has 10k requirements that wouldn't even cross your mind to write down, but 80% of which are useful to you.
https://www.youtube.com/watch?v=V_qzqY1bb7I your sufficiently advanced code generator may generate you a high quality database system for some measure of quality, but it will not have SQLite's reliability over the extremely long tail of edge cases proven through its testing and use in real life
Prompt: “Top language pick for an Android app? Respond with only the language.”
Responses (admittedly via the aichat CLI, not an agent harness; and temp=0):
Kimi K3: Kotlin
GPT 5.6 Sol: Kotlin
Claude Fable 5: Kotlin
DeepSeek v4 Flash 0731: Kotlin
Are there any other “inherently” heavily JVM-biased domains?
I think even in Minecraft game mods Kotlin is possible nowadays…
That's before we get into the entire non-linearity of agentic systems introducing massive decidability problems on this in the first place. A little bit of epistemic humility please.
The cherry on top is that OpenAI and Anthropic brainwashed your coworkers and your company's C-suite into uploading the entirety of the "proprietary" codebase onto their servers thousands of times per day over the last three years.
For that, you'd want the list of requirements and hardware documentation to be written in a precise, formal language. That's no different than writing them in a programming language (though a declarative one, instead of the more common imperative ones).
I've in the past (way before LLMs existed) thought about automatically generating device drivers from hardware documentation. But besides the need for very precise documentation, hardware never works exactly as documented; a human-written device driver can avoid problematic areas (perhaps even by accident), while a computer-written device driver would end up exploiting every corner case of the documentation.
From my perspective, this is a damning conclusion to the argument.
That being said, I can hear him asking Claude to generate a rebuttal as we speak.
There is quite a bit of distance between the exactness of any human language and an sort of programming languages. When you have a language model loaded with software engineering best practices you do not need the exactness of a programming language to describe desired behavior.
Hardware documentation is indeed often lacking in plenty of ways but in a world where writing your own software through agents is commonplace, the hardware manufacturers (or the community) would make testing that documentation to find the problems an important part of hardware development.
This isn't anything new or particularly interesting. It's the entire basis upon which ILP demonstrated generality. A metatheory to synthesize 10 trillion rules isn't even scratching the surface of what you can reasonably do. The key was finding out the tractable semantics for actually computing it in reasonable amount of time, which right now is looking decidedly like informal semantics was the answer the whole time.
They are trained on a narrow set of data and don't really understand the real world.
If you let them run wild with no supervision they'll turn your company into slop.
Someone needs to be there to reign them in, asking the all-important questions like "is suing all of our customers and making them angry REALLY the best option?" (to which they will reply, "You're right to push back on this").
This is typically the role of the CEO, however, as we know, most CEOs of large successful companies are too busy to actually do that. They're mostly training Brazilian Jiu Jitsu, shitposting on their own social media platform, trying to manipulate international politics, or helping run their family nonprofit.
This is why I believe that legal departments need to be demoted. The legal department should not report to the CEO, they should report to a new role that is above both PR and legal. And that person's job should be to force legal and public relations (which we know are natural enemies) to work together. Every press release goes through legal now, and it's only logical that ever legal action should go through public relations review as well.
The fact that public relations at most large companies has atrophied from disuse in our public-equity-owned unicorporate cooperation-over-competition chaebol/zaibatsu/conglomerate world is a problem for another day.
[1] https://www.cio.com/article/4125103/oracle-may-slash-up-to-3...
https://www.theregister.com/ai-and-ml/2026/08/03/as-larry-el...
The register article is about this post:
It is 'OpenJDK Interim Policy on Generative AI' and their lawyers are writing the final version, according to this page.
It sounds like a sensible action, given past scars around Java and copyright, plus this is from a big old corp.
That said I personally don't expect that final proposal will end up any better.
How many of these were self-inflicted?
Especially for a project that runs to many major businesses, this could pose a massive risk.
code is a liability, and they likely have more to lose than gain by allowing AI contribution.
java never had that phase. If anything it had a full binary class (not even source, e.g. long and int are not compatible on binary level) compatibility. Generally you can take an application from '98 and run it nowadays.
It's only recent changes (jigsaw in java9 mostly) that were made them to drop some of that part.
It's shocking given Oracle's AI spending spree over the last couple of years.
If you know how to work with it it's great, but if you use it indiscriminately it does more harm than good.
The biggest use of Java currently is Android, and Google is on a path to migrate to Fuchsia eventually.
The second biggest use is Apache Spark/Kafka, but with LLMs now, you can pretty much re implement all of that functionality in pure C.
They just want a legal angle, since this is how they make most of their money.
One also assumes the people maintaining OpenJDK have had their workload increase to an unmanageable level, like a lot of other free software projects and one thing I'm sure everyone can agree on is that Oracle certainly won't want to pay anyone more or hire more people to deal with that.
A lot of people (including the Register[0]) are pointing out that Oracle leadership are gung-ho about using LLMs for everything, and this seems to go against that. It makes sense to question this from a journalistic angle -- the executives are obviously full of shit and people would do well to remember that the next time one of them opens their mouth. But pointing at the apparent contradiction -- different rules for internal projects vs. open ones -- doesn't seem particularly meaningful on its own:
I don't think the CEO/CTO raving about LLMs should be taken as firm statements about how they actually operate internally. I'm surprised that this does not seem to be the default case. Among other things, Oracle stands to profit from greater adoption of LLM tools.
It would be unrealistic/unreasonable to expect their employees/contractors working on internal projects to be held to the same standard/guidelines/rules as developers contributing to a free software project. This applies either way, whichever side (internal/open) has the worse deal.
I don't know what the rules are for their internal teams. As far as I know, I'm not alone in that. Oracle also don't need to post anything publicly to change those rules.
[0] https://www.theregister.com/ai-and-ml/2026/08/03/as-larry-el...
> If I use a generative AI tool to create 100 lines of code, and then edit ten of those lines myself, may I contribute the result?
> No. Your contribution would still include, in part, AI-generated code.
I totally agree with you on this! The engineer is still actively engaging while getting the performance gains of not having to type out functions. The Engineer is still in charge vs agentic "engineering" a model + harness can spit out whatever and the engineer is left to review tons and tons of code
sips tea
There is also the approach of Microsoft putting AI all over the place on .NET, and CoPilot driven development all over the place.
What if someone generates code with AI and then goes into the IDE and then bathes it, dresses it (including adding comments) etc, in a way a human would? What then? No, here I am not exploring a way to fool the code DNA checking, but rather trying to find out what the real problem is with the AI generated code? (Other than license issues, too many PRs etc)
Too many PRs wouldn't be a problem if they were easily digestible: But they often include these giant refactors with confusing changes that aren't elaborated.
The real underlying problem is with authors who do not understand the code they submit. If you cannot defend the PR, you shouldn't be submitting it.
Will you submit an AI-generated pull request to OpenJDK? No? Voilà, you have successfully self-enforced the OpenJDK policy.
Or do you go to the trouble of finding an interesting issue to work on, get your agent to code it up, manually polish it to make it less AI-looking in case there are doubts, and then submit it? Yes? No. Why would you? To prove some kind of point, to yourself, that you can never disclose publicly? Most people have better things to do. Voilà, the policy is, again, self-enforced. Enjoy your day at the beach instead of trying to trick a project that is politely asking you not to trick it!
And if you catch somebody submitting code they can't explain, you don't know them anymore.
It's really simple. Open source has always had a social layer. Github culture tries to eliminate it, and that's the source of this vulnerability.
I accept @jerf's explanation is the right one but this is an absolutely almighty signal that AI-first developers should heed. And someone should ask Altman about it on the record.
The people you refer to are delusional idiots and should be treated as such.
But people at Oracle below the C suite don't even fart in public without lawyers signing off.
This is Oracle telling quite a lot of customers and partners of one of their most significant products that it doesn't trust AI code to be safe. Oracle, the company that is more exposed to the bubble bursting than basically anyone else apart from Coreweave.
This message could have been a lot shorter, could have said they were pausing accepting AI submissions until the intellectual property situation was clarified, but no — they said AI code can be unsafe and insecure and it's too risky to accept it. So a good question (given that Oracle are fucked if OpenAI even stumbles) is why?
That is a heck of a message to send, and it is visible enough. That is why I think developers (by which I mean individuals and their companies) whose focus is AI-generated code, which is almost everyone in public, should pay a bit of attention.
Also, Oracle knows that it will be stronger with strong IP laws. Other companies will realize that fact soon.
Seems like a better way to do this than ban all AI generated code, but it's probably a matter where they're far more concerned about their own inconvenience than someone else's.
I’m sure someone could still just prompt their way through but it’s a decent soft gate IMO.
I think for Java specifically, it's also very much: which prompts make sense to write. Java has had a storied history of features that made future development hard. That's resulted in a culture where they like to think quite deeply about how to add a feature BEFORE they sanction any sort of programming.
The same problems that bedevil AI-written PRs here bedevil AI-written PRs at my place of work.
I'm sure it'll get better eventually, maybe, but something something Pareto principle.
>I'm sure it'll get better eventually, maybe
Why would it get better?
Google “Oracle org chart meme” and look at image results, only it’s not a meme.
Selling it is one thing. Having to use it one's self is quite another.
Do as Larry does --- not as Larry claims to do.
I would short his stock but it has already lost half it's value over the past year. And his credit rating is one notch above junk.
AI is going to make fools out of a lot of billionaires like Larry.
Even the biggest institutional AI investors have to ban their contributors from spamming them with AI generated slop.
Spark is written in Scala, precisely because distributed systems are best modelled as a set of pure functions with minimal references to state. This is also why part of the reason why it's ascending successor, Polars Cloud, is written in a functional language, Rust, rather than C.
Android development is also increasingly being monopolized by Kotlin, not Java.
The "limited ROI on AI investment" articles will continue to percolate slowly into the brains of the LinkedIn hive-mind until we hit a tipping point, and then we'll finally shut up about how a handy dev tool with some decent use-cases is the dawning of the singularity that will replace all white collar labor and get back to actually building business value.
I haven't written a line of code myself in 6 months, but the point is that _you shouldn't be able to tell_
What I tell people at work: Using AI is good. AI can help you do things that you otherwise wouldn't have. It shouldn't be a crutch for thinking.
In that sense, people shouldn't be able to tell that you're using AI unless the tell is that it's _higher quality_ than if you had implemented it by hand.
A good example here is a well-written document that has a _ton_ of deep research behind it. Or code that has extensive tests that would have taken too long to write by hand for the task at hand.
It's bad when it's just "do this thing and post a PR" but I write my code with AI, look at it as a reviewer, offer suggestions, put those into my steering if needed, and I refine it. The way I see it, I'm the first reviewer on everything now before passing it off to another human for review.
There have been points where having AI rewrite something was slower than me doing it, but I'm already in the harness. The cost of waiting is almost zero. What stops me from contributing to the JVM using AI written code that I understand and spent time manicuring? Can a human really tell that it's AI written?
It would be interesting to see how this effects the copyleft Licences with contributors are using AI for the PRs
Refusing to use AI honestly seems more like vanity to me. Like "no machine could ever do what I do".
That said I would not necessarily accuse people who don't want to use AI of vanity - there are good reasons not to want to use it - but if someone else did it I'd find it easier to understand.
Linux's LLM policy allows AI-generated code because the project was never vulnerable to this DoS attack in the first place. They kept the social defense they've had for decades.
The killzone here is Github culture, where people decided it was normal to accept code from anonymous randoms with anime avatars. They're doomed.
Humans still need to review code for taste, making sure AI is writing sensible, well organized output. That's a core "AI skill" for now, at least until it gets better...
in any of those cases there's no reason to get swamped by cheap PRs for now
Which recent verdict?
> It would be interesting to see how this effects the copyleft Licences with contributors are using AI for the PRs
I know Linux and GCC have been diligent about tagging and tracking LLM based contributions. In the worst case they can chuck it all out and handwrite it back.
There is this old one around contents like images and all, the ripple affect is pretty much everywhere.
Its sort of coupled with the recent penalty of $1.5B on Anthropic. The catch 22 is that some portion of the LLM training data can be classified as IP theft. Even though Anthropic has been fined the data still exists and can be used by LLM to generate code for you. So if you are claiming something as an IP and it has stolen part in it then it leaves you in a hard place.
Language translations may save you in some cases, though the whole definition of cleanroom has been in debate recently too where people are trying to rewrite opensource/famous libraries in different language and claiming IP rights over them.
Settlement not penalty.
Anthropic will have judged the benefits of the settlement, not just the headline cost. It could have been a strategic move by Anthropic: we can't know without information we don't have. https://news.ycombinator.com/item?id=49014389 1.5B looks like ~2% of funding/income.
I’m not saying it’s good for everything, but we’ve gone very far in terms of capabilities in the last 3 years. Thinking otherwise will make me question others’ experience on how much they’ve used it so far.
it would be more convincing if you had concrete examples of technical merit and quality/speed improvements that worked for you or your team that justify going all-in
I manage several teams of developers who use them every day professionally and use them for personal projects privately. Just last week, at the prompting of said devs, we had a working agreement conversation about curtailing the use of AI in our codebases because of rapid erosion of our teams' ability to operate, update and maintain codebases that had started to spill over with slop.
We have had multiple incidents of credential leakage, integration tests wiping live databases, comically broken code that passed vibe-written tests, documentation and code comments that were hallucinated and/or fake, and most importantly developers saying "we no longer know how this code works but it's massively bloated and unreadable and we can't tell you with a straight face that we can maintain it or fix it if it breaks." We have seen a flood of vibe-PRs from engineering adjacent teams that suddenly think they can code shipping prototypes into production that do not work do not scale and cannot be maintained. I am personally writing the tickets to decom one of those today. Which is great, I love telling business "the progress that was reported to you was a lie, this shit never worked, don't shoot the messenger but also don't let this happen again."
I embrace GenAI as a productivity tool for people who know what they are doing. It's a +10-15% velocity boost. That's great! That's a big deal, devs are expensive, and that might push some kinds of business model over the threshold into viability. That's great!
It is not, however, transforming the profession as I know it, it is rather making my job harder and less pleasant and it is making my leadership dumber by the second.
Absolutely no comment on vibe business decisions / vibe OKRs / slop reports or the host of other garbage that has started to creep into professional life. I have had to have some very uncomfortably direct conversations with peers in leadership about using complete bullshit to make decisions, and it is very, very frustrating. Do you know how hard it is to convince someone that metrics their bot hallucinated don't exist and would be meaningless if they did? You can't convince someone of something they are incentivized to not understand. It's been very eye-opening in terms of who I can trust to actually make sense when it matters. I'm grateful for the clarity.
Meanwhile we are rapidly losing brainshare from the top because our principal/staff engineers are pissed off and have the bankroll to just leave. We aren't hiring and training younger engineers to keep the talent pipeline moving. Which boy howdy is THAT going to cost us unbelievable sums of dollars to fix in the medium-term future.
I am in the uncomfortable role of trying to make the best of this but it would be a metric ton easier if the narrative from the c-suite aligned with reality in any meaningful way.
</rant>
when i brought this up with my manager they just said "you need to engineer a better harness" and "aren't you cultivating your agents.md file? thats probably your problem"
which is to say, apparently we're holding it wrong...
To anyone reading this: If you couldn't see this coming three years ago, you don't deserve your job.
> We have seen a flood of vibe-PRs from engineering adjacent teams that suddenly think they can code shipping prototypes into production that do not work do not scale and cannot be maintained.
What a nightmare. Never forget that AI is for idiots.
Secondly the future is still not here yet, the people who came forward to sue are mostly in the category of book publishers/authors. The tech companies are are not actively searching for copyright thefts as of yet, however I am sure its just a matter of time when the big blobs of codes get rediscovered specially in case of any publicly visible code
If that's a reference to Thaler v. Perlmutter, the only thing that's been established is that an LLM can't be considered an author under the Copyright Act, only a human being can. It says nothing about the consequences of a human claiming authorship of LLM-generated code, which would be relevant here.
Thaler v. Perlmutter stands for a much narrower proposition and at any rate is not binding nationally, SCOTUS having denied certiorari.
> Based on an analysis of copyright law and policy, informed by the many thoughtful comments in response to our NOI, the Office makes the following conclusions and recommendations: > • Questions of copyrightability and AI can be resolved pursuant to existing law, without the need for legislative change. > • The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output. > • Copyright protects the original expression in a work created by a human author, even if the work also includes AI-generated material. > • Copyright does not extend to purely AI-generated material, or material where there is insufficient human control over the expressive elements. > • Whether human contributions to AI-generated outputs are sufficient to constitute authorship must be analyzed on a case-by-case basis. > • Based on the functioning of current generally available technology, prompts do not alone provide sufficient control. > • Human authors are entitled to copyright in their works of authorship that are perceptible in AI-generated outputs, as well as the creative selection, coordination, or arrangement of material in the outputs, or creative modifications of the outputs. > • The case has not been made for additional copyright or sui generis protection for AI-generated content. > The Office will continue to monitor technological and legal developments to determine whether any of these conclusions should be revisited. It will also provide ongoing assistance to the public, including through additional registration guidance and an update to the Compendium of U.S. Copyright Office Practices.
Congress or the courts could, of course, override the stance of the copyright office, but I think it would be highly unusual for them to do so (particularly for something like this). It would however be a lot better if congress just stepped in and said no outright, but until then this will have to do.
Only Congress and the courts do. Copyright exists from the moment a work is created, and does not need to be registered with the copyright office.
The law isn't that complicated; if a work was created with a human being with intent, it's probably eligible for copyright protections.
As long as you can convince a court that you did this, the tools you used are not relevant. The vast majority of LLM art falls in this bucket.
The sibling comment lays this out and my original comment above is based on exactly the same link.
> Whether human contributions to AI-generated outputs are sufficient to constitute authorship must be analyzed on a case-by-case basis
It says a plain prompt is not enough but that is not the reality of real software development. People aren't one-shotting complex business apps. The vast majority of software development will trivially pass that bar and end up in the "requires case by case analysis".
If you write a prompt and one-shot a problem and share the source code, that source code is probably not covered by copyright.
If you substantially edit or modify the generated code you would own the copyright.
It's like with a camera. If I set a camera and carefully aim it and somehow trigger the shutter then make adjustments in Photoshop, I own the copyright on that image.
If I stick a Flock camera on a pole somewhere and post the live output, there's been no meaningful human creative involvement in producing those images and so nobody can claim copyright on them.
I don't like this idea that llm code can't be owned by a human, copyrighted. It's just code.
I think your last example with flock camera is relevant here - I can take a picture of a public football as a reporter or something (or a fan I guess) and I can copyright and sell that picture. Newspapers do it every day.
So if I stand on a street corner and take a pic, it's copyrightable. If I take a pic using a flock camera it should also be copyrightable, just like if my nest camera at home takes a pic of something, I can use that.
I guess you are saying "someone else owns the flock camera" so you don't get to own pictures. What if I buy the flock-like camera and put it up, I should own that.
Maybe so, but perhaps the most successful part of the AI marketing pitch has been exploiting management's hatred of labor (and vice versa).
A significant motivator for a lot of the nonsensical AI layoffs and initiatives over the last three years has been that it's a great opportunity for management to get ill on their slaves.
Tech labor got a little too big for their britches during the hiring spree of 2021, and AI was a great opportunity to take them down a peg, even if it offered little in the way of ROI.
Don't underestimate the power of the desire to keep one's job to cause a large organization to behave towards its own self-perpetuation despite the moral lines that must be crossed by individuals to do so.
Lots!
> Don't underestimate the power of the desire to keep one's job to cause a large organization to behave towards its own self-perpetuation despite the moral lines that must be crossed by individuals to do so.
I don't. I don't put anything past people.
In other words, Larry Ellison is going to do what Larry is going to do, and there's no use wondering why.
Usually, that's to make money and sue people, at any expense. It's in his nature.
The lawnmower exists, you might disagree, but it’s not your lawnmower, you can’t uncreate it. But you can and should take care never to put your hand near it.
Lessons:
- The only way to constrain what the lawnmower does, is make sure what do desire is enforced by law (physics)
- A reminder, the lawnmower will chop your hand off and think nothing of it
- The lawnmower is not "evil" but if the laws that constrain it allow for evil behavior, it may do "evil" things.
It's more like that movie with Stephen King where the cars and other machines turn actively evil. Hilarious movie too, not great but hilarious.