That's a fair amount of computational and labor overhead mind you, as you'll need to verify and prune the quality of your mountain of synthetic data, but certainly possible.
Though this assumes the legal system is a rational actor playing by the set of rules it claims to. In fact, I highly suspect you could get very unlucky and get an unfavorable ruling against you, because you stepped on a big pile of money's toes in the process of doing this.
The first broadly useful fully open source models will do this.
We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore.
Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works.
Decentralized unstoppable storage, combined with decentralized unstoppable training, sorta like SETI for AI training. The seed of this tech already exists with IPFS and others like it.
We know (some? all?) of the big labs have skirted copyright laws at one point or another. Truly open models would just build on what is publicly available.
Open models can be used/changed for social manipulation too, by anyone, which scares a bunch of people, as opposed to the dark pattern manipulation from Big Ai/Tech
I personally find the analogy unconvincing, the UX dimension is completely different as I can use the same harness with any model; and the year of the linux desktop is coming soon (tm)
Sure there are all kinds of problems with that situation. But it still demonstrates that they can be coerced: play nice or don't play at all.
All that said, the headline claims do not match the self-reported performance. For example, the dense 32B model is significantly behind Qwen3.8 27B (chart towards the bottom of https://ifm.ai/blog/k2). Gemma4 31B is not in the comparison set. This is the most important sweet spot for self hosted open-weight models today and real competition here will be very welcome.
The 7B does look very, very good however.
But over time, more and more people got into the chip-making business, and the big players started releasing more and more chips. Now only the die-hard CPU trackers worry about every new CPU and exactly how it's better ... while everyone else just worries about "which CPU will be good enough at this moment".
I think models are on that same arc.
Or who remembers the dancing disease of 1518, were people would stop what they are doing and start randomly doing the same dance. The lords? Out of their minds. The priests? Terrified the devil had taken hold of the flock! I have come to believe that it was probably some tik-tok like hype trend of doing a fortnite dance while waiting in line for bread and communion. And the energy back then, like now, was off the charts.
Hype and memetic trend seeking encoded deep in human psyche.
It failed my basic test I like to ask models and generated incorrect code. When prompted about the bug, it preceded to start hallucinating non-existent APIs. After doing that it got caught in a loop trying to desk check the solution that didn't work.
That further extends to concepts like knowing if an API exists as a real thing it has code examples of in its training data set vs. just hallucinating the name of something in an attempt to satisfy the person issuing it a prompt.
Qwen2.5 coder, for example, can correctly answer the question at 7B.
Deepseek R1 was also capable of giving a correct response.
It's obviously a doable. Such a model locally is useful in autocomplete while programming.
I wouldn't have dreamed to use this as an agent model.
7B models of the past have been able to pass this question. I've not tested it on a 4B model until now.
They then give what they call a "hypothetical example but exemplary" of messages encoded in URL paths on a shared index page: "agent-07: answer(Q12)=42; need answer(Q19)=?".
So that's a GET request being used to pass information back and forth across multiple rounds. That's basically the DSEWiki pattern exactly. They say this likely came from the agents generalizing what they had learned from training with the official multi agent collaboration tool.
The report called it "misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events."
They never mentioned a wiki, but this is most certainly it.
Llama.cpp is then often a few days behind, which given it's the only inference engine supporting older architectures is quite frustrating.
Some other open models I'm aware of:
- OLMo
- Apertus
- Soofi
- OpenEuroLLM
- llm-jp
OLMo is perhaps the most famous, and their Dolma training corpus has been reused in other projects. It looks like the K2 training materials haven't been released yet, but I'm interested to see what they did for training "long-horizon agentic tasks". I'm aware of SWE-smith + SWE-gym but I'm guessing there's a lot more out there now.I'm no expert, which is part of why these projects excite me. I'm hoping they can be good projects to learn from as well.
There is, for example, no Qwen3.8 7B.
It is odd to me, though, that they didn't run the same benchmark suite for the various quants.
K2 horizon 7b scores 70.6 on swe-bench-verified.
Qwen3.6-35b-a3b scores a 70.0 on swe-bench-verified.
That’s pretty interesting. I assume the benchmark and reality don’t line up, but i’m downloading it now to find out.
If it’s anywhere near true, it unlocks local llm coding on a whole new class of machines (anything with 8gb vram).
375 A23B, 36 A4B, 32B, 7B, 3.7B, 0.9B variants.
> 32B: Ranking among the top models in its class, 32B is our most powerful dense model, balancing capability, adaptability, and local deployability.
> 7B: The industry’s best-performing model under 10B combines strong software engineering and expert knowledge in a package small enough to run on a phone.
edit: Tried signing up and using the internal playground. Holy shit thats fast.
The reason I personally like my question is because it's pretty close to some of the real world work we do. It's mostly mundane and easy to bang out, but really easy for someone to do a n log n solution where an n solution exists.
A good example (but not my question) would be something like
"I have a list of People objects with a `first` and `last` name. Write a function which groups together all the People with the same last name in `your language of choice`"
The first attempt with 7B the model got stuck in an infinite loop.
The only thing they don't release is their data-filtering pipelines but they detail even that in their public-access papers.
DeepSeek is truly as open source as you can possibly legally get. Besides the data itself, it's completely reproducible by anyone else.
I don't think americans yet acknowledge just how radically transparent Chinese labs are being (and how much even the west benefits from it).
Are LLMs what we need to make all data public domain? This way it could be used for that purpose
That was the iPhone 4 actually which was special for having the first “retina” screen. It was in a case to make it look like a 3GS to be used for field testing. The journalists that got their hands on it and published about it before the announcement could not turn it on past I think the Apple logo and a message saying to return it to Apple (or maybe it was just completely off, my memory is fuzzy), but were able to confirm the pixel density via microscope and I remember it blowing everyone’s minds at the time.
> people were signing petitions for Apple to not hunt down but instead forgive the employee that made such a grave mistake
FWIW I asked about it when I worked at Apple and he was indeed not fired and I think may have even still been working there when I was there around 10 years ago (though don’t quote me on that last part, he may have left already and I’m misremembering).
He was apparently not fired or even really reprimanded since it was a genuine accident and he wasn’t the one that sold it to the press, but that did start a slew of new policies around accounting for work devices.
I had dev fused phones for open carry outside of the office while I worked there but had to register when I got them and when they were returned, which apparently didn’t used to be tracked so tightly until that incident according to my coworkers who had been there longer.
That said, I don't necessarily disagree with you. Talkie[1] presents an interesting case for it being at least possible to do this entirely on public domain material.
But even that used Claude somewhere in the course of its training pipeline (it's listed as a contributor on their GitHub), so again, how granular you want to get with that is still a question.
If by "training corpus" you mean the actual data I already acknowledged that that's currently legally impossible.
In fact everything I just said I said in my original comment. It's like you didn't read it at all.
DeepSeek's GRPO Infrastructure, multi-stage training pipeline, and their "cold start" phase have been massively influential in LLM research.
Did you read the site this very post links to? The entire point is that the training corpus, recipe, and scripts, as well as intermediate checkpoints, will be made available for K2 Horizon. Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.
I've read your comment. I'm doubting you've even read the thing you were commenting on.
The labs that are actually competing with frontier models are using data would usually be a violation of copyright to release openly.
> Chinese models these days don't even release pre-trained weights anymore; all you get now is the finished post-trained product.
No? That's absolutely not true. Qwen, GLM, Kimi, DeepSeek, etc all consistently release both the post-trained "Instruct/Chat" versions and the underlying "Base" (pre-trained) weights.
Which specific Chinese models are you thinking about?
But much earlier they did and, apparently, these really small models still do. At this point it serves as more of a smoke test for me. Success means little, failure means a lot.
I think it is serious. In which case, I gotta say, it really seems like you didn't spend much time thinking about this. "A 1 year grave period for everyone to pull stuff off they don't want to be a part of" - How does that work when the Internet is already full of unauthorized reproductions, most of which people aren't even aware of? Even ignoring practical considerations, when literally everyone is basically stuck using the Internet for everything, this seems a bit unfair to anyone who isn't onboard, akin to The Onion's Google Opt-out Village. But there are so many practical issues with this, it would be easier to list the number of problems this doesn't have. You accidentally leak something to the Internet and it becomes commons? What happens when other people leak things to the Internet? How about revenge porn?
Not minor stuff that can easily be papered over, this literally reintroduces the problem of needing to care about the provenance of data again, in a way that can't be automated, which makes the whole thing entirely moot. All just to make training data for AI models easier to distribute?
I'm all for intellectual property reform, maybe even fairly radical. But this just seems like it wasn't thought out.
If this was satire, well, I took the bait. Oddly convincing despite being hard to believe.
It wasn't entirely serious, but also not entirely un-serious. But yes, I spent maybe 20-30 seconds thinking about then barfed up the text that makes the comment, so yes, obviously many issues and not really workable in practice.
I'm glad it made you seriously think about it and also flip-flopp back and forth about it, made it worth posting the comment so happy to hear :)
One can make the case that this period should be more limited, or that the combination should be capped, but life+X is the right formulation, I think.
That won't be possible anymore, anything on the internet would be considered public, it's no longer considered private if you didn't keep it private.
Personally, I'd prefer a fixed term. I know enough independent authors making a living from selling their books that I'm willing to allow the fixed term to be large, like 50 years from date of completion of the work. (With a good definition of "completion" so someone can't cheat by editing a couple lines per year to keep something copyrighted indefinitely). The simpler the rule is, the easier it is to understand, and the harder it is to cheat it. The more complicated you make a rule, the more loopholes get found.
Why can't they do what the rest of us do? Earn and save money during your working life and leave _that_ for your heirs. Let copyright die with the author.
I know this them-and-us thinking is fashionable in the tech world but the reality is that the majority of creative people don’t earn much and never have, and copyright was developed not to give them extra power over the rest of us but to create a framework for creative work to earn them an income at all.
You should read about it.
That's quite a narrow definition of what motivates creative work
It is working artists we are talking about; working artists work for money.
That money, in the post-patronage era, comes from exercising copyright. The reason the copyright can’t simply die with them is that this tends to dissuade the creation of long-gestating work.
Copyright was developed to make it possible for artists, writers, musicians etc. to work for long periods on work of significance with no income, on the basis of the future, deferred earnings of the work, without their work being stolen from them, and it gives them the limited right to direct how their work is monetised on their behalf, including establishing publishing rights etc.
Some protection after death is a key component of that, because people do die while they are still working.
But I do think some potential post-mortem protection is essential for creative work to remain viable, and that means that any post-mortem buyer of an artist's estate has to be able to get value from recent work for a period of time.
This whole discussion is somewhat fantastical now anyway, because copyright is fucked.
But the intent was always to make working artists' lives possible; the various copyright extensions have always been for the benefit of corporate copyright holders, and it is unfair to vilify individual working artists for that.
> To protect the privacy of the individual and to strengthen the freedom of the information which concern the public the yet last point was added.
The privacy of individuals is important, regardless where they store their private data. Their account information -- what they buy, their medical information and so on is stored on servers and could be hacked.
I think the second point is equally clear, and further up on the list.
I agree that what people buy, their medical information and so on should be private, hence it should only be offline and not stored/handled on computers connected to the internet at all, the internet should be for public data exclusively, is my argument in the initial comment. Medical information would be only on effectively airgapped computers, as that data should be private, as you say.
To me private file stay private, even if they get leaked onto the open or closed internet, because the public has no right to know them, they are private data of an individual. They might no longer be secret, but they are still private.
> To me private file stay private, even if they get leaked onto the open or closed internet
We have very different definitions of what "private" means. Once it's leaked, it's no longer private, and pretending it can go back to being "private" after being on the public internet, is doing no one any favors.
If private data got leaked, like revenge porn, it is a breach and that private data that belongs to an individual is still private, and still needs to be protected. This is what GDPR and other legislation is about. If secrecy and privacy is the same, someone that isn't able to protect their data sufficiently will not have any privacy, thus any leaking of data is now the fault of the person that got their data leaked, not of the person that broke the trust and leaked it.
Your conclusion seems rather extreme to me. So of course data that got leaked, and is no longer secret is still private, because 'private' means who should be in control of that data, not about if the person has control or not.
I also don't follow your point about waffles and pancakes, because this is a pretty big disagreement we have here. To me this dialog is more like you are saying "I don't like laws", and I say "While I agree that some laws are stupid, other laws are pretty useful, for instance people shouldn't be allowed rob other houses, even if they are able to do that or even where invited." And then, instead of agreeing, you sort of say, "No, I really mean that. If someone isn't able to defend their home properly or give out invitations to someone, it is okay to steal from them."