Google is still releasing a lot of llm architecture research. They introduced speculative decoding of LLMs in 2022[1], then released the code to perform sceculative decoding for their Gemma 4 model this year[2]
[1] https://arxiv.org/abs/2211.17192
[2] https://github.com/google-gemma/cookbook/blob/main/docs/mtp/...
Qwen 3.6 shipped with working MTP first, and had working MTP in llama.cpp first.
Revealing optimizations similar to these would pretty much reduce their competitive position.
I suspect their tune will change if they ever take the lead..
That's a lot of words to say it's just capitalist greed.
Hopefully the experts here can offer insight. The above is just my hunch and I’m not a specialist in this field.
Multi-head Latent Attention (MLA), Multi-Token prediction, MoE architecture are some of the most famous examples.
It's more a cultural thing. Sharing progress is just in their blood.
They don't have TPUs or access to the latest Vera Rubin GPUs either to get performance gains for free. All of the optimizations Deepseek have done are in software and it goes down to the PTX assembly level.
Compared to Anthropic who are celebrating in fixing a flickering issue in a terminal app which took months to fix.
DeepSeek are still using NVIDIA (PTX) to train on, but for inference have already transitioned to Huawei Ascend chips, and inference speed is what this paper is addressing.
It's funny, because if you ran Claude Code on a slow terminal, the cause of the flicker was obvious: They kept dumping the entire history of the chat back into the terminal in a number of situations, and relied on the terminal to them end up in the correct state.
What became clear when DeepSeek came onto the scene was that China was seeking to commoditize LLMs. They consider it an issue of national security not to be beholden to US tech companies when it comes to AI. And I, for one, fully endorse this policy.
Another data point on this is the black market for Claude tokens in China [1]. The chat logs themselves are a commodity to train models.
I believe that OpenAI in particular is a bet on a trillion dollar pot of gold that doesn't exist. Google, Microsoft, Amazon and Meta will all be fine. Anthropic is in a far better position than OpenAI (IMHO) but if DeepSeek or some other Chinese open weight model gets as good at coding, they're in real trouble too.
anyone with IQ higher than 130 (thus qualified for actual AI R&D) would be questioning something obvious here -
if they are already doing such dodgy stuff with the aim to maximize profits, why would those resellers have large amount of logs with actual American model responses to sell to those AI labs in the first place. shouldn't they just post train & customize some leading Chinese open source models to pretend to be Opus or GPT for the vast majority of their users (as classified by some models) who don't know much about expected Opus behaviours & not skilled enough to tell the differences?
that is actually the interesting bit not covered in your censored version of the story line, it is also what happens on the ground. your censored version of the story implies that those dodgy resellers using stolen credit cards, pooling accounts with stolen IDs and illegally selling very personal logs would somehow be honest enough to spend extra $ to ensure their victims (aka paying users) can actually use real Opus and GPT. LOL
dude, you failed this IQ test miserably.
I genuinely wonder how it feels to be working your whole life, actual flesh and blood and heart and mind pouring 40 to make something that is a dead-end on the tree of human progress because it’s miserly masters are terrified of sharing knowledge.
Days and nights spent playing pretend human pioneer, when you are a lunatic on an island building towers of coconuts.
https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillati...
We in the United States will never forget!
For all the harm Trump does to the US at least he is helping China!
So many Americans seem to (at least in theory) be ready to sign up for this ongoing confrontation with China. Does anyone think it isn't America who is poking the bear when it comes to the Thucydides trap? Why not try to get along? It occurs to me the only people more Chinese innovation would hurt are the mega cap class in the United States. Elon Musk certainly doesn't want BYD in the United States. Same story all the way down with these super capitalized AI companies. Most average Americans would probably be better off in a world where the United States and China got along. But its those Americans who will be called upon to suffer most of the burden if that trap ever springs.
Why not talk about how China shut out American companies for decades before complaining about BYD?
As an Indian immigrant, the PRC China has engaged in conflict with almost all its neighbors and stated wars in its short history.
China is not so benevolent when they get to the #1 spot:
https://m.economictimes.com/industry/renewables/china-wto-co...
https://yipzap.com/anthropic-accuses-alibaba-of-largest-ai-d...
Imagine if your casio calculator would come with a ToS that says you can't use it to develop a competitor calculator or any other tools. Or that your hammer can't be used to make other tools. Or, closer to the HN crowd, imagine MS in the 90s saying that you can't use their OS to build competing services to MS. They'd be laughed at and be split immediately if they tried that.
The only thing they can do is to refuse serving tokens (and even that's debatable, if we get to tokens being commoditised). But that's gonna be a game of whack-a-mole, and they know it.
And certainly they have no idea whether these outputs (assuming they ever existed and it wasn't made up) were used for training. The article mentions that DS made 150k requests. This isn't much and might have been just an eval or a benchmark to compare their own model against. It's really hard to believe DeepSeek had any Claude outputs anywhere in their training schedule, since it's just too different. Besides training on random vibecode of course, which is mostly written by Claude.
There's no "leader" if, absent someone whose results you're copying, you are an emperor without clothes
Flash: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark
Pro: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark
Excited to see if this makes it into DwarfStar for local inference, have been using the flash model extensively since the 2-bit quants were made available by antirez.
I second ccusage, it's nice
> Local-first session search, analytics, insights, and token use statistics for coding agents, supporting Claude Code, Codex, and more than 20 other agents.
solid piece of software
It's drastically reduced my AI spend. I went from spending $40/day to $10/day.
I've heard others say that Deepseek tends to be smarter on specific problems but that Mimo tends to more well-rounded.
> As with V4-Flash, we treat this point as an indication that DSpark sustains useful throughput under an interactivity target that the baseline cannot efficiently support. At matched system capacities, DSpark delivers 57% to 78% faster per-user generation.
Reminds me of the flawed solution in scaling servers in 2017 that use memory-intensive technologies by adding even more servers to solve the problem. (It just increases costs.)
Rather than doing that, think about which critical parts of your app can be written in a more performant technology.
Fast forward to 2026, now you can see who is just throwing more money at the problem to create even more problems where as DeepSeek is giving us optimized solutions.
I know exactly who I would pay attention to, and it is absolutely not Anthropic.
The last year has shown that’s not true anymore (even for web servers).
Guessing the timing isn't accidental. Demonstrated openness vs harsh regulation
Strange timeline, though this only works because it’s aligned with Xi’s goals.
Mistral...don't fumble this
This paper seems to be an improvement to speculative decoding but I haven't read it yet.
The state-of-the-art nanometer are impossible to achieve but if you have infinite solar energy during business hours does it really matter? Every company has a parking spot so this ASIC-like appliance could be as big as a shipping container.
If it could just run recent open models for a handful of users it would be such a nobrainer to buy.
this is definitely where things are going. the enormous "eat the world" models have extreme diminishing returns by comparison.
It should improve performance on most hardware because most LLMs are memory bandwidth bound during decode.
Can't sell their SOTA models, only slightly better than the open source models for the models they can sell, cost 20x to 50x for good models, a TAM that consists almost solely of developers, with no customer of theirs actually boasting increased profits as a result of AI...
I fear their time to IPO may have passed.
If the business model requires hundreds of billions to get the required quality (R&D but also infrastructure to collect data and train, either purchased or rented to 3rd party) while "only" dozens of billions can be earned back (as costs still exist to earn, it's not free once models are trained), then maybe there NEVER was nor till be a good time for an IPO in a rational market.
Unfortunately the market is often not rational in this way.
Hype within retail market means there are suckers willing to buy. Institutional market knows there are suckers when the hype is high. Both would drive the price up, and retail investors the ones left when it falls.
But I vote for these heroes with my wallet. Just yesterday did again.
A&O no longer have the most to justify their high valuation. The only thing they can do now is to get the government forbid the Chinese models.
I'd also include the other Chinese labs like Moonshot (behind Kimi) and Z.ai (behind GLM). They are innovating and continue openly sharing their research to the public. I believe the founder of Moonshot even shared 40 minute video on Twitter where he goes through techniques that powers Kimi.
More and more I suspect Z.ai just has deeper pockets and access to the Claude traces while DeepSeek is punching way above their class.
They compete with each other by innovating. The innovations result in more utility for the customer, but the technology isn't made public. Trade secrets are secret for a reason.
The reason people may think that DeepSeek is the "most innovative" is because of what they can observe from the outside, much like people may mistakenly conclude models are the "prettiest of the population" because not everyone is photographed for public consumption.
The strategy for the most companies in the US has been for a long time to capture the social audience, whatever the mean is. Quality and innovation is the second factor. Capture the market, lock in the users, influence regulation and lobbying to keep the power.
“Buy every new players threatening their business” should be at #3 in your list.
To compete in that direction, USG needs to learn from CCP to "seize the means of production", which they are sort of doing, but in such an incompetent way that I'm afraid we will probably end up mixing the worst of both communism and capitalism.
In this case, it feels like they are just funding multiple independent pure research projects and letting the chips fall where they may.
Doesn't even really seem like Europe can coordinate that.
No they don't. The U.S. Government is free to launch their own AI labs if they wish -- and even compete with the private sector -- but that doesn't mean they have to confiscate existing investments and capital. But Congress is unlikely to do that, because we've learned in the course of history that in well-functioning competitive markets, publicly-operated services tend to be worse than private ones across multiple dimensions.
Chinese companies are largely where they are not because they're state funded, but because they operate in ways that would be considered criminal in the U.S. If they didn't constantly trespass on OpenAI and Anthropic to try to achieve product and technological parity, they would be too far behind to produce innovative research.
Especially since your 5-day-old account is sus, and thus likely not yet proven not to be a Chinese bot
You can't lead by following the actual leader LOL
The only real innovation I've seen from Deepseek is the out-loud reasoning thing in R1
Did i mention there are only so many memory makers and they are all busy printing money with HBM memory?
Intel is trying with Crescent Island, to make a 160GB GPU that uses LPDDR5X memory.
HBM takes multiple times the resources to make vs basic DDR5 memory. So by going this route, you have more memory, with the disadvantage that its only 700GB/s. VS HBM pumping out Terrabyte numbers like its nothing.
These cards is reasonably priced, may be good alternative to $10k 96GB Nvidia Blackwells... You give up on token generation (heavily memory dependent), for more memory to run larger models at home/office/company servers.
The problem is, again, there are only so many memory makers and its not like the market is flooded with DDR5 memory anymore, as the big 3 moved a lot of production to HBM.
Another approach is Sandisk making HBF ... Flash memory, like your typical NVME but designed around maximum speed. So instead of loading the models into expensive HBM memory, you use the benefits of density in Flash memory, to offload models into that. Cheaper, but slower... But it leaves your expensive HBM memory free for things like KV Cache, Active parameters, etc... So your model will be slower, but your hybrid using it. As in, faster then running a model from system memory with normal DDR memory, but not as fast as HBM.
So yea, there is a lot in development to reduce the dependance of that resource eating HBM memory. For the wafer cost of 1GB HBM, you normally got 4GB normal memory. That is why the world supply of memory dropped. Not just the insane buying but be HBM is just very inefficient in wafer usage.
Can we not use DDR4 production and create some kind of hybrid solution? Sure, but the big 3 moved away from DDR4 in favor of DDR5 a long time ago. We have competition from China with a mix of DDR4/DDR5, but they also need to scale up. Nobody expected to see a large part of the world production vanish into HBM...
Even if its about DDR4 and older nodes, ironically, most companies had been moving away from DDR4. There is only so much wafer capability in the world, to the point that companies are moving to using DDR2 ... Yea, not a typo, like 2007 DDR2! for IOT devices etc, stuff that does not need fast memory. Because even DDR3 got too expensive for them.
Its not like the old nodes are not used anymore ... Like that capacity was sitting idle. It was still in production making other stuff. The only real solution is that we need more fabs, and those take years to build. And the big 3 delayed investing in new fabs for a long time, unsure about the whole AI bubble stuff. Aka, they did not want to make a ton of fabs to end up with over capacity if the AI growth collapsed.
More Crescent Island scale up, although not likely entirely linearly.
But all GPU inference work like this, it’s not specific to Intel. Just Intel promises more affordable cards with big memory so they’re attractive.
BTW, in case there is confusion, we are not talking about CPU speculative execution affecting model inference at all, just about this specific technique: predict tokens via a smaller drafter model, then validate them against the main model in batch.
Well I can't think of even one at the moment, to be honest might be biased but all Chinese research labs are largely oss except Alibaba now.
I am certain there are lots of American labs that claim to do it, but either they are marketting in hype since they aren't even close to the frontier or contrarily just don't make anything of significant value public/oss.
But the word pioneer comes from French pionnier, literally “foot soldier”, a soldier who goes ahead to prepare the way.
If you don't publish you may be advancing, but you're not preparing anyone's way.
Ultimately though the real explanation, I think, is Google doesn't care since for their own purposes (in LiteRT-LM), they do bundle them. As far as I know, anyway.
They are more like a single model that has two separate attention head mechanisms.
I wish people knew more history before using the term “historically”
The Chinese civilisation has been one of the most long living ones with plenty of “innovation” throughout history, both long ago and recently.
Recently, maybe. I see a lot more copycat behavior however.
MTP is from Meta
Another DeepSeek advance that the west are copying is DeepSeek Sparse Attention (DSA)
[1] R. A. Jacobs, M. I. Jordan, S. J. Nowlan, G. E. Hinton, Adaptive mixtures of local experts. (1991)
[2] M. I. Jordan, R. A. Jacobs, Hierarchical mixtures of experts and the EM algorithm. (1993)
[3] L. Xu, M. Jordan, G. E. Hinton, An alternative model for mixtures of experts. (1994)
[4] S. Waterhouse, D. MacKay, A. Robinson, Bayesian methods for mixtures of experts. (1995)
[5] N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, J. Dean, Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. (2017)
As for the rest of it:
It's not in the interests of the vast majority of either country to engage in conflict. American citizens should therefore resist the professed bellicosity of leaders who will never see any consequences, on a battlefield or otherwise. The average American will not be so lucky. Chinese innovation in a world of cooperation need not be less beneficial to Americans than the reverse.
More likely is that an AI generated codename is impossible to fix by humans, and SOTA was not able to figure it out until now.
There is a meteor headed towards all this AI investment that I don't think has been properly accounted for and that is, what happens to all the existing hardware investments when NVidia's next architecture comes out. Blackwell (H100/H200) is the current generation. Rubin (R100, presumably R200) is the next and arrives soon. Now a lot of the investment hasn't been spent yet so will likely be spent on Rubin but at that point, what happens when the next iteration comes out and does 3-4x the compute for the same electricity input and same hardware cost?
Also, what happens when people can run way bigger models on consumer hardware in 5 years? The effective limit for useful local LLMs is currently ~31B parameter models because the RTX 5090 has 32GB of VRAM and Apple's shared memory architecture, which can keep bigger models in memory, just doesn't have the raw processing power.
Anyway, why I argue Anthropic is in a better position (than OpenAI) is that they seem to have captured a market that may well be profitable for them as a company, specifically Claude for coding. So they just haven't burnt quite as much cash as OpenAI so aren't in as deep of a hole.
While I think local models are going to improve maassively over the next few years, running them in a data center at scale is always going to be cheaper for a company. Why? Because they can amortize their costs by running 24/7 and powering them and cooling them is simply cheaper at scale when you're talking about 1000+ engineers who otherwise might only be using their hardware ~40 hours a week.
IMHO Google is in the best position here of all the US companies, even though their models aren't the best, because their data centers are ruthlessly efficient, their homegrown TPUs will eventually catch up (and thus avoid the NVidia tax) and they simply haven't bet the farm on winning AI.
However, Google probably won't catch up. Nvidia has been winning in spite of the fact that their hardware is general purpose rather than tuned for inference.
Rubin has architectural differences I don't understand that are supposed to make inference much cheaper and faster while still retaining those other more generic capabilities. Their next generation after that is going to do even better at being fast for inference and general purpose.
Google is betting that their TPUs won't depreciate faster than the markup they have to pay to Nvidia. I don't think they will be right.
A100s are ~7 years old and going for more than 2 dollars an hour, significantly more expensive than even 2 years ago. This is because anything with 80gb of VRAM or more and made by Nvidia will have economically useful lifespans of like, 10 years.
I could see H100s getting 12 years.
Micheal Berry doesn't know shit about GPUs.
Now jump ahead 2 years and you seem to have a massive jump in performance [1]. The tokens/Watt goes up by at least 2 orders of magnitude. And the B100 is 3-4x that. And we're about to hit the R100 (Rubin) cliff.
That's what this is going to come down. When hyperscalar DCs are getting to Gigawatt power usage, it all comes down to power efficiency. Those A100s aren't far from being sold for scrap.
I've been looking into how different companies are handling depreciation for this. Amazon seems to be saying the life is 3-4 years, Google 4-5 and Meta is saying 8+, which I think is wildly optimistic.
[1]: https://lambda.ai/inference-models/deepseek-ai/deepseek-v4-f...
It’s like with VPN providers. Is Mullvad actually collaborating with law enforcement? They very well could be. It is a calculated risk.
Is DeepInfra actually logging and training or selling the logs? They could be.
So, despite hiring the cream of the crop of math graduates, who could read the papers of free academia, but whose own result the free world could not access - they fell behind.
I have a theory explaining why. I think it's because science is an interactive process. NSA cryptographers could read papers, but they couldn't talk openly with the authors of those papers, because of secrecy demands - even asking question might indicate what they were working on. You can easily imagine them spending months on something they could have avoided by going to the original authors and getting told "Oh, we tried that for a long time, it doesn't work".
Whether that theory is right or not, cryptography is a concrete example of a domain where public research with fewer resources beat private research with a lot more resources.
Is this happening? These open models have been a generation or two behind the closed models for quite a while now. They've been keeping pace but clearly behind.
The American companies, from my impression don’t involve themselves with such lowly “hacks” because they have so much money to just push forward with doing everything on big heavy models that run on the most cutting edge nvidia chips that they can, the moment, kinda sorta get on demand (I say that in some degree of jest).
They don't develop them because they don't collaborate publicly anymore.
Where would the whole industry be if Google never allowed publishing the transformers paper?
It's not a coincidence that the American AI industry grew fastest in capability when it was the most open.
How do you know they aren't doing this stuff? Something has to account for them leading the industry.
They have been raided multiple times, tons of audits, does bleeding edge research on privacy preserving tech, donates to GOS, etc etc. You don't see this kind of VPN company at all because none exists.
Model outputs don't qualify for copyright. They aren’t patented. They aren’t trade secrets - the companies sell them. They aren’t trademarks, obviously. They are nothing, actually.
(and if you argue the US models do produce copyrighted works, then oooops - whose copyright is it huh?)
Don't even try to combine it with any notion of "leadership" then, however, since distillation is literally "copying the actual leader"
1) The CEO himself 2) Tencent 3) CALT (the battery company) 4) NetEase (internet/media company) 5) JD.com (ecommerce) 6) Chinese investment firms
What are they expecting in return? I'd say the same thing that all those investors in OpenAI and Anthropic are expecting - profit.
[0] https://finance.sina.com.cn/stock/vcpe/2026-06-11/doc-iniazi...
Whether that now changes, who knows. It appears the CEO will remain the largest investor/shareholder, though, so I'm hoping not much changes.
https://www.oecd.org/en/data/dashboards/magic-database-indus...
Which will likely help them bolster the sales of the MANY new AI chips in development/use in China to international markets. Dislodging Nvidia.
Kinda the opposite of what Jensen Huang (Nvidia) thinks US is doing: https://www.youtube.com/shorts/u3SY8nvjhQA
Edit: I'm a fan of deepseek and believe it's good to make the technology open/available. And do think that also help business - which I support as well.
Edit 2: No idea why I'm getting downvoted. That's also their official stance https://english.www.gov.cn/news/202601/08/content_WS695f1b55...
???
Profit!
Not suggesting this is it, but you know, one possible angle.
Until that happens this remains an unsolved problem, so my job is to take the description of what needs to be done and find which code is on the cold setup path and can just be some PyTorch or whatever the ML researchers can write themselves, and also which part of the algorithm is where all the FLOPs are. As things get more performance critical and run more, I look at the code closer and closer. In the core of the hottest kernels, where most of the work happens, I might be placing individual instructions by hand, or going even below that and thinking about cache behavior or power characteristics.
A good performance engineer is capable of doing this while also being able to find places where they can automate this process. And there are a lot of things you can automate: layouts, schedules, pipelines. There's a lot of work going for compilers and profilers for all kinds of accelerators. Some of these operate on the "assembly" but there are all kinds of assembly. Some of these tools do almost everything for you; some are a very thin layer over the code they generate. You can see this in the interview that was linked above: it's an assembly optimization task, but you will get better results (in the time provided, at least) if you do compiler-like things. IIRC the assembler already operates on named values and in my submission I had extended the instruction selection algorithm to pack bundles based on hazards.
integrating your own work with the latest public advances takes resources. For one or two small changes this is manageable, but the further you diverge from the public, the cost of maintenance rises exponentially if you want to continue to integrate public advances. when you publish your meaningful advance, you offload the maintenance burden onto everyone else (and they only have to pay a linear cost rather than an exponential one) as it's integrated by default in new work.
In most cases, the (exponential) maintenance cost of integrating public advances with secret ones exceeds the value of the public advances, so most that undertake this strategy of advancing the open frontier in secret don't attempt to integrate continually, but instead try to make a breakaway sprint in isolation to grab a few sticky customers before the unstoppable wave of the public frontier catches up.
This is a pattern commonly seen in university research departments when researchers switch into product development mode, most of these projects are a sprint to advance away from the public frontier once a good idea is found and they do good work and find a few customers for a little while. But if you check back in a few years you won't find an advanced research department but a zombie IP company that brings in a steady income via IP enforcement and a small number of customers for whom switching is too expensive.
And regarding the dataset:
> Unlike most OECD databases, which rely on government data provided at country-level, the OECD MAGIC database uses firm-level data. The subsidy estimates included in the database are based on raw data obtained from firms’ annual reports, financial statements, bond prospectuses, IPO prospectuses, etc. The data are collected and verified manually by the OECD to maximise accuracy, consistency, and comparability. In some cases, additional information is also obtained from government databases, either to verify the firm-level information or to complement it. Care is taken to avoid double-counting where the data mix corporate and government sources.
One would have to conclude that there is little correlation b/w openness and progress speed. Sometimes open is faster, sometimes it isn't.
Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think. Normally it's exposing fraud, but here we get the really fortunate side benefit of what could eventually amount to the most significant contribution to the general software community since Linux.
Its revealing that they always seem to publish after some big announcement by American AI companies. But regardless, this is one of the benefits of a duopoly.
There's always an announcement by one of the frontier labs.
Wikipedia is altruistic, and serves humanity quite well.
Any individual that provides free labor cannot survive off of said free labor. He must work for money to survive or get donations from someone who earned that money from incentive based labor in order to even buy the food he needs to exist as a living human being. Much of the time that labor is actually closed source.
This is a logistical reality. A lot of open source advocates are unable to get their brains out of the whole mentality that open source literally cannot exist without incentive based software supporting it. Who pays for GitHub to exist? Who pays for the food swes eat? I just code for open source all day and money falls out of the sky.
My smart friend says there are jobs that pay you to work on open source exclusively. Smart guy. In this case you follow the money trail. How does that company get enough money to pay a guy to work exclusively on open source?
[†] Another problem with altruism: we don't all agree on whether a goal is altruistic, and what's altruistic in the enactor's eyes might not be in yours. Curating a fountain of human knowledge like Wikipedia? Probably altruistic. Protecting humanity from itself by installing your company as the stewards of frontier LLMs? Not so altruistic in my view.
Markets don't run on altruism.
I say this because we see the same thing used as an argument against China. "If they overtake us, they'll do imperialism (like us)." Again, it says more about us than them.
A better reading (IMHO) Of the situation is that China believes that AI shouldn't be used simply to mint a few more trillionaires but the benefits should be shared with society. Why do I say this? Because we now have 70+ years of China doing exactly that. The transformation in China all the way from rural villages to Tier 1 cities has been utterly astounding. China has lifted ~800M people out of extreme poverty.
In some ways we're at a similar point to the late 1990s and 2000s when Microsoft execs complained that Linux, being free, destroyed intellectual property value. Linux should be a perfect example of how people can and do act altruistically, or at least not in a way to bait-and-switch to enrich themselves.
[1]: https://www.reddit.com/r/AskHistory/comments/1d26grm/in_the_...
US labs in Google, Meta and SpaceX are not leading, none of them managed to build something on par with GLM 5.2.
Care to explain to me why they still don't collaborate and still choose to do it in private?
Meanwhile we in the US are blocked from buying Huawei GPUs and retirees are boasting about the nvidia in their portfolios.
Even if they're ahead they don't have enough GPUs to scale. Open sourcing is hence a good strategy to at least get market share (even if not $).
Genuine question.
Meanwhile, Xi Jinping has published his 5th book on how governance in China works and what they're after. These are not books written for a western audience: they're compilations of speeches that he already gave to the Chinese party and state apparatus, so the contents are not sanitized for foreign audiences. But there are no English reviews of summaries of this 5th book at all by the usual China experts that distribute what western audience know about China.
This extends to beyond the government. Even though "for the people but only against the government" is an often-heard mantra, nobody seems to listen to what Chinese AI companies themselves say about why they publish open models. DeepSeek and GLM have said multiple times publicly what their motivations are, yet people on HN still speculate like they usually do.
Truly mind-boggling. I get that a lot of people don't like China. But setting aside the question of whether their dislike is justified, it would at least be rational to properly understand China, even if it's to defeat it. And listening to what China says themselves is absolutely essential for proper understanding. But people don't bother to? And they seem mostly happy with sticking to speculations that match preconceived notions, even if that hurts their chances of defeating China.
If you simply take what the Chinese government says at face value, you will be correct way more often than 95% of Western policy wonks, media talking heads, "analysts" and so forth. Because, like you say, they tell you everything they're doing.
In the recent US-China summit, Xi Jinping just came out and used the Thucydides Trap metaphor, which tells you everything about where China thinks it is and where it sees the US going, which is to become increasingly belligerent as their power declines. Now whether or not you agree with that assessment (I do agree), it still tells you China wants to avoid open hostilities, it sees itself as continuing to rise and it fears what a declining US might do.
...you've almost achieved enlightenment on the nature of the majority of HN comments: vibes-based, off the cuff braindumps, where an idea is examined as it is being typed. It's great for tech and software discussions where many commenters have good knowledge or even mastery, but on "exotic" topics, an incorrect take can be voted to the top because it sounds right by affirming the biases of the majority. If you're a practitioner in a non-tech field and a topic in your field comes up for discussion on HN, be ready to be disappointed - and be ready to question all the other correct-sounding comments in areas unfamiliar to you.
In communist or post-communist countries like China and Russia, the percentage of government workers is extreme. To them, all social action is political action, and that includes economic. Since there is only one party allowed, any economic action (and thus political action) which threatens the CCP is unacceptable. They leave other private companies alone.
The CCP is bad and this conclusion is justified not only in theory as a distant observation of ideological concepts, but from their behaviors around the world which echo some of the causes of World War 2. That doesn't mean Chinese AI companies are bad, but the CCP will certainly find ways of using it for its purposes. For now, they're quite far behind on AI, but they deserve credit for optimizing for some use cases which masks poorer generalization.
I copied it into DeepSeek because I figured who's better to teach me about greatness of Chinese government policies if not the most popular Chinese LLM?
Anyhow, it must have detected _something_ in your comment because Chinese censorship policy kicked in and DeepSeek refused to talk about it. Funny because I would wager the overall sentiment about China's abilities to govern in your comment was positive but okay.
I literally asked it "expand on it so I can learn more about Chinese policies" and that was enough to get censored!
Anyhow, after saying few times "huh? but I want to learn what's great about Chinese policies!" it finally gave me response in... Mandarin. So I asked it to provide me that information in English and... it refused. Talk about difficulty of finding any materials in English ;-)
After starting a new chat and using plenty of positive adjectives to make sure I don't want to learn a single bad thing about China, I finally got a list. Looks like China is 100% successful in everything they do! How neat!
So I said "That's cool. Does this set of policies have one name? Can you recommend any books about it? In English of course" and...
...request denied again.
I mean, maybe it says more about how horrible DeepSeek is for this kind of research but boy, it was so ironic I now have stack of iron at home.
Russia, China and Iran all make public statements as if they abide by international law and everyone else are the law breakers, while their measurable actions are shockingly contradictory.
It is their actions that have caused the US reactions, but many people present them as if they occur in a vacuum.
The AI race sits within this context, with a constellation of concerns that most people do not think about. Of course the AI engineers have their own motivation, the companies will share some of that motivation combined with their business trajectory and governments will get involved with it for the justifications they see.
Or because they're human and that's what humans have always done. If the US is no longer a check on China, what will happen to Taiwan?
Frankly, you seem to be arguing that the US is somehow uniquely bad when, in actuality, The US has been hegemonic during a time of incredible peace and minimal imperialism.
The issue is one of incentives. The US needs cheap foreign labor because of deindustrialization policies in the 60s and 70s. These were arguably passed as a check on labor power since socialism was still looking potentially ascendant at the time. Whatever the reason though, the contemporary US is reliant on keeping foreign wages down and domination of the oil trade. Imperialism grows out of this need for external resources to maintain economic growth. It would be less relevant to us if we had better fundamentals, but we traded those away to avoid letting certain demos get wealthy and powerful.
China's play is more mercantile. They benefit most from stable trade conditions. They get richer the more customers they have. They benefitted massively from becoming an industrial trade partner with the US in the 60s and 70s. Because of this, they have completely different foreign policy objectives. All they need to do to win is normalize relations and build trade infrastructure. Its way cheaper than imperialism.
But western politicians keep raising this metaphor. So at some point they're like "okay we'll speak your language". They then used this metaphor to make the point "our rise isn't the threat, your fear of it is. If you resist it you're walking right into the trap Thucydides warned about". So your conclusion is still right, they don't want open hostilities, a stable world is in their interest.
Then western media ran away with this and were like "OMG Xi mentioned the Thucydides Trap", completely ignoring his point.
For something shorter, you can see Arnaud Bertrand's recent review. https://arnaudbertrand.substack.com/p/the-book-the-west-refu... The review is behind a paywall, but not expensive.
If you want to read policy documents directly (primary source), try the State Council / Chinese government policy database: https://www.gov.cn/zhengce/ and https://sousuo.www.gov.cn/zcwjk/policyDocumentLibrary
They also provide official translations: https://english.www.gov.cn/policies/
For Central Party documents: https://news.cn/politics/zywj/. It lists recent Central Committee / General Office / joint Party-State documents, e.g. 2026 documents on township duty lists, Party member development rules, carbon evaluation, long-term care insurance, and SOE leadership rules.
I think the author wrote a twitter post with a summary of the content, and someone on twitter who had read the original Chinese source also chimed in with a summary
The War on Drugs had the purpose (not just in its origin but in its perpetuation) of inflicting harm on elite-disfavored subsets of the population that could not be openly targeted for Constitutional reasons, which is about as far from an altruistic reason as it possible to get.
For the world. Compare the those 70 years to the previous 70 years. Regime change and intervention is significantly better than full scale invasion, total war and colonization of other people.
> China's play is more mercantile.
Because they are held in check by US power. Remove the US and China is taking what it wants.
I mean maybe, but China's industrial strategy isn't well served by war. I just don't see the structural incentive.
In Netherlands, the government wants to reduce emission, so it incentivizes people to isolate their homes better, and to use heat pumps instead of gas heating. On the other hand, if you actually try to install a heat pump, you'll run into all sorts of regulation issues: the unit can't be too big, there are only a few specific places where you're allowed to place it, a ton of people can object to it, permitting takes years if at all. Oh and if you isolate your house, then voila, during the current heat wave it's a constant 35 C in your home. So you try to install AC and you run into permitting/regulation issues. So you then use a super inefficient portable AC that just barely lowers the temperature by 3 C and uses 4x more energy, and that's fine. facepalm
And the government and banks also want to combat money whitewashing, so they incentivize people to use digital payments and discourage cash. Police could look at you suspiciously merely for having too much cash on hand. On the other hand, NATO and also a bunch of government agencies are warning about war and encouraging people to have lots of cash at home for emergencies.
"They" do not "clearly" want one or the other. Different government branches can have different, conflicting priorities.
The Netherlands is tiny. China has 1.4 billion people, and its state apparatus is orders of magnitude bigger. Forget about coordinating the population, even coordinating the tens of thousands of local government bodies has always been a huge problem. All the previous dynasties have said that governing such a large country is a nightmare.
Xi is not personally in charge of the censorship bureau. The top government sets broad direction and KPIs, while local governments and government agencies are given a lot of leeway for implementation as they see fit. And frankly you cannot run a large organization any other way — there is no large company in the world where the CEO micromanages everything without burning out. The KPI is "social stability", and as long as this is kept and there are no grave problems like corruption, it's not the top government's job to dictate how the censorship bureau do their work. Of course, you may be of the opinion that something like "freedom of speech" is more important than "social stability", but the point is that they value "social stability" more, and that they're motivated by that, and by not some idea of "suppressing freedom". This ties directly into my point of properly understanding them.
Furthermore, many people tend to be risk averse, and would rather instinctively deny something than to take chances. There was a famous scene in the Jiang Zhemin days in which Jiang said something frank in some meeting with a foreign politician. Then the cameraman was like "uuh should we record this?" and his boss was immediately like "no, cut it away". Then Jiang said "why shouldn't we record this? of course this should be recorded!" This risk-averse attitude is still pervasive in a lot of places. It's not just DeepSeek that's "paranoid", everybody implementing censorship rules is paranoid similarly. On Xiaohongshu/RedNote they don't want you to talk about societal issues at all, even "positive" things like "I think Taiwan belongs to China" — they recently banned a Taiwanese's account for saying stuff like that, they want you to focus on travel and food or whatever. This attitude likely won't change until the current censorship bureau generation retires, and gets replaced with the next generation that's more confident.
Finally, whatever Poland did pre-1989 has absolutely nothing to do with China. There are no similarities in motives or circumstances. You can't just lazily lump random Soviet-era countries together with China just because you give them both the "communist" label. China's adaptation of and motivation for adoption of communism is wholly different from the Soviet Union.
From a practical POV having all the training data, training infrastructure, and training know-how wouldn't help you either unless you could afford to spend the millions of dollars (hundreds of millions for a SOTA model) in compute to train it each time they released a new training set, in which case you're only talking about the big commercial companies. "open source for the people" just does not apply.
Particularly when it comes to tool use w/ self-distillation it can be done without any data... have a tool the model doesn't know? a teacher model RTFMs and the source code, and helps the student learn to get it right.
But even if I can't build it from source locally, being able to see what went into the model is an important part of what open source is about.
Yes, but for whatever reason this usage seems to have stuck. Open weights is definitely a better name. I assume the reason "open source" has stuck is because you can download and use it for free, but "open source" was always intended to be about "free as in speech", not "free as in beer". That said, I remember when the term "open source" was invented, and it was always a bit different, more commercially aligned, than the goals of the FSF.
> But even if I can't build it from source locally, being able to see what went into the model is an important part of what open source is about.
True. Unfortunately LLMs have become such a big money and closed enterprise (the opposite of OpenAI and Anthropic's altruistic founding principles) that it's hard to see these commercial models releasing their training data, especially since this data is the closest thing they have to a moat other than the cost of training.
The most valuable training data right now seems to be "reasoning data", and the need for this at least may disappear as AI moves beyond pre-trained language models to smarter systems capable of learning for themselves, and that can actually reason, not need to parrot reasoning data.
I don't see an inconsistency. money is pragmatic, the mission needs money
The real mission statement for most companies is to make as much money as possible.
OpenAI might have the same flaw, but you need to demonstrate it, not just assume it
Put it another away: if we removed your ability to do incentivized labor and all you can do is charity work… you would run out of money and die from starvation. If we did the opposite and we removed your ability to do charity work… you’d be fine.
All of this re-emphasizes the point of this thread: In our objective reality, the world is driven by incentive based work while altruism is a side effect of surplus wealth generated by incentive based work. That is the fundamental reality.
My one other point is that incentivized labor is not the same as the value it creates. Indeed, it must be less. Otherwise, our economic system could support only a fixed number of people (subsistence), which would decay inevitably because there is no margin for error. But my point is that margin in reality isn’t fully realized, even by trillionaires, because then there would be no growth to support more people growing in their standard of living. There must be slack in this distributed system and the slack wasn’t valued: it’s free labor. It’s mixed in with incentivized labor, so I understand if you reject the premise entirely, but I do believe this is the essence of modern (specialized) capitalism. If skilled workers try to optimize or invent, more resources will be available for distribution for the same incentive (i.e. “worker productivity”). You can say “yeah that’s their job,” and I can say “that productivity wasn’t fully monetized because otherwise productivity would be lower overall.”
So, incentivized labor presupposes free labor, and economic productivity is a mix of free and monetized labor.
Contributing to it might not necessarily be. Most open source development is funded by large companies after all and from their perspective it can function as a cost saving measure. Allowing them to focus on their core products and removing the possibility of their rivals from getting a competitive advantage due to having a superior low level stack under their product.
Which is why open source is so successful in areas where software is a cost-center but mostly failed for consumer products (since spending resources on them would actually be altruistic unlike e.g. Linux kernel development)
any altruistic act can be perceived as self serving
Software engineers need money to survive. If they exclusively work on open source stuff where are they getting money from to survive? Follow the money trail… even a donation… eventually it leads to an incentive based source or action.
From open source. You can earn money from open source. Open source is not opposed to capitalism, idk where you got that idea.
You’re right. One of its big uses is for big companies to destroy the ability of individual developers or small companies to make a living selling software. It’s also used to harm larger competitors whose revenue depends on some non open source software. Essentially, open source is the new “dumping”.
That’s not to say open source is bad. But in a capitalist context, it’s certainly a very double-edged sword.
I said open source is derivative to capitalism. Meaning open source cannot exist without capitalism. I never said they oppose each other.
Second I said you need to follow the money trail. Money given to people who work on open source comes from non-open source places.
These are orthogonal. One can have a paid job while contributing to open source for entirely altruistic reasons.
> Follow the money trail… even a donation… eventually it leads to an incentive based source or action.
BS. Humans do things for altruistic reasons devoid of individual reward all the time.
I, myself, maintain multiple OSS projects entirely for fun and with the hope that others will find it useful. That's it, that's all. I also donate entirely anonymously to charities simply because I believe others deserve support and dignity.
This form of cold, American libertarianism you espouse is pure poison in the body politic, both in this US and globally. It degrades all of human interaction to transactions. Its no wonder that the US is where sociopaths like Zuck were birthed.
Are you not implying below, with your words, that working exclusively on open source cannot bring money?
> If they exclusively work on open source stuff where are they getting money from to survive?
You seem to imply that open source is incompatible with making money. You seem to believe that if someone is making money they are not doing "exclusively open source" but something else in addition to open source.
> you need to follow the money trail. Money given to people who work on open source comes from non-open source places.
Money spent on coffee comes from non-coffee places, mostly. Does that mean one cannot make money exclusively selling coffee?
I get your point, it is very uncommon to live only of open source. That I can agree with. It is the exaggeration and dogmatism that is untrue.
No I am implying that it in itself does not generate money because it is given away freely. Open source can draw donations but that is not part of a transaction and the source of that money is from incentivized labor.
> You seem to imply that open source is incompatible with making money. You seem to believe that if someone is making money they are not doing "exclusively open source" but something else in addition to open source.
Usually yes but not necessarily. They may exclusively do open source but then in this case they must be supported by some kind of monetary source. That monetary source will be a for profit incentive. For example: donations. Donation money must be generated by for profit ventures. Or salary money paid by companies for employees to work on open source. That salary money is generated from for profit ventures.
> Money spent on coffee comes from non-coffee places, mostly. Does that mean one cannot make money exclusively selling coffee?
Dude. You’re missing the point and making a wrong analogy. It is not about where the money comes from but whether a transaction occurred. When you sell coffee you make a transaction. When you give away software… no transaction occurs. You lose time effort and money for nothing in return. Thus the effort must logically be supported by something else. I am exclusively talking about a logistical issue with a logistical consequence. That’s it.
> That I can agree with. It is the exaggeration and dogmatism that is untrue.
Where did I exaggerate or dogmatize anything? I am talking about a logistical reality and you’re getting emotionally worked up over it. Don’t fucking accuse me of shit I didn’t do.
The problem is you define open source as giving away software, and any money making aspect as "not open source". That is the exaggeration.
If one releases some software as open source and charges for support they are living of open source (my pov), not doing some open source + something else on the side (your pov). I disagree with your definition of open source.
Imagine I create my own illustrations and make them available for free (with some conditions GPL, or even almost without conditions MIT) and print tshirts with the illustrations, and sell them for money. Would it be fair to say I live of being an illustrator? or would you argue I am not living of my illustrations, that I am merely giving them away and really selling tshirts is the business?