The lack of ASML EUV machines certainly hurts, and pushing DUV so hard results in abysmal yields of good chips, but you can compensate by running more wafers or making smaller chips, and the net result is that Huawei's Ascend production volume is limited by CXMT's HBM capacity not processor dies.
The problem is that HBM manufacture requires many steps (die thinning, via drilling, plating, alignment) where the equipment used by everyone else (Samsung, SK Hynix, Micron) is also blocked by sanctions, so the Chinese are having to develop all of this themselves too, which they have, but yields are currently low, even when using shorter HBM stacks.
I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?
These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That's what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.
Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.
HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.
The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.
So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!
These are not normal requirements, other than for a GPU.
But in reality we also already have unified memory architecture systems, integrated graphics etc.
People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.
They have managed to pull this sort of thing off many many times. https://en.wikipedia.org/wiki/Reality_distortion_field
No, Mac laptops use LPDDR, currently LPDDR5X.
And the worst thing is that is the best possible strategy for them. It's essentially win-win for everyone but consumers.
- Buyer (AI) has crazy money, so will pay whatever
- Seller doesn't have to build anything new, as the buyer is willing to pay whatever
- Generate ridiculous profits from crazy money
- No oversupply risk in case of reversal
- Return from producing HBM to DDR5 in a single quarter if reversal does happen.
Hold on to your existing hardware, people, and be on lookout for your local deals.
Everything in their statement can be true and it be a bad thing for consumers of non-HBM RAM.
"HBM capacity to expand to 250,000 wafers a month"
So let's say their current HBM capacity is 100k wafers/month (pure speculation/random number for illustration), and their total RAM capacity (including HBM( is 300k wafers/month, then non-HBM capacity reduces from 200k to 50k.
I.e., you get a locked down device with access to their AI and maybe a way to run third party apps. Of course the developers of those apps will have to pay a percentage of their revenue.
Sound familiar?
Which governments, and what exactly would you want these governments to do? The demand is global. There are no levers a single government can pull to meaningfully influence the global demand without fully committing to an protectionist economic policy, in which case the U.S. doesn't have the facilities to magically pop up world class fabs overnight, and South Korea and Taiwan don't have the market demand that the U.S. generates to justify their investments in making these chips and China lacks the IP to be able to build anything comparable to Nvidia's silicon at the moment.
No one has all the cards and no one controls all the levers.
This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.
I'm as fanatical a free-market fundamentalist as you'll ever meet, but if someone were to argue that government has a role in preventing bullshit like OpenAI's unilateral 40% attack on the entire DRAM market, backed by nothing but funny money, I would have a hard time coming up with defensible counterarguments.
Can that share of production be allocated to consumers, and the AI fights over the rest?
Bastards.
If I get 10% more performance for 50% more cost it really depends on one's needs, for example.
It's the address space that's unified, not always the physical hardware.
The data movement (when needed) is handled transparently in the background by page faults and other tricks.
The reality distortion is that people seem to believe it's HBM, or somehow it gives you extraordinary amounts of vram. Neither are really true.
This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.
So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.
Especially not for consumer goods.
It may turn out that demand for SOTA frontier models isn’t so limitless after all.
We're not seeing the progress in those "frontier models" that we have previously seen. There's certainly still gas left in tank tank, but we're way into the diminishing returns by now.
Cloud inference still beats hardware investments by orders of magnitude of course, but that's only if your data doesn't really matter to you.
I agree that datacenters are not going to go away, but I have doubts that the buildup that has happened is really going to pay off for most operators.
But I suspect returns may have already diminished into negative territory for at least some other use cases. One of my least favorite job responsibilities in this brave new era is figuring out how to avoid performance and behavior regressions when an older model were using for some application reaches end of life. It’s getting uncommon for me to look at our benchmark results and say, “Oh, good, it does better on one of the newer models!”
Before that, we had 0. After that, we had more than 1.
A leap as far as that is hard to recreate.
But that wasn't my point. That's just trolling.
The actual point is that LLMs aren't gaining new capabilities anymore. They just get more reliable at the ones they already have; turning what was a coin flip to some higher probability.
That's (intuitively speaking, not strictly mathematically speaking) kinda the mathematical definition of diminishing returns.
Alternatively, the money invested can be rational because it prices out competitors and establishes a monopoly, after which point the monopolist earns their absurd investment back with complete control of the market. This is also bad.
The problem you have here is hundreds of different companies are doing the "gambling" in a non collusionary manner so it's going to take decades in court to prove it. Are you saying governments should do an authoritarian take over of RAM allotment?
Part of the reason harnesses work well is you can run a lot of agents in parallel. That doesn't slow down demand.
LLMs are amazing tech, but they're terrible without oversight. More agents faster just makes reality collapse on them quicker.
But yeah, you're right, temporarily, this will still push demand. But the topic was about "diminishing returns" as in "tech getting better". Not as in "customer spending".
New models keep being able to use more and more agents on longer time frames. Your hypothesis doesn't look like what we're measuring.
And memory is already expensive. It's downright hard to even get it though - you frequently would prefer not what's cheapest, but whatever is in largest scale production.
Consider the 'MMA N matrices' primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there's nothing preventing it from being 1024 or larger if we have more memory bandwidth.
All this with a single instruction.
https://news.skhynix.com/en/fab-facility-investment-2026/
It's not an overnight thing obviously.
What Altman did, and the way he did it, was not OK. Unless you believe that humanity's adoption of AI-based intellectual work is best served by concentrating an enormous amount of power in one company at the expense of everyone else, turning the market into a literal zero-sum game. Do you?
Long-term contracts, not building new fabs - they're in full on hedging mode right now.
Their choice, their consequences.
I expect something similar for AI. It's useful tech... just not very profitable tech, and certainly not to the tune of trillions of dollars worth of public demand. The promises of superintelligence, replacing everyone, and the rest of the hype will die with the companies who made the promises, but the tech will survive and thrive.
The hype will die but there are many reasons why super intelligence and robots are a separate entity from said hype.
Look, back a few decades and tell someone our gdp would be in the trillions and it's likely they'd have a hard time believing you. A huge portion of our products that we use day to day would be complete science fiction to them.
All we're negotiating at this point is the timescale it will take.