AMD Ryzen AI Halo – $4k AI Dev Kit(lttlabs.com) |
AMD Ryzen AI Halo – $4k AI Dev Kit(lttlabs.com) |
why is everything a BOX ?????????
why not some other platonic solid
Satire if you can’t tell…
So it's usable only in data center conditions but then why make it in this form factor?
With a desktop your system memory is slow and your fast graphics memory is limited in size.
To me it seems like the best bang for your buck in the BYO desktop PC space is to get a board with dual PCIe slots then find some old generation 24GB GPUs like RTX 3090.
But you’re not getting access to more than 48GB of fast memory without something similar to this or a Mac Studio.
On the high end side, it is too slow. On the low end size, it waste money on VRAM.
I don't think I'd pay $4k for it today though, 2 years ago and less than half the price feels like a good machine. I'd be very disappointed in it today for $4k.
It loses to Intel in CPU, and NVIDIA in GPU, in case of scientific libraries and HPC-worthy libs, tools.
I think people who want an "AI Dev Kit" will lean towards Intel + NVIDIA setup.
I am not a fan of Intel, but their MKL, MPI, etc. are not paralleled. Same goes for CUDA with NVIDIA.
As traditionally AMD was a supplier of parts.
Microsoft = yes, they care enormously, as Surface has taken away many sales. Albeit they sold some ChromeBooks
With the current RAM and SSD prices... I rather a bit later.
PS6 "undertaker of physical media" will supposedly be priced >$1k: https://youtu.be/-F1JS-4Abjo
https://github.com/pettijohn/corsair-ai-workstation-performa...
Depends how long this market can remain utterly loony though.
Probably way, way longer than we'd ideally like. Every now and then, you hear that this is the new normal, that it'll last until 2032 or something, and I can practically smell the paid advertising or cult-like messaging behind attempts to destroy any speck of hope consumers have so they'll just pay the toll. But man, it really could be exactly that. There's just too much money to be made for all the involved parties.
Even setting that aside, depending on the pace of datacenter buildout even a round of new fab capacity coming online might not be enough to bring prices down. I think there's a decent chance we have to wait for investors to get tired of building new datacenters which could easily be 5+ years.
Hardware is the exact same as what used to be available for $2K last year (and is still $1K cheaper from Chinese OEMs).
LTT Lab's LLM testing is getting more sophisticated, which is great - I think it's worth noting that ROCm/Vulkan versions and llama.cpp build versions are going to have some big differences for numbers.
For those wanting to get the most out of their Strix Halos, there's both kernel tweaks and utilities like ryzenadj that can help you get the most out of it. ( http://strixhalo.wiki/ has most of that documented). Also, if you're running for coding or agentic work, if you model supports MTP, that's mature and should give you a decent (30%?) decode boost.
just a few weeks ago they backported a kernel oops amdgpu null dereference into stable, it's still not fixed.
Whenever I rebuild llama.cpp, I wind up using the Vulkan build anyway.
It has the same 256 GB/s memory bandwidth limit as every board previously, not sure why this is even being released right now as if it's some new fangled thing - you can go get a Framework Desktop for roughly the same price or a GMKtec EVO-X2 for a bit cheaper.
The memory shortages won't last forever; when companies start adding capacity I wouldn't be surprised to see massive sticks of RAM being sold for consumers.
(I agree that you almost certainly don't need this)
Absolutely no reason that these need to be capped at 256 GB/s other than shortsighted design from three years ago.
I fear that by the time the RTX Spark comes out it'd have to be $6k, and by the time a 128gb or more machine with 700+ GB/s comes out it'd be at $10k, way out of most consumers' hands.
Edit: capitalized gb/s to GB/s.
But when they cost the same price (unless the Spark has shot up too), there's no reason to buy this over a Spark.
The Spark is literally a faster version of this, with better software support.
Edit: And I say that as an owner of a Ryzen AI Max 395 device.
These days, you can get a DGX Spark for $4.7k, so yes, the price has risen, but Strix Halo (with a few exceptions like the Bosgame and Corsair systems) $4k (or more!) is simply not a very good deal. If I were buying new right now, I'd 100% go DGX Spark without even thinking about it.
Gorgon Halo (releasing this fall/winter) is allegedly coming with 3GB memory chips, enabling a 192GB maximum unified memory SKU, alongside minor (100MHz) clock bumps in the iGPU and CPU, plus memory bumping from 8000 MT/s to 8533 MT/s (matching the MBW of the DGX Spark), and is otherwise unchanged. I fear these will be $5000+. At $3000, these would be awesome. At $5000, not so much.
Though, I will say that Nvidia ships a dogshit custom Ubuntu on their hardware that's hard to deal with. Nvidia is not good at software. I keep thinking they'll get better at it, but I've been dealing with their Jetson line for a couple of years now, and it still sucks. Still a clumsy custom Ubuntu, and it's not as easy as simply installing a different Linux version as it's a complicated image-based thing and no UEFI. At least, I assume they ship Ubuntu on the big devices; I've only dealt with the little embedded Jetson machines. The AMD stuff, being a regular x86_64 PC, you can install pretty much any Linux. I immediately put Fedora on mine.
However for local single-user setups, it's often better to have access to more capable/bigger MoE models at reasonable speeds and lower concurrences, which is enabled by these platforms.
I do (and have historically done) quite a work with both local LLMs and local diffusion models. I have an M3 Max MBP at 400 GB/s and also a desktop with a RTX 4090 with 1,008 GB/s
While the M3 Max MBP can serve up MoE reasonably fast (~60 token/sec)the RTX 4090 is an entirely different experience (~170 token/sec). I also do a fair bit of experimentation and am currently running a custom decoder that requires expensive look-ahead, but I'm still able to get a usable 25 token/s on the RTX.
The raison d'etre for the DGX spark is not practical home inference, but rather offering the same fundamental architecture as data center cards for a affordable CUDA prototyping. If you want to build software to run on H100s, you probably can't justify buying (and running) a single card. The DGX spark solves this by having the same fundamental setup as what those cards have.
That makes these non-NVIDIA DGX-like devices confusing to me. The entire benefit of the DGX series is the NVIDIA architecture itself.
Anyone interested in home LLMs should decide whether a Mac or a dedicated GPU is the more sensible path based on their budget and other computer use. Each has their own benefits.
https://community.frame.work/t/was-there-no-possible-way-to-...
> A shame, really, as the Ryzen 7640U, 7840U, 7840HS, and 7940HS all support 256GB of RAM.
To be fair, those platforms support dual dimms per channel, which Strix Halo would not, at least not at it's high speeds.
But reciprocally Gorgon Halo 400 just launched and it supports... 192GB. And is the exact same APU.
Memory chips did finally have their first big doubling per chip semi recently (available last February), with 48 & 64GB dimms becoming available. There is some reasonable lag here, that Strix Halo & Gorgon Halonuse lpddr5x, which perhaps had some lag, that 32GB (x4) was the best available. But now with Gorgon Halo being 192GB capable but not 256GB, it sure feels looks & seems like this is just bad spirited fuckery from AMD. https://forum.level1techs.com/t/where-are-the-ddr5-unbuffere...
128 bit: 96 GB?
256 bit: 192 GB
512 bit: 384 GB?
1024 bit: 768 GB?
The only problem, you need 8 or 16 memory controllers. Memory controllers are not that expensive: Intel Core i3-14100F has 2 channel controller and costs $110, so we can estimate that 16-channel controller should cost not more than $880, and 8-channel controller should cost $440.
So isn't it better to make a cheap CPU with 16 DRAM controllers instead of this $4K gear having only 128 Gb? Or maybe 2 CPUs each having 8 RAM channels?
DDR5 costs 2 times more ($360 for 32 Gb) while not even having 2 times the bandwidth so it is not worth buying. It is more reasonable to make more RAM channels and stuff them with DDR4.
if you want inference, go mac with much higher memory bandwidth. given the price premium here, or the little there is, you might as well.
if you want to finetune and experiment, cuda still has the moat and the kit is not much cheaper, if at all, than dgx spark.
from personal experience, i had access to amd developer cloud with a fair bit of credits. however, even doing inference outside of their supported use cases (which are often dated btw) using vllm was a pain. in the end, despite great computing potential on paper, i decided to not spend more time than its worth on it. if their enterprise cloud continue to have these grievances, i am not optimistic about this kit.
it might be down to skill issue on my end. perhaps if these sell and it gives amd enough motivation to add more software staff in-house, more power to them.
otherwise, good article from labs as usual. nice to know that other kits based on this soc are more or less the same (unsurprisingly).
I designed this:
https://github.com/phkahler/mellori_ITX
It's 195 x 190 x 60mm and takes a standard ITX board. You'll need to relocate the hole for the fan depending on your motherboard, but CAD files are available and you only need to change 2 parameters (X,Y of the hole center).
BTW mine was upgraded to 64GB RAM and a 5700G (zen 3 APU) but it died and I'm still trying to bring up a newer board - still socket AM4.
BTW to get the center coordinates of the hole, measure from the edges of the board to the top and bottom of the metal plate under the CPU socket and take their average distance. That plate is symmetric and centered under the hole.
Open, cheap & good enough will win the race.
Only thing that separates them is the build quality and the extra 20W of boost the framework desktop and this variant support.
They have a note on the thermals but no measurement of noise. Doesn't matter if it's stricly a whoosh or a whine, only if they bother people in the same room. And the small ones like Bosgame get a consistent complaint about the noise in in-depth youtube videos.
For this to be compelling it would need to be eg 256GB minimum or something
"The Apple Silicon Mac Studios outperform the AMD Ryzen AI Max+ 395 machines"
The framework desktop even has a usable PCIe 4x slot available if you put the board in a different case. They sell the 128GB board on its own for $3150.
Waiting for the market to be less insane is somewhat akin to waiting for the s&p500 to drop a decent amount so you can buy in.
There was a post recently, that showed the DGX spark is twice as fast as the M3 Ultra at prompt processing, but half as fast at output tokens [2]. They used gps-oss-120 for that test with a small context.
[1] https://gpuquicklist.com/apus?models=GB10%20Grace%20Blackwel...
[2] https://aimultiple.com/dgx-spark-alternatives , https://news.ycombinator.com/item?id=48732679
I like my Strix Halo and keep it chewing on stuff, mostly non-interactive workloads (security audits of software mostly, training experiments, etc.), I get a lot of use out of it. If you want to experiment with AI, it is a good platform for that, though at $4k you can get an Nvidia-based Asus Ascend GX10, which is probably better. But, if you want a local model for interactive agentic use, you're going to be running either Qwen 3.6 or Gemma 4, which will fit comfortably on 2x64GB GPUs (even old GPUs will run them faster than the Strix Halo...I have dual Radeon Pro V620s which are faster, and they're six years old), or snugly on 32GB. A 48GB or 64GB Mac would run them well. Two Radeon AI Pro R9700 GPUs is probably the sweet spot, right now for GPUs. Not the cost of a good used car, like a 5090 or 4090, but plenty of memory and performance for local inference. Also, not finicky and weird and needing custom 3D printed fan shrouds like the old server GPUs on eBay.
At the moment, there just isn't a model that works better on a 128GB inference machine like this that don't also work fine on 64GB machines, which may be faster (very few 32GB GPUs will be slower, though I wouldn't recommend buying any GPU that isn't currently actively supported by the vendor drivers and CUDA or ROCm...so probably don't buy an MI50 or V100 or whatever).
ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,kvcache_bytes 2048,2048,202.02,128,15.31,52184460 4096,2048,211.03,128,14.64,80373132 6144,2048,208.04,128,14.59,108561804 8192,2048,200.78,128,14.43,136750476 10240,2048,203.04,128,14.37,164939148 12288,2048,200.82,128,14.27,193127820 14336,2048,198.62,128,14.22,221316492 16384,2048,196.14,128,14.20,249505164 18432,2048,189.48,128,14.13,277693836 20480,2048,186.59,128,14.06,305882508 22528,2048,183.88,128,13.99,334071180 24576,2048,183.38,128,13.92,362259852 26624,2048,181.57,128,13.87,390448524 28672,2048,183.46,128,13.80,418637196 30720,2048,181.80,128,13.73,446825868 32768,2048,175.93,128,13.55,475014540 34816,2048,175.42,128,13.46,503203212
I think once someone comes up with a machine which has both it will easily sell for $10000 and people will be queueing to buy it.
You'll need a custom-built distro image, but that goes for like 90% of ARM hardware on Linux.
For anyone considering these devices, the only reason I would recommend against them is if you plan on getting multiple to link together - the DGX Spark has a much, much faster interconnect bandwidth ceiling than the AMD devices do.
Otherwise, they're great!
[1] https://www.microcenter.com/product/699008/nvidia-dgx-spark
The CPU of Ryzen is better than that of DGX Spark, especially for modern programs that have been updated to use AVX-512 (i.e. it has a significantly higher multithreaded performance).
Only for GPU applications the NVIDIA system is likely to be better.
Difference is on the Spark you can use and learn CUDA, vLLM, SGlang etc which is the industry standard.
I bought mine (ASUS version) back in December with the intent to learn that stack of stuff. I've been on and off with it but it seems to have paid off, looks like I might be getting work in the inference serving space.
The big advantage of the DGX Spark over the Strix Halo is much faster prefill. Like 5x the speed. Also the networking hardware on it is insanely powerful, though I and 99% of other Spark users, are unlikely to use it to its full capacity.
Let's just say if I had my druthers, I would not choose Ubuntu, and I really wouldn't choose the Nvidia/ARM spin of Ubuntu. The Strix Halo has the benefit of being entirely a normal x86_64 PC that happens to also have a big chunk of unified memory. You can put pretty much anything on it. Any Linux, regular Windows, probably even a BSD (though good luck getting AI stuff working there).
But, as I said, if you're spending $4000 exclusively for inference and AI workloads, you might as well get the Nvidia-based unit. It is better for that.
And if your jobs do fit onto a 24GB card, then you are not the target user for the "AI mini PC" niche that these guys are trying to carve out
It's great to get lots of tokens, but being able to handle and extent context is why it'll continue to be a great machine compared to any of the small graphics cards.
it allows you to run smaller models much better
imo 3090s make the most sense if you can buy at least 2x ideally 4x but of course we're talking about a completely different budget at that point
This is one of the main reasons (the other is the number of PCIe lanes) why high end desktop and server CPUs have like double the number of pins and so much bigger sockets as compared to consumer desktop CPUs.
In addition to cost and possibly market segmentation, presumably the additional power consumption wouldn't be worth the tradeoff for laptop parts. If you want high channel count you've always been able to purchase power hungry datacenter gear. You can also pick up surplus 10 GbE or even 100 GbE fiber NICs to link up your beowulf cluster. You'll probably have to run a few new electrical circuits and a small bit of HVAC if you're installing this in a residential home but it's entirely doable to cram 10 kW in a closet if you feel like it.
Alternatively you could just rent cloud instances by the hour to avoid the hassle.
Isn't adding pins kind of expensive?
M5 Pro has as slightly higher memory bandwidth, but currently it is available only with up to 64 GB of DRAM. Even with this small memory, it is more expensive than either AMD or NVIDIA, especially when you do not want a puny SSD, but a normal-sized SSD, i.e. big enough to put there an LLM if you want to compute a quantization yourself (i.e. more than $5600 with a 4 TB SSD & 64 GB DRAM).
If you want to do LLM inference, I do not see Apple as a competitive solution, as their price is much, much higher, while also having limited expansion for SSDs, where you need a lot of space if you want to store a few LLMs.
The only thing that is correct is that the AMD Strix Halo system used to be much cheaper than NVIDIA, but now it has the same price. The CPU of Strix Halo is better than that of NVIDIA, but the NVIDIA GPU is likely to be better than the AMD GPU and CUDA is guaranteed to work fine.
Now I think it's totally fine to have a less capable offering, and the Strix Halo is still a mighty capable machine for inference on mid-size MoEs. At 2k it was a tinkerer's dream. But the performance difference should be reflected in the price. This is roughly a doubling of the price compared to less than a year ago without adding any notable features, it's appalling.
Well, almost. Maybe. Not all units. But probably.
There will be a basic M6.
I mostly run them clustered for DS4 and am quite happy with the performance, and the cost isn’t that much more for two than the MBP while giving me double the unified memory.
I’ll probably pick up a third to run multiple smaller models. I don’t understand why people would buy a halo over a spark at comparable prices, particularly because if you want to cluster, the cx7 be beats the shit out of them when it comes to latency and throughput
Its very unlikely that with 192GB Gorgon Halo that anything really blocks 256GB Gorgon Halo, or Strix Halo. Higher than 32GB chips are possible and exist.
And after all, that demand from consumers.
Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.
To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.
If you DIY your own SSD, you can spec a Framework Desktop for below $4k; but not much below. Roughly the same price.
They have made a few attempts at investing in the hardware, but the software side is letting them down hard, and that part is almost entirely their own fault. They have underinvestment in their own stack, and into popular standard and Community libraries that would make it easy to use their gear.
Also, when you do Python work, you have to remember to install dependencies like pytorch from an alternate repo, otherwise you wind up with the CUDA versions that only use the CPU and not the GPU through ROCm.
I haven't finished comprehensive tests but I found:
- Vulkan is up to 2.25X faster for most coalesced, strided and interleave variants for memory-side scheduling/access shapes
- 3.3X faster on specific dot-path sweeps, including for scalar-dequant
- For matched LDS, Vulkan can be 8-14X+ faster (!!!) than matched HIP LDS
HIP doesn't always win against RADV/ACO, but on dispatch/runtime, it does appear to be quite a bit faster than HIP/LLVM on gfx1151 (Strix Halo). I'll be publishing sharing full data once I also run vs gfx1100...
And as for DRAM channels, typical cheap motherboard has 2 channels and 4 slots, it should not be super difficult to add 2 more channels.
If you have a desktop CPU with 2 memory channels and 2 DIMMs per channel, then the 2 DIMMs on each single channel share all of the address and data lines, there's just a chip select difference to pick which DIMM you're addressing. There is some extra loading on the lines, since you have 2 DIMMs, hence the supported speeds are usually slightly lower, but you only add 1-2 more actual PCB traces to have 2 DIMMs per channel vs 1 DIMM per channel. In contrast, adding another memory controller would add upwards of almost 100 additional PCB traces from the CPU.
And that's where AMD CPU shines vs ARM.
Yes they spent those costs to switch from DDR4 to over-priced DDR5 and I suggested the cost could be spent on adding more DDR4 channels instead.
Which is a huge problem? Even using 2x memory controllers in typical consumer motherboard can make system very unstable.
Regretting not buying the Framework back when it was around $2k.
[0] https://www.gmktec.com/products/gmktec-evo-x3-ai-mini-pc-amd...
Absolute insanity
The differences are basically, sparks require ARM and sparks allow interconnects; so if you do have dreams of electric sheep to chain them together, you're not gonna get the AMD halo units.
But if you just want to putz around with a dev machine and do other things, not sure you'd want a spark.
Personally I’ll never buy AMD again.
[0] https://en.wikipedia.org/wiki/Apple_M3 [1] https://en.wikipedia.org/wiki/Apple_M4
Apple M3 Ultra chip 819GB/s memory bandwidth
They have 96, 128, 256, and 512GB variants my friend.
With DeepSeek-V4-Flash-DSpark (Deepseek's new speculative decoding scheme), which is still barely supported anywhere, we're seeing a more steady 45-55 tok/s with bursts into the 60s.
Regular Ubuntu or Fedora ISOs do work out of the box
Just needs the Nvidia GPU driver install afterwards
(And the realtek 10gbe module oot, or blocklist if not used)
For Jetson Orin and later:
You can download an ISO from https://developer.nvidia.com/embedded/jetpack/downloads and that'll work. Other distributions are still a bit of a mess though but Yocto is supported now.
That's actually much better than I would have thought!
Thanks for the answer and it does make this approach make more sense as a budget solution to running larger models locally.
I said less insane, not sane. If prices go down 20% but are still up 150% over a few years ago, that would still be an improvement from now.
Or it could work the other way: if new hardware comes out in 2027 such that the tokens/$ ratio works out better, that would also be less insane.
lol this is so wrong it's funny - equities go up in price, commodity goods go down in price. the two markets are literally diametrically opposed.
So, you should get into RAM futures if you believe this is more than a transient arbitrage sort of situation. All extant RAM will become obsolete as the demand shifts to newer, fancier versions.
I do wish didn't buy the one that only had the 1TB nVME in it though. It's one of those expensive half-sized dealies, which makes upgrading expensive. Esp now.
I basically treated it as a stock Ubuntu machine but on Aarch64, and ignored the stuff they installed. And then went looking for docker images for vllm that made sense and were up to date and hopefully tuned for NVFP4 on the hardware and...
... that was more the disappointing part.
Still, I'd rather have the blazing prefill speed of the Nvidia over my Strix Halo, but I wasn't willing to spend nearly twice as much for it (when I bought, the Strix Halo was $2100 and the GX10 was going for $3700, I think). Now that the difference between the two is much smaller, there's no reason to get the AMD.
I am not sure why NVIDIA couldn't have released a cheaper version of the thing with just standard Ethernet on it and leave it at that. Hardly anybody can afford two or more of them to cluster via ConnectX anyways.
I'm running vllm + GoModel on k3s, works like a charm. Even wrote some CUE last night to generate the values file for my Helm chart. It calculates the GPU percentage for me from the memGB I assign to a model.
xAI effectively did and lucked out to cover their losses and more with it.
Yes but on what timescale? Replacing the > 5 year old sticks in my laptop would currently cost well over half what the machine ran me when it was brand new.