Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
> and 23 000 for deepseek
How did you calculate it? Based on per 5 hours max request allowance?
Would be cool if they added it.
There are some quirks if your harness use unsupported features of course.
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
> if one looks at the CoT, it's evident that it's way way stupider than frontier models
Frontier models don't show the full CoT
Can't you just say "shrank to 1/437th the size"? It's not that hard.
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
I do wonder how long it'll be before a us-hosted offering is available via bedrock, copilot, etc.
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
So I ask again, what are you basing your assertion on?
BUT. they are employed to do / deciding-to-do authentic (if often meaningless) stuff.
here's a short list of inauthentic activity that claude and openai refuse to do:
- chat services that, when you ask them, say they are not chatbots when they are
- code to work around software licenses or DRM
- code to scrape or download copyrighted material
- directly cheating on homework
- adopting a persona in social media that spreads misinformation or propaganda
this is but a short list. but ask me, "are there enough inauthentic activity demands such that someone who CANNOT USE claude or gpt as the LLM would use dsv4.1 on openrouter instead?" yes. i mean there are whole countries right now where the culture can be summarized as, "bottom to top, inauthentic activity." i am surprised it is not more usage!
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).
• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
Memory companies have price fixed multiple times. They've paid hundreds of millions in fines. wikipedia even has a page on it. https://en.wikipedia.org/wiki/DRAM_industry_price_fixing.
Look at the financials of these companies, they're all making obscene margins and do they plan to increase production? No. Micron is doing a stock buy back to pump the price of their share.
The Micron CEO just recently said this is the exact plan https://www.theregister.com/systems/2026/10/01/ram-supply-se...
There's sanctions, tarrifs, and a DOJ who doesn't give a shit. Until we can fix that the insanity will continue. Phones will be unaffordable. Laptops will be obscene. Gaming consoles will be thousands of dollars. Desktops will be dead.
If you're waiting for some David Ricardo equation to happen, tough cookies, it's not coming.
The market is legally locked down and we're in hostage pricing mode.
And what's the story? You can't afford electronics because we're using it to build robots to take your job? I mean ...
Nobody is coming to save us. That's our job.
Micron has 3 brand new fabs currently under construction, 2 Boise, 1 in New York as the first of 4 planned for a campus.
Plus expanding other existing facilities.
These things take ~3-5 years from breaking ground to full production. You'd have had to anticipate the current demand years before it happened in order to be bringing production on-line before 2030 or so.
Samsung and HK Hynix also have fabs under construction and planned.
CXMT started 11 years ago and only now is reaching any real volume. If they decided a year ago to react to the current demand cycle they'd be 6-7 years out.
Not much you can really do to wish for more fabrication to exist on any timeline not measured in fractional decades.
Could they do more and react quicker? Probably, but everything I've read on the subject seems to point to 3 years is absolute bare minimum if you happen to have a shovel ready project with the land bought, local permitting completed, infrastructure extended to the site, and a skilled workforce already in place. They could suspend buy-backs/dividends today and dump it all into building production and there would be no material impact until around 2030.
> The Micron CEO just recently said this is the exact plan
CEO simply stated the demand pressure will not go away through 2027, and supply will not increase until around 2028 when currently under construction fabs start shipping volume. The article does not support your statement.
wonder what voting would be like?
gamer vote ++
datacenter hater vote --
datacenter lobby ++
micron lobby --
I know it's just a figure of speech, but damn. I laughed out aloud in public just reading this.
I guarantee Micron & friends are not intentionally orchestrating their business such that they would suffer a massively reduced chance of yielding on a per-die basis. Unless someone is actually buying HBM devices, they are not going to be making them. These are not a commodity that can be speculatively manufactured in any economically rational way.
Also the released weights are pre-quantised and presumably QATed, so your "Full Precision" and INT8 are simply not a version of the model that actually exists.
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
1070ti launch MSRP was $450 ish. 5070 could be had in the last year for 5xx-6xx range easily.
All things considered - (inflation being about 30%~ (guess)) between these two timelines. You are looking at 300% performance difference at a cost dollar for dollar that is cheaper then when they purchased their cards.
Might be a bit of a stretch blaming it on "severely overpriced for too long..."
Allow a question from someone who’s only got a very vague idea of how this kind of stuff works behind the scenes: say I rent usage of this model through one of the many LLM hosting providers out there, and let‘s assume I use it extensively through something like Pi or OpenCode and vibe code away all the time, keeping the hosted model occupied as much as I can, happily burning my credits.
Does that mean that there is a hardware cluster as described by you above that is crunching away just for me?
So at FP16, I alone keep a 1,664 GiB system occupied all the time?
It has a set of n-gram tables which you can stream from system RAM or even NVMe
That said it’s still quite big! I can’t fit it on my DGX Spark, though I believe you can if you have two?
Because it's an open model so providers compete on price.
He gave demand signal so many times years ago and was mocked for it and now we have the consequences of industry not taking him seriously.
I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense because my implementation uses float32 precision anyways)
I'm currently learning how to distill reasoning traces (check my other github repositories) but I think that a locally selfhostable deepseek is possible with my mixture of experts sharding mechanism. I decided to optimize everything for CPU parallelization, with the idea that the KV cache and meta model have to run from CPU RAM anyways, so the experts can also be loaded/unloaded at runtime if needbe, to save more RAM.
My assumption is that the KV cache optimizations in combination with the CED and compressed attention features are the reason why v4.1 flash has so few hallucination problems and such a strong self-lookup/thinking behavior. But that's more a gut feeling, need to evaluate and test this more thoroughly.
Anyways, would love to see someone train this on their own datasets. Currently my pipeline is kinda optimized for parquet and zim files.
No, you won't get frontier-level intelligence on a 1070Ti. Yes, it should be illegal to do what Altman did. Since we clearly don't live in the best of all possible worlds, we need to settle, and DS4.1 Flash is a good place to do that.
For tasks that don't require vision I personally like the NVFP4 quant of GLM 5.3 from Local Inference Lab better than DS4.1F, but they are both well beyond awesome.
1660 ti, 4790k, 16gb ddr3
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
Plus you can also get dsv4.1f subsidized. OpenCode Go gives 4x if I understand their pricing correctly. Anecdotally, I feel like I get way more out of my $10/mo OpenCode Go sub for the price than my $20/mo ChatGPT, even using gpt-6.1-sol high which is very cheap, and I have yet to convince myself dsv4.1f is a worse model.
It blows frontier API pricing out of the water, but again, look at cost per task, not token usage. Still easily wins though for my work.
I do think it's the most viable alternative I've seen so far, and that applies pressure to the frontier models. Should subscription prices hike or become unavailable for some reason, I know what I'll be using.
When pricing this, it's important to consider whether or not you want to opt out of data training. You won't get the advertised rate. Also the dsf 4.1 subscription providers are throttled af... and of course they are, because otherwise they'd be haemorrhaging money.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
The sub-agent separation is still valuable to keep context clean for the orchestrator, but I just have no reason to use Sonnet as the grunt-work implementer because I'm finding it hard to run out of tokens with Opus 5.5 on a $200 subscription plan. It's really really good at subjective quality of work per token used.
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?
I think there would be a market and I think it would be large.
So, I agree.
Unfortunately, it looks like Cursor doesn't support DeepSeek... but I can say that for lots of tasks, the cheapest models do just fine, too. Of course there are cases where they just keep spinning for 30 minutes without being able to figure things out, but in those cases, I just retry on a slightly more expensive, hopefully better model.
Some of the models are so good and cheap, not sure about the frontier labs and their multi-trillion dollar evaluation. Sometimes "good enough" is really good enough.
I don't even bother checking how much I spent on API any more, its well under $30 over the past 2 months despite daily constant use. Who even needs a subscription at these numbers?
On that note I’ve been subbing in MiMo-2.6-pro when cost is an issue, which is super cheap and also performing really well.
The reasoning and the result document were done after less than 1 or 2 seconds.
Have Ollama suddenly bought GPU capacity?
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
Currently have auto compaction turned off. When the orchestrator's context is getting close to full, I have it write a handoff markdown file and point a fresh agent at it.
I do feel like I'm getting close to the point where I might be ready for something more sophisticated, especially wrt to subagents communicating with the orchestrator.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
Theres already models that outdo DS 4.1 flash in cost/performance. Luna 6 on max effort for example. Luna also doesn't care what time of the day it is for cost calculation.
And I'm sure by the time people ask why Luna 6 is being slept on there will be another cost/performance king
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at end.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
see https://artificialanalysis.ai/models/releases/comparisons?co...
The token-equivalent monthly spend is > $5K+. If Deepseek's token cost is 20x cheaper, that's $250/mo, and I'd be spending a lot more of my brainpower babysitting it and getting worse results.
For business/team accounts that pay per-token, maybe I can see the "freaking out" being warranted on the part of the fronter labs. But as long as they're willing to subsidize their end-user subscriptions, I'm not going to move off of them until the alternatives are truly at their level.
And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
> By default Mjolnir sends recent prompt and reply text and help-search text to TypeSafe's hosted Jev classifier through a public proxy
Is OpenAI coming in $20B under a sign of "freaking out"?
(And what are the preferred providers?)
are you worried about sending all your data to third parties, especially if they're in different countries?
nobody here is talking about running frontier level intelligence locally so if you’re Chinaphobic and prefer layers of corporations siphoning your data in between you and the party there are plenty of options instead of directly to the party
> Today's models are now good enough for high-quality unattended tasks. Chasing the latest and greatest is silly. It is fun to see the new Fable capabilities, but the tasks we throw at them are usually ridiculous (maybe even insulting) if you believe in LLM sentience. It's like asking a math PhD to organize the files on your desktop.
I'm using DS V4.1 Flash as my main model since their release and it works great for all my coding tasks. My setup is OpenCode Go subscription and obra/superpowers skill.
The only times I try to change models are on general planning tasks (like research this codebase for tech debt mitigation opportunities) or if I need deep research which would benefit from searching the web, in which I still think Gemini is still the best because of the speed and access to google search index. But these are not even 20% of my daily tasks.
but DS 4.1 Flash is good enough for most tasks
Anthropic and OpenAi are in the news, so they get the press and people go and try out their product. Large enterprise businesses are going to make larger, longer-term contracts with them and are only going to pivot if they think switching costs are easy or if they think the provider won't deliver.
The other inference producers are less well known or you need to get your cloud sales rep to tell you how to switch to them as a provider rather than Anthropic or OpenAI.
I use OpenRouter, I know switching is easy, but larger businesses tend to work in yearly cycles. DeepSeek v4 Flash came out in late April.
I agree OpenAI and Anthropic are going to struggle when the median price of running a smart-enough model keeps falling.
Edit: I also think demand for hardware will be rapidly absorbed by other companies if Anthropic or OpenAI stumble. We've finally turned hardware directly into runnable intelligence and people are not going to go back to the old ways.
People tend to conflate the question "is AI a useful technology?" with "are the AI companies going to do well?" but they're surprisingly separated in practice, with either one able to be true while the other is false. There is a lot of money tied up in a lot of hardware with a lot of loans made against that hardware as collateral all based on the assumption that AIs are going to need more and more and more and more hardware and whoever has the hardware wins. If a much better model comes out that requires vastly less hardware, or even more accurately, merely charges vastly less than the current AI companies, then to a first approximation (barring Jevon's paradox, and bearing in mind there's no timeline guarantee on that) all that hardware becomes much less valuable for being grotesquely oversupplied relative to what is necessary, and even though that would generally make AI objectively more useful than it was before, it would cause mass financial chaos in the markets.
The markets need a very particular rate of progress. It isn't entirely clear to me that it's even a possible rate of progress, it may be overconstrained, but they certainly don't have plans for the AI models to get commoditized on the timeframes of these vast, vast array of loans being made against hardware as collateral. Spend a metric shit ton of money to kill all your competition then charge monopoly rent on the one thing absolutely everyone needs doesn't work if you can't economically "kill all your competition" because the economics favor them in the spending spree.
And then, based on the fact that this is not even remotely complicated logic, there are plenty of people who are fully aware that they have a lot of money tied up in not running around telling everyone how wonderful the cheap models have become.
Or perhaps they consider the upside from cheap Chinese models to hedge the effect that OpenAI/Anthropic collapsing would have on their portfolios. This would make sense for (hedge funds holding) most companies: they don't really care about who supplies the AI, as long as they get it at roughly the same price as their competitors.
Costs did go nuts, but there are signs of easing in the market of late. CXMT is starting to have an impact and priced will probably fall in 2027.
Some of my family is pretty happy, though, with the job security as they are pretty convinced these projects are all going to take much longer than what's being stated publicly. Micron is saying the first chip from the new fab will be in 2027... though they also predicted it'd be 2026. The date seems pretty slippy.
You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.
I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.
It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.
Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!
Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.
But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.
But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.
This has happened in electronics markets before. Many times.
When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.
This also assumes heavy utilization, though. If there's heavy utilization, it might mean they're doing well. If they're all spinning, it's time to raise prices.
The model engine provider might be ZDR, but the service as a whole isn't.
Let's say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase.
You have a data center and you're in debt, this is all financed on the premise that you have these special valuable expensive computers.
Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or whatever.
You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that's unlikely. It's likely going to drop.
On market if you were to sell some of that ram you have 100% profit right now.
What do you do?
Historically the answer has been "sell that shit"