MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai) |
MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai) |
I get way more usage for way less money without any quality or performance degradation. My $200 Codex Pro plan allowance is depleted in 2-3 days. Sometimes Tibo announces a usage reset. But GPT-5.6 models are really not good for coding. Sol has been making increasingly more mistakes in the past two weeks even in the reviewer and advisor roles. Astra is usable for coding but slow and very expensive. In the past two days I’ve used up over 70% on simple copy editing, with dedicated short specs and short sessions. Really little one can do to make it more efficient. Similar work took 20% at most just a month ago. I’m looking to use Astra for milestone reviews/advisory. Perhaps a $100 Pro downgrade will be enough. But my main work is now on open-weight models. And you don’t need to depend on someone to send you a reset. And it’s cheaper by the end of the month too.
With Claude the limits are not even fun anymore - my weekly $100 Max plan quota is gone in one day on merely review invocations, no coding. And my $200 Pro quota is gone in two with some coding. Sonnet 5 is not usable for coding. And Opus 5 tends to always make a couple avoidable mistakes on every task. Fable 5.1 is ok but tends to ignore skills and to work around explicit instructions. Completely canceled all my Claude subscriptions.
With Qwen 3.8, DeepSeek 4.1 Flash, GLM 5.3 I’ve been getting Opus 5-level performance, with less blah blah and no overengineered churn. Public benchmarks are really not telling the real story. The models are more dependable and more predictable. They have their own failure modes. Sometimes DeepSeek 4.1 Flash is quite stubborn but it fails in a good way. Bad for full autonomy - I need to intervene, but it sticks to the rails and instructions - other than Fable and Opus that try to outsmart you and your harness.
Grok is interesting but has been a bit underwhelming on Grok plans - my SuperGrok allowance is depleted in a single session overnight. SuperGrok+ gives more but it’s still about the same as with OpenAI, Claude is way less now.
Since the allowance volume has been shrinking with the major model providers, to me, open-weight alternatives are really necessary now to at least maintain the momentum and budget.
But at work it’s really an uphill challenge - it’s become impossible to convince the tech leadership once they got hooked on Anthropic. They No facts will help. Some people underestimate how expensive Claude really is after getting used to the subscription plans with allowance resets. OpenAI models are expensive too.
I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.
And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.
I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.
Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.
I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.
But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.
Early results with MiMo 2.6pro are quite encouraging for anything that's non-UI work so likely switching spend for the time being
Serious question: does anyone have evidence of this?
It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.
By and large they don’t. I have seen this drop a few times, eg before fable came out opus dropped a lot probably due to less compute available.
My guess is it’s a combination of getting used to the new cliff models fall off on and forgetting that model performance drops significantly when context fills up.
So new model comes out, people try it and it’s amazing on a task or two. Then they start using it, context window fills up and it gets a lot worse.
The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
Why?
W what?
KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
1 - https://bench.killswitch-lang.org/Interested to see if it also beats DeepSeek V4.1 Flash.
Human error means this wasn't just stopped together by some bot.
For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.
Mostly due to lower cost of living; Shenzhen is way cheaper than SV
I wonder what dangoodmanUT is using! This is the time to compare!
When I’m doing work on a repo where I’m implementing a standard and the agents have to read the standard to keep from hallucinating my usage skyrockets.
Hell this changes depending on which language I’m working with.