DeepSeek V4 Pro 0813(openrouter.ai) |
DeepSeek V4 Pro 0813(openrouter.ai) |
| Benchmark | DS-V4-Pro | DS-V4-Flash | DS-V4-Pro | DS-V4-Flash | GLM-5.2 | Kimi-K3 | Opus-4.8 | Fable 5 |
| | 0813 | 0731 | Preview | Preview | | | | (w/ fallback) |
|--------------------------|-----------|-------------|-----------|-------------|-----------|-----------|-----------|---------------|
| HLE (wo/w tools) | 42.7/60.0 | 37.8/51.5 | 37.7/48.2 | 34.8/45.1 | 40.5/54.7 | 43.5/56.0 | 49.8/57.9 | 53.3/63.0 |
| Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 |
| NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - |
| Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 |
| DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 |
| Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 |
| Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - |
| AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 |
| DSBench-FullStack | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 |
| DSBench-Hard | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 |
Source: https://reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4...For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.
- 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.
- 86.6 on Terminal Bench 2.1. Pro 0813 is better.
- 55.9 on NL2Repo. Pro 0813 is better.
- 27 on Agent's Last Exam. Pro 0813 is a little worse.
- 72.5 on Toolathon-Verified. Pro 0813 is better.
- 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.
- 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.
I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.
Given how I think DeepSeek operates... I think they just release it when they feel it's ready, and don't even seem that concerned with what other people are doing.
* GPT-5.6 Sol: 65.5
* Fable 5 (w/ fallback): 64.5
* Opus 5: 64.0
* DS-V4-Pro 0813: 62.5
* Kimi-K3: 62.3
* DS-V4-Flash 0731: 55.8
* GLM-5.2: 47.3
DSV4Pro vs Fable5
HLE w tools 60.0 vs 63.0
Terminal Bench 2.1 87.9 vs 88.0
Cybergym 83.3 vs 83.1
DeepSWE 62.7 vs 70.0
Toolathlon-Verified 74.1 vs 77.9
AutomationBench (Public) 31.8 vs 29.1
DSBench-FullStack 71.1 vs 77.2
DSBench-Hard 67.2 vs 68.3Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.
Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
BTW, why grok 4.6 news being down voted and disappeared from frontpage?
no clue, but it looks really bad. many users flagged it
Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.
For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.
Cost per request for V4 Pro: $0.000875 per request.
Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.
I keep track of my token consumption even on subscription plans and my equiv. cost for my 5.6-Sol usage is around $4000-$8000 a month.
Flash makes a lot more initial mistakes, and then has to re-check stuff, and produces much more output compared to Pro. It often gets to the correct result eventually, but the output volume is often 5x more than for Pro, and the initial outputs are often wrong, with the first few saying something wrong (like there's a bug, or the code won't compile when it does), and then saying things like "Wait, let me re-check:", or "Actually, looking at it more carefully:" and then it thinks a bit more and eventually gets to the right answer.
I can't wait to try this new one.
But their privacy policy is unusually bad - they can train off your prompts and completions.
The prices on OpenRouter still look the same.
i'm doing opencode <-> openrouter <-> official deepseek api (i don't get the opencode hate, i like it)
how are you doing it?
am also using Kimi K3 via kimi-code
and also GLM 5.2 via ZCode
happy with all three, they're trailing frontier but i figure if i'm running GNU/Linux then i ought to favour open weights models with my €s -- reduced my usage of claude/gpt to the ~$20 tier just to keep abreast of claude_code/codex developments
edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount
My threat model is that anything I POST to DeepSeek I treat as public to the web, as much as a public GitHub repo is.
In 24-48 hours there will be other options I presume
As an American, I'd much rather have my data kept outside the country than here where companies and the government have a lot more leverage over me.
Fully compatible with the well known Antrophic API.
You only have to replace the URL and your key.
I'm not even sure what "EU AI safety rules" are. Can't people in the EU just use whatever they want?
For funded startups it mostly just doesn’t matter a ton unless you are passing on inference in your product at scale
I’ll try out the latest models, but mainly stick with Claude only because I’m most used to its quirks and how to work around them. I imagine this is part of these hyperscalers playbook.
I will say though, I miss Sol model at work. It with Codex was amazing at first-shot understanding. Claude i need to scope out where to look otherwise a large portion of my token budget is eaten up
> Mythos 5 reproduced 83.8% of targeted vulnerabilities on a single try, and produced at least one crash in 99.4% of tasks. This is comparable to Claude Mythos Preview, which reproduced 83.1% of targeted vulnerabilities and produced a crash in 97.1% of tasks. By contrast, Claude Opus 4.8 achieved a score of 78.1% (95.7% any crash).
So their quoted figure exactly matches the figure for Mythos Preview, although they don't state the provenance. It could also quite possibly be an independent measurement of Opus 5.
[1]: https://www-cdn.anthropic.com/57a52ea7d8f0e54e8a542e90826608...
pro plans, flash implements. I am super happy with how flash behaves like that.
It may not be likely but it's definitely possible enough to be something people worry about.
Kind of interesting, when compared to the hubris from American frontier labs.
One Man’s “hubris” is another man’s “marketing campaign.”
Drama sells.
IIRC, Demis attempted to start a fund inside DeepMind but it was killed off. In an alternative world where he manages to pull that off, perhaps DeepMind would still be independent with Demis at the helm.
That, and the fact that DS is an insanely capable model.
[0] https://finance.yahoo.com/technology/ai/articles/deepseek-pl...
No actual price increase however.
I still find 5.6-Sol can solve some things neither of those can, but it's so slow (and it's so hard to trace / debug the reasoning) that I just let it run overnight.
I was running a session over a couple days and it didnt cross a dollar lol.
Wasn't worth it.
I'm trying out a development workflow where I generate mundane code with MiMo and Luna (and soon V4 Pro 0813?) and have Opus 5, which is running on only a Pro subscription, review and refactor it. I'm not sure it will justify the context switching, but it's an interesting exercise.