From HF: 284B parameters (13B active), 1M context window.
This is indeed some kind of compressed context and the quality goes down as the context grows. IIRC the V4 paper had some numbers on this
https://www.reuters.com/world/china/openai-accuses-deepseek-...
But China?
I actually canceled my Claude Code plan a few months back after trying out some of the "lesser" models on openrouter. They seem to work as just as well (or just as bad) for my coding tasks.
Really confused how people make these claims. Are you just basing this off benchmarks or your own personal work? Are you an experienced dev or just doing vibe coding?