Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels(nanduruganesh.github.io) |
Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels(nanduruganesh.github.io) |
Has anyone used the new Minimax M3 model? I’m curious how it compares with Deepseek V4 and GLM 5.2 and other larger open weights models.
[Update: their cheapest token plan has been removed, I guess its back to GLM now]
Right now M3 is not far behind DS4, but I belive DS4 will improve much more with each round of training. It simply has a bigger brain, it just needs to fill it with more information.
Such lazy, much farming
https://github.com/fla-org/native-sparse-attention?utm_sourc...