MiniMax-M3-2x-DGX-Spark-36-tok-s. MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three serving lanes: speed / balanced / long-context.

github.com/tonyd2wild/MiniMax-M3-2x-DGX-Spark-36-tok-s

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.