GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s. Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster

github.com/tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.