GLM-5.2-Abliterated-NVFP4-316K-4x-DGX-Spark. Abliterated GLM-5.2 with a true 4-bit NVFP4 KV cache on 4x DGX Spark: 316K context, 317,279-token pool (+58.6% vs fp8), 41.4 tok/s peak — with honest content-dependent decode measurements.

github.com/tonyd2wild/GLM-5.2-Abliterated-NVFP4-316K-4x-DGX-Spark

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.