GPTQModel. Production ready LLM model compression/quantization toolkit with accelerated inference support for both cpu/gpu via HF, vLLM, and SGLang.

github.com/liguodongiot/GPTQModel

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.