GPTQModel-1. Production ready LLM model compression/quantization toolkit with hw accelerated inference support for both cpu/gpu via HF, vLLM, and SGLang.

github.com/smpanaro/GPTQModel-1

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.