Rare find

neural-compressor. SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

github.com/intel/neural-compressor

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.