Rare find

fp6_llm. An efficient GPU support for LLM inference with x-bit quantization (e.g. FP6,FP5).

github.com/usyd-fsalab/fp6_llm

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.