Rare find

marlin. FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.

github.com/IST-DASLab/marlin

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.