Rare find

cider. W8A8/W4A8 inference + optimized SDPA on Apple Silicon — unlocking unused INT8 TensorOps in M5 for 1.2–1.9× faster LLM prefill, plus FlashInfer-inspired GQA decode attention for up to 1.6× SDPA speedup, built as MLX custom primitives.

github.com/Mininglamp-AI/cider

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.