Rare find

decoding_attention. Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.

github.com/Bruce-Lee-LY/decoding_attention

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.