MP-Sparse-Attn. MP-Sparse-Attn provides Triton kernels for Diagonal-Tiled Mixed-Precision Attention, targeting efficient low-bit MXFP inference for Transformer models. It combines tile-level mixed-precision computation and kernel fusion to accelerate attention on modern GPUs.

github.com/yifu-ding/MP-Sparse-Attn

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.