gqa-flash-attn-rtx3060. Triton GQA flash attention kernel for NVIDIA RTX 3060 (sm_86). Matches or beats flash-attn2 on 9/10 workloads; up to 2.95x faster on long-context decode.

github.com/tonbistudio/gqa-flash-attn-rtx3060

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.