Rare find

SageAttention-for-windows. Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossing end-to-end metrics across various models.

github.com/sdbds/SageAttention-for-windows

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.