Rare find

flash_attention_inference. Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.

github.com/Bruce-Lee-LY/flash_attention_inference

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.