MInference. To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

github.com/qhjqhj00/MInference

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.