FlexGen-PurnedInference. Running large language models like OPT-175B/GPT-3 on a single GPU. Focusing on high-throughput generation.

github.com/Yifei-Zuo/FlexGen-PurnedInference

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.