Rare find

vLLM-2080Ti-Definitive. The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode with support of FP8 weight

github.com/weicj/vLLM-2080Ti-Definitive

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.