llama.cpp-turboq-mtp. Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090

github.com/Indras-Mirror/llama.cpp-turboq-mtp

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.