llama.cpp. llama.cpp fork with a patch for RYS-duplicated Qwen3.5/Qwen3Next models (non-uniform full_attention pattern). See rys-qwen35 branch.

github.com/DJLougen/llama.cpp

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.