qwen35x. qwen35x is a focused C++/CUDA inference engine for Qwen3.5 models, built for correctness-first bring-up and fast GPU decode. It includes a native tokenizer, CPU reference inference, CUDA-hybrid inference, and reproducible benchmark scripts.

github.com/Danmoreng/qwen35x

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.