yolo.cu. 1,100+ FPS YOLOv8/YOLO11 inference on a $400 GPU — every kernel hand-written in one CUDA file. No cuDNN, no TensorRT, no Python at runtime. All scales, all task heads, same detections.

github.com/Soju06/yolo.cu

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.