LLMProxy. Seamlessly route requests to your LLM backends—whether you're using stream=false for standard JSON responses or stream=true for real-time token streaming via Server-Sent Events (SSE). LLMProxy handles both modes out of the box, with zero buffering on streams, intelligent load balancing, and OpenAI-compatible API routing.

github.com/aiyuekuang/LLMProxy

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.