llm-inference-gateway. OpenAI-compatible LLM gateway: routes to the cheapest model meeting latency/quality targets, with SSE streaming, semantic caching, per-key rate limiting, and cost tracking.

github.com/edaywalid/llm-inference-gateway

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.