This is your work, valued
club-3090. Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
★ 1.8kqwen36-27b-single-3090. Shell
★ 117qwen36-dual-3090. Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving
★ 56benchlocal-cli. CLI port of BenchLocal quality bench packs — runs LLM behavioral evals against any OpenAI-compatible endpoint. Companion to club-3090.
★ 10buun-llama-cpp. LLAMA Turboquant implementation with CUDA support
★ 721vllm-ampere-optimized. Optimized VLLM for Ampere
★ 11ODS. Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
★ 3.8kT3MP3ST. autonomous red teaming platform; multi-agent offensive-security meta-harness
★ 5.3kAgent-Reach. Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
★ 62kshard. Pipeline-parallel LLM inference across GPUs on separate machines.
★ 432VibeThinker. Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
★ 1.5kik_llama.cpp. llama.cpp fork with additional SOTA quants and improved performance
★ 3kskills. Skills for Real Engineers. Straight from my .agents directory.
★ 194kdeep-swe. Measuring frontier coding agents on original, long-horizon engineering tasks
★ 1.3kllama-cpp-turboquant. LLM inference in C/C++
★ 2.2kLlama-Optimizer. Automatically find the fastest possible settings for running large language models on your machine. Supports llamma.cpp and ik_llama.cpp
★ 1beellama.cpp. KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM
★ 818train-llm-from-scratch. A straightforward method for training your LLM, from downloading data to generating text.
★ 8.7kpi. AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
★ 80kwhichllm. Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
★ 6kmarin. Open-source framework for the research and development of foundation models.
★ 1.2klongctx. Open long-context inference stack: retrieval + open weights, no closed parts. pip install longctx.
★ 7turboquant_plus. Python
★ 7kMLS-Bench. Python
★ 76club-3090-server. Single-file installer for a club-3090 webserver providing an admin control panel, a reverse proxy that automatically routes requests to the correct containers via a single endpoint with configurable vLLM setting presets, live log management, GPU power and fan speed controls, multi-instance GPU orchestration and API access control for multiple users
★ 21benchlocal-cli. CLI port of BenchLocal quality bench packs — runs LLM behavioral evals against any OpenAI-compatible endpoint. Companion to club-3090.
★ 13club-3090. Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
★ 1.8kdflash. DFlash: Block Diffusion for Flash Speculative Decoding
★ 5.5ksndr_core_engine. SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.
★ 128BenchLocal. Test LLMs on real tasks. Compare models side-by-side.
★ 387ai-legal-claude. AI Legal Assistant skill for Claude Code. Contract review, risk analysis, NDA generation, compliance auditing, negotiation strategy, and PDF reports — 14 skills, 5 parallel agents. If you want to learn how to sell this to real businesses, check out the Skool community
★ 1.6kToolCall-15. TypeScript
★ 340whatmodels. Figure out what models you can run locally
★ 41picoclaw. Tiny, Fast, and Deployable anywhere — automate the mundane, unleash your creativity
★ 30kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 384kCodexBar. Show usage stats for OpenAI Codex and Claude Code, without having to login.
★ 19kmcp_agent_mail. Asynchronous coordination layer for AI coding agents: identities, inboxes, searchable threads, and advisory file leases over FastMCP + Git + SQLite
★ 2kvscode-antigravity-cockpit. VS Code extension for monitoring Google Antigravity AI quotas. Features Webview dashboard, QuickPick mode, and quota grouping.
★ 4.8kproxypal. A desktop app that lets you use your AI subscriptions (Claude, ChatGPT, Gemini, GitHub Copilot) with any coding tool. Wraps CLIProxyAPI with a clean UI for managing connections and tracking usage.
★ 1.2kccs. Switch between Claude accounts, Gemini, Copilot, OpenRouter (300+ models) via CLIProxyAPI OAuth proxy. Visual dashboard, remote proxy support, WebSearch fallback. Zero-config to production-ready.
★ 2.8kclaude-code-proxy. Run Claude Code on OpenAI models
★ 3.7kdyad. Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
★ 21kclaude-code-switch. One-command model switcher for Claude Code (Only for Anthropic API).
★ 654MiniMax-M2. MiniMax-M2, a model built for Max coding & agentic workflows.
★ 2.6kvibeproxy. Native macOS menu bar app to use your Claude Code & ChatGPT subscriptions with AI coding tools - no API keys needed
★ 3.2kCLIProxyAPI. Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.5, Grok 4.3, Claude model through API
★ 45kCL4R1T4S. LEAKED SYSTEM PROMPTS FOR CHATGPT, CLAUDE, GEMINI, GROK, PERPLEXITY, CURSOR, LOVABLE, REPLIT, AND MORE! - AI SYSTEMS TRANSPARENCY FOR ALL! 👐
★ 47kK2-Vendor-Verifier. Verify Precision of all Kimi K2 API Vendor
★ 581LongCat-Flash-Thinking.
★ 287Ai-speedometer. Benchmark AI model speeds with TTFT, tokens/sec, and performance metrics .
★ 21qwen-code. An open-source AI coding agent that lives in your terminal.
★ 26kUI-TARS-desktop. The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38kBMAD-METHOD. Breakthrough Method for Agile Ai Driven Development
★ 51kcamoufox. 🦊 Anti-detect browser
★ 11kTurnstile-Solver. Python-based turnstile solver using the patchright library, featuring multi-threaded execution, API integration, and support for different browsers.
★ 882security-guide-for-developers. Security Guide for Developers
★ 21k