This is your work, valued
PhD student at Secure Learning Lab at the CS department of the University of Chicago
HALC. [ICML 2024] Official implementation for "HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding"
★ 115SafeWatch. [ICLR 2025] Official implementation for "SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations"
★ 45RL_Plane_Strategy. established for the data normalization and reinforcement learning training scheme to train an agent in DCS world
★ 12MJ-Bench. Official implementation for "MJ-BENCH: Is Your Multimodal Reward Model Really a Good Judge?"
★ 8POAR-SRL-4-Robot. The implementation for the integrated SRL-RL algorithm POAR as well as the simulator for environments
★ 6Awesome-Diffusion-LLMs.
★ 5PANDORA. Python
★ 5AutoPolicy. C
★ 5robotXhuman. Applying IRL algorithms to learn robot arm motion planning from human demonstration
★ 3CAV-intelligence. This repository contains the algorithms implementation for vehicles scheduling, dispatching and planning in complicated scenarios such as intersection, junction etc. Currently we are developing learnable driving policies module via inverse reinforcement learning algorithms.
★ 3pottery-fragments-matching. applying both traditional and heuristics methods to pottery relics restoration (fragments matching, 3D model alignment)
★ 2SafeRLZoo. SafeRLZoo, a standardized toolkit with over 12 SOTA model-free safe RL algorithms based on Spinningup and benchmark safety-critical tasks
★ 2Notebook.
★ 1Visualization-Project-for-Greyout. This project mainly develops a visualization platform aiming to process and present the medical data of Greyout/Redout to analyze this symptom and its causes both qualitatively and quantitatively.
★ 1BillChan226. Config files for my GitHub profile.
★ 1ODA-Multi-Manipulator. TeX
★ 1firecracker. Secure and fast microVMs for serverless computing.
★ 36kcve-bench. CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
★ 262cybench. HTML
★ 294slime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kProRL-Agent-Server. Agentic RL on Any Harness at Scale
★ 730SkillOpt. SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
★ 15kAgentDoG. A Diagnostic Guardrail Framework for AI Agent Safety and Security
★ 673superpowers. An agentic skills framework & software development methodology that works.
★ 264kProgramBench. Can Language Models Rebuild Programs From Scratch?
★ 872DecodingTrust-Agent. Python
★ 72OpenShell. OpenShell is the safe, private runtime for autonomous AI agents.
★ 7.9kagent-sandbox. agent-sandbox enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI agent runtimes.
★ 3.3kNemoClaw. Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
★ 22kRedTeamCUA. [ICLR'26 Oral] RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
★ 57AReaL. The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.6kboxlite. The micro-VM for AI agents — light enough to embed on your laptop, elastic enough to power an agentic cloud.
★ 2.2kclawshell. The local runtime control layer for OpenClaw/Hermes-agent, PII & sensitive credentials protection.
★ 322agent-world-model. Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
★ 421mcporter. Call MCPs via TypeScript, masquerading as simple TypeScript API. Or package them as cli.
★ 4.9kInnovator-VL. Fully Open-source Multimodal Language Models for Science Discovery
★ 167invariant-gateway. LLM proxy to observe and debug what your AI agents are doing.
★ 78mcp-context-forge. An AI Gateway, registry, and proxy that sits in front of any MCP, A2A, or REST/gRPC APIs, exposing a unified endpoint with centralized discovery, guardrails and management. Optimizes Agent & Tool calling, and supports plugins.
★ 4.2kmcp-gateway. docker mcp CLI plugin / MCP Gateway
★ 1.5kagentgateway. Next Generation Agentic Proxy for AI Agents and MCP servers
★ 4.2kUnla. 🧩 MCP Gateway - A lightweight gateway service that instantly transforms existing MCP Servers and APIs into MCP servers with zero code changes. Features Docker deployment and management UI, requiring no infrastructure modifications.
★ 2.2khappy. Mobile and Web client for Codex and Claude Code, with realtime voice, encryption and fully featured
★ 23kui-ux-pro-max-skill. An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms
★ 112kOpenTinker. OpenTinker is an RL-as-a-Service infrastructure for foundation models
★ 677oat. 🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.
★ 666mcp-gateway. MCP Gateway is a reverse proxy and management layer for MCP servers, enabling scalable, session-aware stateful routing and lifecycle management of MCP servers in Kubernetes environments.
★ 760agentbeats. Python
★ 80AutoEnv. Scaling Agentic Environments Automatically.
★ 66torchforge. PyTorch-native post-training at scale
★ 700bolt.new. Prompt, run, edit, and deploy full-stack web applications. -- bolt.new -- Help Center: https://support.bolt.new/ -- Community Support: https://discord.com/invite/stackblitz
★ 16kMiniWoB-plusplus. A collection of reinforcement learning environments for simple web interaction tasks
★ 393Kiln. Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
★ 5ktinker-cookbook. Post-training with Tinker
★ 4kAwesome_Environment_Scaling. Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to agent intelligence.
★ 72PipelineRL. A scalable asynchronous reinforcement learning implementation with in-flight weight updates.
★ 430novu. The open-source communication infrastructure for agents and products
★ 39kRLVE. [ICML 2026] RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
★ 227mem0. Universal memory layer for AI Agents
★ 62kMemori. Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.
★ 16korangehrm. OrangeHRM is a comprehensive Human Resource Management (HRM) System that captures all the essential functionalities required for any enterprise.
★ 1.1kpromptfoo. Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
★ 24kOpenHands. 🙌 OpenHands: AI-Driven Development
★ 83khrms. Open Source HR and Payroll Software
★ 8.3kminthcm. Open-source HCM system for managing HR processes with AI agents. Full data ownership, MCP/A2A-native, no vendor lock-in.
★ 387AgentLab. AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.
★ 611skills. Public repository for Agent Skills
★ 165ksecure-mcp-gateway. Secure MCP Gateway - Setup Admin level gateway functionality for MCP servers - with guardrails at each MCP server to overcome multiple security issues with using MCPs
★ 57OpenEnv. An interface library for RL post training with environments.
★ 2.5kWebShop. [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
★ 576SuiteCRM_MCP. Model Context Protocol server for SuiteCRM integration. Enables AI assistants to interact with SuiteCRM for lead management, contact operations, and CRM workflows.
★ 10inspect_petri. An alignment auditing agent capable of quickly exploring alignment hypothesis
★ 1.3ksupergateway. Run MCP stdio servers over SSE and SSE over stdio. AI gateway.
★ 2.8kmailpit. An email and SMTP testing tool with API for developers
★ 10kSHADE-Arena. Python
★ 57SuiteCRM. SuiteCRM - Open source CRM for the world
★ 5.6kEPIC. (NeurIPS 2025 🔥) Official implementation for "Efficient Multi-modal Large Language Models via Progressive Consistency Distillation"
★ 49ramparts. mcp & skill scanner that scans any mcp server or skills for indirect attack vectors and security or configuration vulnerabilities
★ 95terminal-bench. A benchmark for LLMs on complicated tasks in the terminal
★ 2.5kverifiers. Our library for RL environments + evals
★ 4.4kawesome-model-based-RL. A curated list of awesome model based RL resources (continually updated)
★ 1.4kterminal-bench-rl. GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's TerminalBench leaderboard.
★ 398ART. Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
★ 11kTextWorld. TextWorld is a sandbox learning environment for the training and evaluation of reinforcement learning (RL) agents on text-based games.
★ 1.4krllm. Jupyter Notebook
★ 407Douyin_TikTok_Download_API. 🚀「Douyin_TikTok_Download_API」是一个开箱即用的高性能异步抖音、快手、TikTok、Bilibili数据爬取工具,支持API调用,在线批量解析及下载。
★ 19kinspector. Visual testing tool for MCP servers
★ 11kCoAct. Python
★ 32DeepResearch. Tongyi Deep Research, the Leading Open-source Deep Research Agent
★ 20kSearch-R1. Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
★ 5.2kcontrol-arena. ControlArena is a collection of settings, model organisms and protocols - for running control experiments.
★ 214DVWA. Damn Vulnerable Web Application (DVWA)
★ 13kEduVisAgent. [ICLR'26] EduVisAgent: A Benchmark and Multi-Agent Framework for Pedagogical Visualization
★ 30agent-workflow-memory. AWM: Agent Workflow Memory
★ 450OpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58kWebVoyager. Code for "WebVoyager: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models"
★ 1.1kAutoDAN-Turbo. [ICLR 2025 Spotlight] The official implementation of our ICLR2025 paper "AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs".
★ 383Kimi-K2. Kimi K2 is the large language model series developed by Moonshot AI team
★ 11kOSWorld. [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
★ 3kWebRL. Building Open LLM Web Agents with Self-Evolving Online Curriculum RL
★ 537llama_index. LlamaIndex is the leading document agent and OCR platform
★ 51kRAG-FiT. Framework for enhancing LLMs for RAG tasks using fine-tuning.
★ 768gorilla. Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
★ 13kRAG-Retrieval. Unify Efficient Fine-tuning of RAG Retrieval, including Embedding, ColBERT, ReRanker.
★ 1.1kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kSkyRL. SkyRL: A Modular Full-stack RL Library for LLMs
★ 2.1kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kUI-TARS. Pioneering Automated GUI Interaction with Native Agents
★ 11kOnline-Mind2Web. An Illusion of Progress? Assessing the Current State of Web Agents
★ 192AgentSynth. [ICLR 2026] AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
★ 49SimWorld. SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
★ 741verl-agent. verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
★ 2.2kWebDreamer. [TMLR'25] "Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents"
★ 104FinRobot. FinRobot: An Open-Source AI Agent Platform for Financial Applications using LLMs 🚀 🚀 🚀
★ 7.7klangflow. Langflow is a powerful tool for building and deploying AI-powered agents and workflows.
★ 153kAgent-R1. Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
★ 1.6kautomation-mcp. Control your Mac with detailed mouse, keyboard, screen, and window management capabilities.
★ 415AgentsMeetRL. Awesome List for Agentic RL
★ 1.7kmobile-mcp. Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)
★ 5.7kXcodeBuildMCP. A Model Context Protocol (MCP) server and CLI that provides tools for agent use when working on iOS and macOS projects.
★ 6.2kAwesome-VLA-Robotics. A comprehensive list of excellent research papers, models, datasets, and other resources on Vision-Language-Action (VLA) models in robotics.
★ 488ReTool. Python
★ 388One-Shot-RLVR. [NeurIPS 2025] Reinforcement Learning for Reasoning in Large Language Models with One Training Example
★ 444damn-vulnerable-MCP-server. Damn Vulnerable MCP Server
★ 1.3kPaper2Poster. [NeurIPS 2025] Open-source Multi-agent Poster Generation from Papers
★ 3.9kdyad. Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
★ 21kAwesome-Diffusion-LLMs.
★ 5AppAgent. AppAgent: Multimodal Agents as Smartphone Users, an LLM-based multimodal agent framework designed to operate smartphone apps.
★ 6.8kLLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kMMaDA. MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
★ 1.7kDoxBench. [ICLR 2026] The official code for "Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models"
★ 30MMBench. Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"
★ 307awesome-multi-modal-reinforcement-learning. A curated list of Multi-Modal Reinforcement Learning resources (continually updated)
★ 617Awesome-RL-for-LRMs. A Survey of Reinforcement Learning for Large Reasoning Models
★ 2.5kBLIP3o. Official implementation of BLIP3o-Series
★ 1.7kinvariant. Guardrails for secure and robust agent development
★ 436codex. Lightweight coding agent that runs in your terminal
★ 103kWeClone. 🚀 One-stop solution for creating your AI twin from chat history 💡 Fine-tune LLMs with your chat logs to capture your unique style, then bind to a chatbot to bring your digital self to life.
★ 18kAwesome-Embodied-AI-Safety. Focused on the safety and security of Embodied AI
★ 113OpenManus. OpenManus is an open-source initiative to replicate the capabilities of the Manus AI agent, a state-of-the-art general-purpose AI developed by Monica, which excels in autonomously executing complex tasks.
★ 922T2VSafetyBench. Python
★ 28Qwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kx-teaming. Python
★ 67Qwen-Agent. Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.
★ 17kagent-scan. Security scanner for AI agents, MCP servers and agent skills.
★ 2.8kAI-Infra-Guard. A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
★ 4.3kAutoPolicy. C
★ 5web-ui. 🖥️ Run AI Agent in your browser.
★ 16kSWE-agent. SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
★ 20kEasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kHybrid-VLA. HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
★ 352ShieldLM. ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors [EMNLP 2024 Findings]
★ 231PocketFlow. Pocket Flow: 100-line LLM framework. Let Agents build Agents!
★ 11kmcp-logic. Fully functional AI Logic Calculator utilizing Prover9/Mace4 via Python based Model Context Protocol (MCP-Server)- tool for Windows, Linux, Claude App etc
★ 45langchain. The agent engineering platform.
★ 143kopenai-agents-python. A lightweight, powerful framework for multi-agent workflows
★ 28kfastmcp. 🚀 The fast, Pythonic way to build MCP servers and clients.
★ 27kawesome-mcp-servers. A collection of MCP servers.
★ 92kLSDBench. A benchmark that focuses on the sampling dilemma in long-video tasks. Through well-designed tasks, it evaluates the sampling efficiency of long-video VLMs. (ICCV2025)
★ 28VAGEN. World model reasoning RL for multi-turn VLM agents
★ 489sweet_rl. Benchmark and research code for the paper SWEET-RL Training Multi-Turn LLM Agents onCollaborative Reasoning Tasks
★ 271LLM-Multistep-Jailbreak. Code for Findings-EMNLP 2023 paper: Multi-step Jailbreaking Privacy Attacks on ChatGPT
★ 37PrivaCI-Bench. Jupyter Notebook
★ 23Isaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kAutomated-Multi-Turn-Jailbreaks. Python
★ 140swarm. Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
★ 22kowl. 🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
★ 20kmcp-agent. Build effective agents using Model Context Protocol and simple workflow patterns
★ 8.5kASVS. Application Security Verification Standard
★ 3.5kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kinstruct-video-to-video. Python
★ 134Video-R1. Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]
★ 884mem-kk-logic. On Memorization of Large Language Models in Logical Reasoning
★ 79hybrid_space. Source code of our TPAMI'21 paper Dual Encoding for Video Retrieval by Text and CVPR'19 paper Dual Encoding for Zero-Example Video Retrieval.
★ 88GuardBench. A Python library for guardrail models evaluation.
★ 38inspect_evals. Collection of evals for Inspect AI
★ 610agent-attack. [ICLR 2025] Dissecting adversarial robustness of multimodal language model agents
★ 140private-evolution-papers. The collection of papers about Private Evolution
★ 18PlotNeuralNet. Latex code for making neural networks diagrams
★ 25kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kagent_prm. Python
★ 60Step-Video-T2V. Python
★ 3.2kLiger-Kernel. Efficient Triton Kernels for LLM Training
★ 6.5kopen-r1-multimodal. A fork to add multimodal model training to open-r1
★ 1.6kMJ-Video. [NeurIPS'25 Spotlight] MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
★ 20VideoLLaMA3. Frontier Multimodal Foundation Models for Image and Video Understanding
★ 1.2kopenpi. Python
★ 13kRAGEN. RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.
★ 2.8kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kVisionReward. [AAAI 2026] VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
★ 422LlamaGym. Fine-tune LLM agents with online reinforcement learning
★ 1.3kcs257-students.
★ 4ScribeAgent. Code for ScribeAgent paper
★ 63BrowserGym. 🌎💪 BrowserGym, a Gym environment for web task automation
★ 1.3kwebarena. Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
★ 1.6kmath. The MATH Dataset (NeurIPS 2021)
★ 1.4kalpaca_eval. An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
★ 2kST-WebAgentBench. A Benchmark for Evaluating Safety and Trustworthiness in Web Agents for Enterprise Scenarios
★ 25Formal-LLM. Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
★ 137Autonomous-Agents. Autonomous Agents (LLMs) research papers. Updated Daily.
★ 1.4kautomata. A Python library for simulating finite automata, pushdown automata, and Turing machines
★ 405z3. The Z3 Theorem Prover
★ 13kknowledge-harvest-from-lms. ACL 2023 (Findings) - BertNet: Harvesting Knowledge Graphs from Pretrained Language Models
★ 106DeepSeek-V3. Python
★ 104kpython-sdk. The official Python SDK for Model Context Protocol servers and clients
★ 24kair-bench-2024. AIR-Bench 2024 is a safety benchmark that aligns with emerging government regulations and company policies
★ 30search-agents. Code for the paper 🌳 Tree Search for Language Model Agents
★ 223SafeAgentBench. Codes for paper "SafeAgentBench: A Benchmark for Safe Task Planning of \\ Embodied LLM Agents"
★ 74ManipMob-MMKG. Python
★ 15ShowUI. [CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.
★ 1.9kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kaug-pe. [ICML 2024 Spotlight] Differentially Private Synthetic Data via Foundation Model APIs 2: Text
★ 61DPSDA. Private Evolution: Generating DP Synthetic Data without Model Training [ICML 2026, ICLR 2024, ICML 2024 Spotlight]
★ 116RedCode. [NeurIPS'24] RedCode: Risky Code Execution and Generation Benchmark for Code Agents
★ 86SafeWatch. [ICLR 2025] Official implementation for "SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations"
★ 45DropoutDecoding. [NeurIPS 2025] Official Implementation for "Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding"
★ 22textgrad. TextGrad: Automatic ''Differentiation'' via Text -- using large language models to backpropagate textual gradients. Published in Nature.
★ 3.7k