verl-agent. verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
2.2kDrMAS. Dr. MAS is an end-to-end RL training framework for multi-agent LLM systems, supporting the co-training of multiple (heterogeneous) LLMs.
144TimeMaster. Official code for paper "TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning"
68AgentOCR. [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting RL-driven self-compression
38MLF-DSResNet. This is the implementation of MLF & spiking DS-ResNet
16CoSo. Official code for paper "Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning"
15tree-diffusion-planner. Code for the paper "Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation Tree"
7DQN-PG-A2C-DDPG-PPO-SAC-Pytorch. Implementation of DQN-PG-A2C-DDPG-PPO-SAC with Pytorch
3DDPG-with-pytorch. Implementation of DDPG based on Pendulum-v1
3SAC_with_Pytorch. Implementation of soft actor critic (SAC) based on Pendulum-v1
1FasterRCNN-expanded-VOC2007. Python
1