This is your work, valued
compute-optimal-tts. Official codebase for "Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling".
288Awesome-Process-Reward-Models. A comprehensive collection of process reward models.
176GenPRM. [AAAI 2026] Official codebase for "GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning".
102MRN. [NeurIPS 2022] Official codebase for "Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning".
26AttnRL. [ICLR 2026] Official codebase for "Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models"
14RL_parking. HTML
14genrm-critiques. GenRM-CoT: Data release for verification rationales
2OpenRLHF-fork. An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & Mixtral)
2RLHF-Reward-Modeling. Recipes to train reward model for RLHF.
2dpss-exp3-VC-BNF. Voice Conversion Experiments for THUHCSI Course : <Digital Processing of Speech Signals>
2Awesome-RL-Reasoning-Recipes. Awesome RL Reasoning Recipes ("Triple R")
2epic. Implements the Equivalent-Policy Invariant Comparison (EPIC) distance for reward functions.
1evaluating-rewards. Library to compare and evaluate reward functions
1CLIP4Clip. An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
1LLaMA-VID. Official Implementation for LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
1factor-world. Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (2023)
1VTF_PAR. [CVPR-2023 Workshop@NFVLR] Official PyTorch implementation of Learning CLIP Guided Visual-Text Fusion Transformer for Video-based Pedestrian Attribute Recognition
1eai-vc. The repository for the largest and most comprehensive empirical study of visual foundation models for Embodied AI (EAI).
1rl-teacher-tf. Open source implementation of "Deep Reinforcement Learning from Human Preferences", updating with evolutionary strategies and augmented morphologies from "Reinforcement Learning for Improved Agent Design"
1LLaVA. Large Language-and-Vision Assistant built towards multimodal GPT-4 level capabilities.
1diffuser. Code for the paper "Planning with Diffusion for Flexible Behavior Synthesis"
1mPLUG-Owl. mPLUG-Owl🦉: Modularization Empowers Large Language Models with Multimodality
1PyRep. A toolkit for robot learning research.
1rl-teacher-pytorch. Python
1VideoChat. Python
1Video-LLaMA. Video-LLaMA: An Instruction-Finetuned Visual Language Model for Video Understanding
1metaworld. Collections of robotics environments geared towards benchmarking multi-task and meta reinforcement learning
1RLBench. A large-scale benchmark and learning environment.
1open_flamingo. An open-source framework for training large multimodal models.
1