This is your work, valued
Exploiting physical rewards @periodic. Prev: RL @allenai @huggingface.
cleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
10kppo-implementation-details. The source code for the blog post The 37 Implementation Details of Proximal Policy Optimization
945portwarden. Create Encrypted Backups of Your Bitwarden Vault with Attachments
632lm-human-preference-details. RLHF implementation details of OAI's 2019 codebase
198invalid-action-masking. Source Code for A Closer Look at Invalid Action Masking in Policy Gradient Algorithms
168summarize_from_feedback_details. Python
164cleanba. CleanRL's implementation of DeepMind's Podracer Sebulba Architecture for Distributed DRL
125PPO-Implementation-Deep-Dive. DEPRECATED - please visit https://github.com/vwxyzjn/ppo-implementation-details
46gym-microrts-paper. The source code for the gym-microrts paper.
44a2c_is_a_special_case_of_ppo. A2C is a special case of PPO!
23jupyter_disqus. Add Disqus to your Jupyter notebook.
14SC2AI. Integrated Tensorforce and OpenAI Gym to train SC II game agents.
13costa-utils. Python
10gym-pysc2. Gym wrapper for pysc2
10envpool-cleanrl. Python
9free-mujoco-py. MuJoCo is a physics engine for detailed, efficient rigid body simulations with contacts. mujoco-py allows using MuJoCo from Python 3.
9LeanRL. LeanRL is a fork of CleanRL, where selected PyTorch scripts optimized for performance using compile and cudagraphs.
8ppo-atari-metrics. Python
8benchmark-ci. Python
7action-guidance. Python
6notablog-starter. The official starter project for Notablog.
5lm-human-preferences. Code for the paper Fine-Tuning Language Models from Human Preferences
4cleangpt. Python
4trl. Train transformer language models with reinforcement learning.
4minimal-adam-difference. Python
4quickchat. Python
3microrts. Java
3vectorized-value-methods. [WIP] Vectorized architecture for value-based methods such as DQN and DDPG
3direct-preference-optimization. Reference implementation for DPO (Direct Preference Optimization)
2envpool-xla-cleanrl. Python
2entity-ppo-demo. Python
2transformers. 🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
2alignment-handbook. Robust recipes for to align language models with human and AI preferences
2minimal-uv-deep-ep-gemm-installation. Python
2Open-Reasoner-Zero. Official Repo for Open-Reasoner-Zero
2optax. Optax is a gradient processing and optimization library for JAX.
1hfblog. Public repo for HF blog posts
1verl. verl: Volcano Engine Reinforcement Learning for LLMs
1cleanba-test. Python
1nanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.
1zero3_min_repro. Python
1envpool_bug. Python
1PokemonRedExperiments. Playing Pokemon Red with Reinforcement Learning
1Sentiment-Analysis-LSTM. Used neural network to classify movie reviews based on sentiment
1embedding_projector. Python
1notion-blog. TypeScript
1CS583. Python
1awesome-vue. A curated list of awesome things related to Vue.js
1validate-new-gym-mujoco-envs. Python
1