Menlo Park, CA

Costa Huang

Elite
@vwxyzjn

Exploiting physical rewards @periodic. Prev: RL @allenai @huggingface.

cleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)

10k

ppo-implementation-details. The source code for the blog post The 37 Implementation Details of Proximal Policy Optimization

945

portwarden. Create Encrypted Backups of Your Bitwarden Vault with Attachments

632

lm-human-preference-details. RLHF implementation details of OAI's 2019 codebase

198

invalid-action-masking. Source Code for A Closer Look at Invalid Action Masking in Policy Gradient Algorithms

168

summarize_from_feedback_details. Python

164

cleanba. CleanRL's implementation of DeepMind's Podracer Sebulba Architecture for Distributed DRL

125

PPO-Implementation-Deep-Dive. DEPRECATED - please visit https://github.com/vwxyzjn/ppo-implementation-details

46

gym-microrts-paper. The source code for the gym-microrts paper.

44

a2c_is_a_special_case_of_ppo. A2C is a special case of PPO!

23

jupyter_disqus. Add Disqus to your Jupyter notebook.

14

SC2AI. Integrated Tensorforce and OpenAI Gym to train SC II game agents.

13

costa-utils. Python

10

gym-pysc2. Gym wrapper for pysc2

10

envpool-cleanrl. Python

9

free-mujoco-py. MuJoCo is a physics engine for detailed, efficient rigid body simulations with contacts. mujoco-py allows using MuJoCo from Python 3.

9

LeanRL. LeanRL is a fork of CleanRL, where selected PyTorch scripts optimized for performance using compile and cudagraphs.

8

ppo-atari-metrics. Python

8

benchmark-ci. Python

7

action-guidance. Python

6

notablog-starter. The official starter project for Notablog.

5

lm-human-preferences. Code for the paper Fine-Tuning Language Models from Human Preferences

4

cleangpt. Python

4

trl. Train transformer language models with reinforcement learning.

4

minimal-adam-difference. Python

4

quickchat. Python

3

microrts. Java

3

vectorized-value-methods. [WIP] Vectorized architecture for value-based methods such as DQN and DDPG

3

direct-preference-optimization. Reference implementation for DPO (Direct Preference Optimization)

2

envpool-xla-cleanrl. Python

2

entity-ppo-demo. Python

2

transformers. 🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.

2

alignment-handbook. Robust recipes for to align language models with human and AI preferences

2

minimal-uv-deep-ep-gemm-installation. Python

2

Open-Reasoner-Zero. Official Repo for Open-Reasoner-Zero

2

optax. Optax is a gradient processing and optimization library for JAX.

1

hfblog. Public repo for HF blog posts

1

verl. verl: Volcano Engine Reinforcement Learning for LLMs

1

cleanba-test. Python

1

nanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.

1

zero3_min_repro. Python

1

envpool_bug. Python

1

PokemonRedExperiments. Playing Pokemon Red with Reinforcement Learning

1

Sentiment-Analysis-LSTM. Used neural network to classify movie reviews based on sentiment

1

embedding_projector. Python

1

notion-blog. TypeScript

1

CS583. Python

1

awesome-vue. A curated list of awesome things related to Vue.js

1

validate-new-gym-mujoco-envs. Python

1
49
Apply