This is your work, valued
Macaron AI CRO, Mind Lab Director. Focusing on scaling agentic RL.
pytorch_car_caring. Reinforcement Learning for Gym CarRacing-v0 with PyTorch
160dsac. Distributional Soft Actor Critic
63simple-pytorch-rl. Reinforcement Learning Methods with PyTorch
38apo. Average-Reward Reinforcement Learning with Trust Region Methods
11msvpo. The official implementation of "Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning"
2xtma.github.io. SCSS
1vimrc. The ultimate Vim configuration: vimrc
1pytorch-a2c-ppo-acktr-gail. PyTorch implementation of Advantage Actor Critic (A2C), Proximal Policy Optimization (PPO), Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation (ACKTR) and Generative Adversarial Imitation Learning (GAIL).
1ray-maddpg. MADDPG implementation with Ray
1