Reinforcement Learning and Reasoning for Coding Agents
ma-gym. A collection of multi agent environments based on OpenAI gym.
632muzero-pytorch. Pytorch Implementation of MuZero
356minimal-marl. Minimal implementation of multi-agent reinforcement learning algorithms
59mmn. Moore Machine Networks (MMN): Learning Finite-State Representations of Recurrent Policy Networks
52visTorch. Interacting with Latent Space of AutoEncoder
21dream-and-search. Code for "Dream and Search to Control: Latent Space Planning for Continuous Control"
12conformal. Conformal prediction is a framework for providing accuracy guarantees on the predictions of a base predictor
9gym-cartpole-continuous. CartPole env. with continuous action space
7marl-pytorch. Pytorch Implementations of Multi Agent Reinforcement Learning(marl) algorithms
6gym_x. Gym environments for capture properties of hidden states(hx) of recurrent networks.
5variable-td3. Learning n-step actions for control tasks
4opcc. Benchmark for "Offline Policy Comparison with Confidence"
3deep-conformal. Applying Conformal Prediction over Deep Neural Nets
3policybazaar. A collection of multi-quality policies for continuous control tasks.
3maze-world. Random maze environments with different size and complexity for reinforcement learning research.
2opcc-baselines. Baselines for "Offline Policy Comparison with Confidence"
1pfa. Policy Fusion Architecture (PFA): We investigate policy gradient approaches for reward decomposition in reinforcement Learning
1tensorboard2seaborn. Plot Tensorflow Summary Event in a Beautiful Way 🌈
1vpn. PyTorch implementation of Value Prediction Network (VPN) :construction: :construction_worker:
1