This is your work, valued
Self-Distillation. Python
★ 664Value-Augmented-Sampling. Python
★ 20PReF_code. Python
★ 6Bi_Manual_Robot. Jupyter Notebook
★ 2sbon. Python
★ 2multi_ref. Python
★ 2rl_razor_rebuttal.
★ 1PoggioAI_MSc. Our math research agent
★ 34OpenClaw-RL. OpenClaw-RL: Train any agent simply by talking
★ 5.6ktrl. Train transformer language models with reinforcement learning.
★ 19kAReaL. The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.6kMath-Verify. Python
★ 1.2klerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kwebarena. Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
★ 1.6kValue-Augmented-Sampling. Python
★ 20alpaca_farm. A Simulation Framework for RLHF and alternatives
★ 3alpaca_farm. A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
★ 845vqtorch. Python
★ 145cleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
★ 10kllm-foundry. LLM training code for Databricks foundation models
★ 4.4kDeepSpeedExamples. Example models using DeepSpeed
★ 6.8kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kawesome-RLHF. A curated list of reinforcement learning with human feedback resources (continually updated)
★ 4.4kStableLM. StableLM: Stability AI Language Models
★ 16kUniverSeg. UniverSeg: Universal Medical Image Segmentation
★ 588DeformableObjectsGrasping. Python
★ 59Implicit-Language-Q-Learning. Official code from the paper "Offline RL for Natural Language Generation with Implicit Language Q Learning"
★ 213DiffHand. [RSS 2021] An End-to-End Differentiable Framework for Contact-Aware Robot Design
★ 160nanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.
★ 62kSelfGPT. A Whatsapp-bot that allows access to GPT3 while also serving as your own memory backup
★ 131awesome-deep-reinforcement-learning. A collection of resources about deep reinforcement learning
★ 26BOReL. Official implementation for the paper "Offline Meta RL - Identifiability Challenges and Effective Data Collection Strategies", NeurIPS 2021.
★ 31rl. A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
★ 3.5kslac.pytorch. PyTorch implementation of Stochastic Latent Actor-Critic(SLAC).
★ 94Deep-Learning-in-Hebrew. ספר מלא בעברית על למידת מכונה ולמידה עמוקה
★ 1kclearml. ClearML - Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution
★ 6.8k