This is your work, valued
Research Scientist
nlp-gym. NLPGym - A toolkit to develop RL agents to solve NLP tasks.
★ 203pytorch-optimize. A simple black-box optimization framework to train your pytorch models for optimizing non-differentiable objectives
★ 11spsa-optimization. Repository to implement SPSA
★ 5google-word2vec-demo. A demo of using google's pre-trained word2vec model
★ 3Reacher-2D. To train an agent to reach an object using geometrical solution
★ 2spsa-policy-rl. Python
★ 1minimalistic-ml. A collection of basic ML programs
★ 1novelty-guided-rl. Python
★ 1rllm. Democratizing Reinforcement Learning for LLMs
★ 5.8kART. Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
★ 11ksearch-and-learn. Recipes to scale inference-time compute of open models
★ 1.1kwildguard. Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
★ 132MoRA. MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning
★ 362drpo. Dateset Reset Policy Optimization
★ 30outlines. Structured Outputs
★ 15kllm-swarm. Manage scalable open LLM inference endpoints in Slurm clusters
★ 291tril. Python
★ 132al-folio. A beautiful, simple, clean, and responsive Jekyll theme for academics
★ 16kInteractiveTextGeneration. Python
★ 34BIG-bench. Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models
★ 3.2kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kcomposer. Supercharge Your Model Training
★ 5.5kfluidml. FluidML is a lightweight framework for developing machine learning pipelines.
★ 22dialog-eval. Evaluate your dialog model with 17 metrics! (see paper)
★ 97remote-jobs. Source for remoteintech.company — a community-maintained directory of remote-friendly tech companies
★ 41klanguage. Shared repository for open-sourced projects from the Google AI Language team.
★ 1.8kWebShop. [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
★ 577pixel. Research code for pixel-based encoders of language (PIXEL)
★ 348pomdp-baselines. Simple (but often Strong) Baselines for POMDPs in PyTorch, ICML 2022
★ 347zenml. ZenML 🙏: One AI Platform from Pipelines to Agents. https://zenml.io.
★ 5.5kcog. Containers for machine learning
★ 9.5karxiv-latex-cleaner. arXiv LaTeX Cleaner: Easily clean the LaTeX code of your paper to submit to arXiv
★ 7kRL4RS. A Real-World Benchmark for Reinforcement Learning based Recommender System
★ 237GEM-metrics. Automatic metrics for GEM tasks
★ 69leaqi. Active Imitation Learing with Noisy Guidance
★ 10metadict. MetaDict is a powerful dict subclass enabling (nested) attribute-style item access/assignment and IDE autocompletion support.
★ 38reinforcement-learning-an-introduction. Python Implementation of Reinforcement Learning: An Introduction
★ 15khaystack. Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search, and conversational systems.
★ 26kannoy. Approximate Nearest Neighbors in C++/Python optimized for memory usage and loading/saving to disk
★ 14kTextRL. Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
★ 564awesome-rl-nlp. Curated Reinforcement Learning Resources for Natural Language Processing
★ 401tianshou. An elegant PyTorch deep reinforcement learning library.
★ 11kurl_benchmark. Python
★ 368esn-for-named-entity-recognition. Python
★ 2SimCSE. [EMNLP 2021] SimCSE: Simple Contrastive Learning of Sentence Embeddings https://arxiv.org/abs/2104.08821
★ 3.7kapplied-ml. 📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.
★ 30ktorchtyping. Type annotations and dynamic checking for a tensor's shape, dtype, names, etc.
★ 1.5kSSRL. Self-Supervised Reinforcement Learning
★ 8dl-visuals. Over 200 figures and diagrams of the most popular deep learning architectures and layers FREE TO USE in your blog posts, slides, presentations, or papers.
★ 1.5kalfworld. ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
★ 815rl-exploration. Reinforcement Learning papers on exploration methods.
★ 19Chrome-T-Rex-Rush. A python remake of the famous T-Rex game from Google chrome
★ 91RL4Recsys. paper list in the area of reinforcenment learning for recommendation systems
★ 25dataset-sts. Semantic Text Similarity Dataset Hub
★ 730awesome-public-datasets. A topic-centric list of HQ open datasets.
★ 78ktrl. Train transformer language models with reinforcement learning.
★ 19kecco. Explain, analyze, and visualize NLP language models. Ecco creates interactive visualizations directly in Jupyter notebooks explaining the behavior of Transformer-based language models (like GPT2, BERT, RoBERTA, T5, and T0).
★ 2.1kdeep-learning-project-template. Pytorch Lightning code guideline for conferences
★ 1.3kTD3. Author's PyTorch implementation of TD3 for OpenAI gym tasks
★ 2.1kautonomous-learning-library. A PyTorch library for building deep reinforcement learning agents.
★ 656pfrl. PFRL: a PyTorch-based deep reinforcement learning library
★ 1.3kprocgen. Procgen Benchmark: Procedurally-Generated Game-Like Gym-Environments
★ 1.2kcode-for-paper. Python
★ 113SLM-Lab. Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book "Foundations of Deep Reinforcement Learning".
★ 1.4kGriddly. A grid-world game engine for game AI research
★ 259simple-bert. A simple PyTorch implementation of BERT, complete with pretrained models and training scripts.
★ 44mlgym. Python
★ 20outlierhub. Python
★ 4PPO-PyTorch. Minimal implementation of clipped objective Proximal Policy Optimization (PPO) in PyTorch
★ 2.4kdoccano. Open source annotation tool for machine learning practitioners.
★ 11kcaptum. Model interpretability and understanding for PyTorch
★ 5.7kstable-baselines3. PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
★ 14kdatastack. a stream-based file storage solution for machine learning datasets.
★ 11papers-with-annotations. Research papers with annotations, illustrations and explanations
★ 826pg-is-all-you-need. Policy Gradient is all you need! A step-by-step tutorial for well-known PG methods.
★ 1kminimalRL. Implementations of basic RL algorithms with minimal lines of codes! (pytorch based)
★ 3.2krainbow-is-all-you-need. Rainbow is all you need! A step-by-step tutorial from DQN to Rainbow
★ 2ktext-dedicom-paper. Source code to the paper "Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM"
★ 2Miniworld. Simple and easily configurable 3D FPS-game-like environments for reinforcement learning
★ 773dashifyML. A lightweight tool to manage and track your large scale machine leaning experiments
★ 7cleanrl. High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
★ 10kPySyft. Perform data science on data that remains in someone else's server
★ 9.9kRankNet. My (slightly modified) Keras implementation of RankNet and PyTorch implementation of LambdaRank.
★ 247reinforcement-learning. Personal experiments on Reinforcement Learning
★ 119pytorch-A3C. Simple A3C implementation with pytorch + multiprocessing
★ 659packing-unpacking-pytorch-minimal-tutorial. Minimal tutorial on packing and unpacking sequences in pytorch
★ 207gym-gridworlds. Gridworld environments for OpenAI gym.
★ 78Minigrid. Simple and easily configurable grid world environments for reinforcement learning
★ 2.5kPyTorch-RL. PyTorch implementation of Deep Reinforcement Learning: Policy Gradient methods (TRPO, PPO, A2C) and Generative Adversarial Imitation Learning (GAIL). Fast Fisher vector product TRPO.
★ 1.3krllab. rllab is a framework for developing and evaluating reinforcement learning algorithms, fully compatible with OpenAI Gym.
★ 3.1kai-safety-gridworlds. This is a suite of reinforcement learning environments illustrating various safety properties of intelligent agents.
★ 636pytorch-goodies. PyTorch Boilerplate For Research
★ 607trpo. Trust Region Policy Optimization with TensorFlow and OpenAI Gym
★ 364ml-agents. The Unity Machine Learning Agents Toolkit (ML-Agents) is an open-source project that enables games and simulations to serve as environments for training intelligent agents using deep reinforcement learning and imitation learning.
★ 20kdm_control. Google DeepMind's software stack for physics-based simulation and Reinforcement Learning environments, using MuJoCo.
★ 4.7kpractical-pytorch. Go to https://github.com/pytorch/tutorials - this repo is deprecated and no longer maintained
★ 4.5kevostra. A fast Evolution Strategy implementation in Python
★ 272reinforcement-learning. Minimal and Clean Reinforcement Learning Examples
★ 3.7kstochsearch. Stochastic search algorithms using python multiprocessing
★ 11spsa. Simultaneous perturbation stochastic approximation Python code
★ 32climin. Optimizers for machine learning
★ 185pytorch-examples. Simple examples to introduce PyTorch
★ 4.9kpytorch.rl.learning. for learning reinforcement learning using PyTorch.
★ 64PyTorch-Tutorial. Build your neural network easy and fast, 莫烦Python中文教学
★ 8.5k