This is your work, valued
PhD student in ML @ GT
Adaptive-Preference-Scaling. [NeurIPS 2024] Adaptive Preference Scaling for Reinforcement Learning with Human Feedback
AdaSketch-Newton. [ICML 2023] Julia implementation of AdaSketch-Newton
CLNR. Paper: https://arxiv.org/abs/2305.00623
Think-RM. [NeurIPS 2025] Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
TorchCode. 🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
strands-env. A framework for building agent environments for RL training and evaluation with Strands Agents.
strands-sglang. SGLang model provider of Strands Agents for on-policy agentic RL training.
OpenKimi. [ICML2026] Reproduce Kimi K1.5/K2 RL algorithm and rollout system
SDPO. Reinforcement Learning via Self-Distillation (SDPO)
uncertainty-router. [NeurIPS 2025] Ask a Strong LLM Judge when Your Reward Model is Uncertain
LoftQ. Python
triton_tutorial. Tutorials for Triton, a language for writing gpu kernels
es-at-scale. This repo contains the source code for the paper "Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning"
verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
math-evaluation-harness. A simple toolkit for benchmarking LLMs on mathematical reasoning tasks. 🧮✨
DFT. Discriminative Fine-tuning of LLMs without reward models and human preference data
RM-Bench. [ICLR 25 Oral] RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
DFT. Discriminative Finetuning of Generative Large Language Models
OpenRLHF. An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
sweet-watermark. Official repository of the paper: Who Wrote this Code? Watermarking for Code Generation (ACL 2024)
LMFlow. An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.