This is your work, valued
NorMuon. Official Implementation for NorMuon paper
SMASH. Code for Paper "Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process"
SMURF-THP. Python
imitation. Python
Awesome-ML-SYS-Tutorial. My learning notes for ML SYS.
mini-swe-agent. The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
SWE-QA-Bench. Python
qwen_3_chat_templates. Alternative chat templates for Qwen 3 8B. Useful for multi-turn RL
GaLore. GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
Muon. Muon is an optimizer for hidden layers in neural networks
QeRL. [ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
opt-robust. [TMLR] Understanding the robustness difference between stochastic gradient descent and adaptive gradient methods.
modded-nanogpt. NanoGPT (124M) in 90 seconds
ArchScale. Simple & Scalable Pretraining for Neural Architecture Research
ProLong. Homepage for ProLong (Princeton long-context language models) and paper "How to Train Long-Context Language Models (Effectively)"
LongPPL. Code for ICLR 2025 Paper "What is Wrong with Perplexity for Long-context Language Modeling?"
oat-zero. A lightweight reproduction of DeepSeek-R1-Zero with indepth analysis of self-reflection behavior.
Adaptive-Preference-Scaling. [NeurIPS 2024] Adaptive Preference Scaling for Reinforcement Learning with Human Feedback
LoftQ. Python
FinBERT-MRC. Python
ckiptagger. CKIP Neural Chinese Word Segmentation, POS Tagging, and NER
torchstat. Model analyzer in PyTorch
InFi. [MobiCom 2022] InFi is a library for building input filters for resource-efficient inference.
attention-is-all-you-need-pytorch. A PyTorch implementation of the Transformer model in "Attention is All You Need".
gsoc2019.
primal. Parametric Simplex Method for Sparse Learning