This is your work, valued

Liangyu Wang

Advanced
@liangyuwang

LLM, CUDA/System, CPU offloading, Distributed training

zo2. ZO2 (Zeroth-Order Offloading): Full Parameter Fine-Tuning 175B LLMs with 18GB GPU Memory [COLM2025]

208

Tiny-FSDP. Tiny-FSDP, a minimalistic re-implementation of the PyTorch FSDP

112

Tiny-DeepSpeed. Tiny-DeepSpeed, a minimalistic re-implementation of the DeepSpeed library

53

Tiny-Megatron. Tiny-Megatron, a minimalistic re-implementation of the Megatron library

32

Tinytron. A minimal, hackable pre-training stack for GPT-style language models

10

FSDP-Canzona. Bring Muon-style optimizers to PyTorch FSDP training.

8

train-large-model-from-scratch. A minimal, hackable pre-training stack for GPT-style language models

7

LLM-Trainers. Trainers for LLM Training. Accepts Huggingface models and datasets

5

Tiny-LLM-Libs.

3

Streaming-Dataloader. A memory-efficient streaming data loader designed for LLM pretraining under limited CPU and GPU memory constraints

3

Tiny-transformers. Python

3

Flash-Attention-Implementation. Implementation of Flash-Attention (both forward and backward) with PyTorch, CUDA, and Triton

3

MultimodalWCE.

3

LLM-Length-Estimation. Python

3

NanoPD. NanoPD, a minimalistic implementation of the disaggregated PD serve

3

LLM-efficient-learning. Find the most efficient way for a specific large language model to learn a specific task

2

optimization-project. Python

2

MetaProfiler. MetaProfiler is a lightweight, structure-agnostic operator-level profiler for PyTorch models that leverages MetaTensor execution to simulate and benchmark individual ops without loading the full model into GPU memory.

2

simple_cuda_kernel. A collection of ultra-simple yet high-performance CUDA kernels.

1

core_scheduler. CoreScheduler: A High-Performance Scheduler for Large Model Training

1

OpenRollout.

1