18
Mon
Tue
Wed
Thu
Fri
Sat
Sun
Active 4d ago
San Francisco

Phil Wang

Top 3%
@lucidrains

Phil ships production-ready PyTorch implementations of cutting-edge transformer research, from vit-pytorch and denoising-diffusion-pytorch to x-transformers and PaLM-rlhf-pytorch. His repos have accumulated 182k stars, making him one of the most prolific open source contributors globally.

Vision Transformer. I built image recognition models that work better and faster than older approaches.

25k

PaLM RLHF. I built the code structure to train an AI model using human feedback.

7.9k

X-Transformers. I built a complete AI model toolkit with experimental research features included.

5.9k

Diffusion. I built a working image generation model using a noise-removal approach.

11k

Titans. I built a memory system for AI models that handles long documents efficiently.

2k

Vector Quantize. I built a compression tool that AI image and music systems use.

4k

BYOL. I made a tool that teaches image models from unlabeled photos.

1.9k

Lion. I made a better optimizer for training AI models that Google Brain discovered.

2.2k

Perceiver. I made an AI model that understands images, video, and audio the same way.

1.2k

Band Split RoFormer. I built an AI model that splits songs into separate instrument tracks.

880

Transfusion. I built one AI model that predicts text and generates images together.

1.4k

Rotary Embedding. I made a tool that adds position awareness to AI transformers.

819

Rectified Flow. I built a PyTorch implementation of rectified flow for image generation.

476

Alphafold 3. I made Google's protein-folding AI available for anyone to use.

1.7k

Enformer. I made Deepmind's gene expression AI runnable and customizable in Python.

572

EMA. I built a tool that maintains a smoothed copy of your AI model.

657

Imagen. I built a tool that turns text descriptions into realistic images.

8.4k

MMDiT. I made the multi-modal AI layer from Stable Diffusion 3 as reusable code.

552

Tab Transformer. I made an AI model that learns patterns from spreadsheet data.

1.1k

Slot Attention. I made an AI model that automatically isolates objects in images.

493

Pi Zero. I built a working version of Physical Intelligence's robot foundation model.

582

Clinical Calculator. I trained AI models to apply real clinical calculators from patient data.

315

StyleGAN2. I made an image generator you can train yourself with no coding.

3.8k

Autoregressive Diffusion. I built an image generation tool that creates pictures step by step.

438

Native Sparse Attention. I implemented a faster attention pattern for AI models that skips unnecessary calculations.

811

E2 TTS. I made software that turns text into realistic spoken audio.

516

MLP Mixer. I implemented Google's all-MLP image recognition model in PyTorch.

1.1k

SoundStorm. I built a tool that generates audio in parallel instead of sequentially.

1.5k

AudioLM. I made an AI that generates realistic audio from text descriptions.

2.6k

Ring Attention. I implemented Ring Attention to let AI models process documents that are millions of tokens long.

546

GigaGAN. I built Adobe's latest image generation software in Python.

1.9k

MinGRU. I built a PyTorch implementation of the minGRU neural network.

325

Lightweight GAN. I made an AI image generator that trains in hours on one computer.

1.7k

iTransformer. I built an implementation of iTransformer for time series forecasting.

537

nGPT. I built a PyTorch implementation of Nvidia's normalized GPT architecture.

300

Local Attention. I built a tool that makes AI models process text faster by focusing locally.

503

DALLE2 PyTorch. I built DALL-E 2 as open-source code you can run yourself.

11k

MeshGPT. I built an AI model that generates 3D shapes by learning patterns in how geometry works.

861

Q-Transformer. I built a robot learning system trained on recorded experience.

407

SE3 Transformer. I built an AI model for understanding 3D molecular geometry.

331

Soft MoE. I implemented Soft MoE, a mixture-of-experts technique for AI models.

348

Make-A-Video. I built a text-to-video generator based on Meta's research.

2k

MusicLM. I built a tool that writes original music from text descriptions.

3.3k

Classifier Free Guidance. I wrote a tool that adds text control to AI image models.

544

MagViT2. I implemented MagViT2, a video tokenizer that achieves state-of-the-art results.

668

MEGABYTE. I built a PyTorch implementation of MEGABYTE for processing million-byte sequences.

655

Speculative Decoding. I built techniques to make AI text generation run faster with less wasted computation.

307

Self-Rewarding LM. I built code that trains AI models to evaluate and improve themselves.

1.4k

Video Diffusion. I built an AI tool that generates videos from text descriptions.

1.4k

Voicebox. I made Meta's Voicebox text-to-speech model work in Pytorch.

699
50
Apply