This is your work, valued

PyTorch researcher implementing cutting-edge transformer architectures and attention mechanisms from academic papers.

Transformer architectures and attention mechanismsWorld models and model-based reinforcement learningDiffusion and generative modelsOptimization and training techniquesRobotics and embodied AIVector quantization and discrete representations

inferred from the public footprint

21
Wed
Thu
Fri
Sat
Sun
Mon
Yesterday
Active 1d ago
San Francisco

Phil Wang

Top 4%
@lucidrains

Phil ships production-ready PyTorch implementations of cutting-edge transformer research, from vit-pytorch and denoising-diffusion-pytorch to x-transformers and PaLM-rlhf-pytorch. His repos have accumulated 182k stars, making him one of the most prolific open source contributors globally.

Vision Transformer. I built image recognition models that work better and faster than older approaches.

25k

PaLM RLHF. I built the code structure to train an AI model using human feedback.

7.9k

X-Transformers. I built a complete AI model toolkit with experimental research features included.

5.9k

Diffusion. I built a working image generation model using a noise-removal approach.

11k

Titans. I built a memory system for AI models that handles long documents efficiently.

2k

Vector Quantize. I built a compression tool that AI image and music systems use.

4k

BYOL. I made a tool that teaches image models from unlabeled photos.

1.9k

Lion. I made a better optimizer for training AI models that Google Brain discovered.

2.2k

Perceiver. I made an AI model that understands images, video, and audio the same way.

1.2k

Band Split RoFormer. I built an AI model that splits songs into separate instrument tracks.

881

Transfusion. I built one AI model that predicts text and generates images together.

1.4k

Rotary Embedding. I made a tool that adds position awareness to AI transformers.

819

Rectified Flow. I built a PyTorch implementation of rectified flow for image generation.

476

Alphafold 3. I made Google's protein-folding AI available for anyone to use.

1.7k

Enformer. I made Deepmind's gene expression AI runnable and customizable in Python.

572

EMA. I built a tool that maintains a smoothed copy of your AI model.

658

Imagen. I built a tool that turns text descriptions into realistic images.

8.4k

MMDiT. I made the multi-modal AI layer from Stable Diffusion 3 as reusable code.

552

Tab Transformer. I made an AI model that learns patterns from spreadsheet data.

1.1k

Slot Attention. I made an AI model that automatically isolates objects in images.

493

Pi Zero. I built a working version of Physical Intelligence's robot foundation model.

582

Clinical Calculator. I trained AI models to apply real clinical calculators from patient data.

315

StyleGAN2. I made an image generator you can train yourself with no coding.

3.8k

Autoregressive Diffusion. I built an image generation tool that creates pictures step by step.

438

Native Sparse Attention. I implemented a faster attention pattern for AI models that skips unnecessary calculations.

811

E2 TTS. I made software that turns text into realistic spoken audio.

516

MLP Mixer. I implemented Google's all-MLP image recognition model in PyTorch.

1.1k

SoundStorm. I built a tool that generates audio in parallel instead of sequentially.

1.5k

AudioLM. I made an AI that generates realistic audio from text descriptions.

2.6k

Ring Attention. I implemented Ring Attention to let AI models process documents that are millions of tokens long.

546

GigaGAN. I built Adobe's latest image generation software in Python.

1.9k

MinGRU. I built a PyTorch implementation of the minGRU neural network.

325

Lightweight GAN. I made an AI image generator that trains in hours on one computer.

1.7k

iTransformer. I built an implementation of iTransformer for time series forecasting.

537

nGPT. I built a PyTorch implementation of Nvidia's normalized GPT architecture.

300

Local Attention. I built a tool that makes AI models process text faster by focusing locally.

503

DALLE2 PyTorch. I built DALL-E 2 as open-source code you can run yourself.

11k

MeshGPT. I built an AI model that generates 3D shapes by learning patterns in how geometry works.

861

Q-Transformer. I built a robot learning system trained on recorded experience.

407

SE3 Transformer. I built an AI model for understanding 3D molecular geometry.

331

Soft MoE. I implemented Soft MoE, a mixture-of-experts technique for AI models.

348

Make-A-Video. I built a text-to-video generator based on Meta's research.

2k

MusicLM. I built a tool that writes original music from text descriptions.

3.3k

Classifier Free Guidance. I wrote a tool that adds text control to AI image models.

544

MagViT2. I implemented MagViT2, a video tokenizer that achieves state-of-the-art results.

668

MEGABYTE. I built a PyTorch implementation of MEGABYTE for processing million-byte sequences.

655

Speculative Decoding. I built techniques to make AI text generation run faster with less wasted computation.

307

Self-Rewarding LM. I built code that trains AI models to evaluate and improve themselves.

1.4k

Video Diffusion. I built an AI tool that generates videos from text descriptions.

1.4k

Voicebox. I made Meta's Voicebox text-to-speech model work in Pytorch.

699