This is your work, valued
Research Scientist, Google DeepMind
mtan. The implementation of "End-to-End Multi-Task Learning with Attention" [CVPR 2019].
★ 731clarity-template. Clarity: A Minimalist Website Template for AI Research
★ 225maxl. The implementation of "Self-Supervised Generalisation with Meta Auxiliary Learning" [NeurIPS 2019].
★ 174reco. The implementation of "Bootstrapping Semantic Segmentation with Regional Contrast" [ICLR 2022].
★ 170auto-lambda. The Implementation of "Auto-Lambda: Disentangling Dynamic Task Relationships" [TMLR 2022].
★ 140minimal-isaac-gym. A Minimal Example of Isaac Gym with DQN and PPO.
★ 112shape-adaptor. The implementation of "Shape Adaptor: A Learnable Resizing Module" [ECCV 2020].
★ 71vsl. The implementation of "Learning a Hierarchical Latent-Variable Model of 3D Shapes" [3DV 2018].
★ 36facade-urban-analysis. Building Facade-based City Classification from Aerial-View Images
★ 6TorchJD. Library for Jacobian descent with PyTorch. It enables the optimization of neural networks with multiple losses (e.g. multi-task learning).
★ 391EditCtrl. [CVPR 2026] EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
★ 46Captain-Safari. Offical Implementation of Captain-Safari [CVPR 2026]
★ 46giscus. A commenting system powered by GitHub Discussions. :octocat: :speech_balloon: :gem:
★ 12kLTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kUCPE. 📷 [CVPR'26] Camera-controlled text-to-video generation, now with intrinsics, distortion and orientation control!
★ 213GEN3C. [CVPR 2025 Highlight] GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
★ 1.4kmin-pi-flow. Python
★ 562d-gaussian-splatting. [SIGGRAPH'24] 2D Gaussian Splatting for Geometrically Accurate Radiance Fields
★ 3.3kScaling-Diffusion-Transformers-muP. [NeurIPS 2025] Official implementation for our paper "Scaling Diffusion Transformers Efficiently via μP".
★ 100Sana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kcolor-alignment-in-diffusion. Official Github repository for the CVPR 2025 paper "Color Alignment in Diffusion"
★ 17ControlVideo. [ICLR 2024] Official pytorch implementation of "ControlVideo: Training-free Controllable Text-to-Video Generation"
★ 864Geo4D. [ICCV 2025 Highlight] Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
★ 437CFG-Zero-star. Official repo for CFG-Zero*
★ 715Go-with-the-Flow. The official implementation of CVPR'25 Oral paper "Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise"
★ 1.1kfast3r. [CVPR 2025] Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
★ 1.6kLeffa. [CVPR 2025] Learning Flow Fields in Attention for Controllable Person Image Generation
★ 1.7kslam-handbook-public-release. Release repo for our SLAM Handbook
★ 4.6kopen-oasis. Inference script for Oasis 500M
★ 2.1kStableDelight. StableDelight: Revealing Hidden Textures by Removing Specular Reflections
★ 410transfusion-pytorch. Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
★ 1.4kwhisper-medusa. Whisper with Medusa heads
★ 861flux. Official inference repo for FLUX.1 models
★ 26kmar. PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
★ 1.9kMultiDiffusion. Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
★ 1.1kfast-DiT. Fast Diffusion Models with Transformers
★ 954MOFA-Video. [ECCV 2024] MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model.
★ 764LLM101n. LLM101n: Let's build a Storyteller
★ 38kchameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kGaussianFlow. GaussianFlow: Splatting Gaussian Dynamics for 4D Content Creation (TMLR)
★ 192Real3D. Code for ICCV'2025 "Real3D: Scaling Up Large Reconstruction Models with Real-World Images"
★ 2113dgs-mcmc. [NeurIPS 2024 Spotlight] Implementation of the paper "3D Gaussian Splatting as Markov Chain Monte Carlo"
★ 677Unique3D. [NeurIPS 2024] Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
★ 3.6kChatTTS. A generative speech model for daily dialogue.
★ 40k3DitScene. [ICLR 2025] 3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting
★ 260Diffusion4D. [NeurIPS 2024] Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
★ 344NVS_Solver. Source code of paper "NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer"
★ 320gaussian_surfels. [SIGGRAPH'24] Implementations for "High-quality Surface Reconstruction using Gaussian Surfels".
★ 682unsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek, GLM and other models.
★ 69kcustom-diffusion360. CustomDiffusion360: Customizing Text-to-Image Diffusion with Camera Viewpoint Control
★ 171ohara. Collection of autoregressive model implementation
★ 84OpenCLAY. CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
★ 980VAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kcomo. [ECCV 2024 Oral] COMO: Compact Mapping and Odometry
★ 224CityDreamer. The official implementation of "CityDreamer: Compositional Generative Model of Unbounded 3D Cities". (CVPR 2024)
★ 698super_primitive. [CVPR'24, Demo Track Honourable Mention] SuperPrimitive: Scene Reconstruction at a Primitive Level
★ 204MorpheuS. [CVPR'24] MorpheuS: Neural Dynamic 360° Surface Reconstruction from Monocular RGB-D Video
★ 164Den-SOFT. JavaScript
★ 11GeoWizard. [ECCV'24] GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
★ 938flatten. Pytorch Implementation of FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing (ICLR 2024)
★ 212multimodal. TorchMultimodal is a PyTorch library for training state-of-the-art multimodal multi-task models at scale.
★ 1.7kV3D. [T-PAMI 2025] V3D: Video Diffusion Models are Effective 3D Generators
★ 522GaLore. GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
★ 1.7kMonoGS. [CVPR'24 Highlight & Best Demo Award] Gaussian Splatting SLAM
★ 2.1kmagvit2-pytorch. Implementation of MagViT2 Tokenizer in Pytorch
★ 668TCD. Official Repository of the paper "Trajectory Consistency Distillation"
★ 360T5-Textual-Inversion. Textual Inversion for DeepFloyd IF
★ 61dust3r. DUSt3R: Geometric 3D Vision Made Easy
★ 7.3kDSINE. [CVPR 2024 Oral] Rethinking Inductive Biases for Surface Normal Estimation
★ 920Medusa. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8kkubric. A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
★ 2.8klaion-3d. Collect large 3d dataset and build models
★ 298ml-hypersim. Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
★ 2kMVDiffusion_plusplus. MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction
★ 144Dream2Real. [ICRA 2024] Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
★ 69SiT. Official PyTorch Implementation of "SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers"
★ 1.2kMDT. Masked Diffusion Transformer is the SOTA for image synthesis. (ICCV 2023)
★ 596VBench. [CVPR2024 Highlight] VBench - We Evaluate Video Generation
★ 1.7kDataset. News: the 10k dataset is ready for download.
★ 655PUG. This is the repository for the Photorealistic Unreal Graphics (PUG) datasets for representation learning.
★ 239StableCascade. Official Code for Stable Cascade
★ 6.5kgta. [ICLR'24] GTA: A Geometry-Aware Attention Mechanism for Multi-view Transformers
★ 160ml-mgie. Python
★ 3.9kEscherNet. [CVPR2024 Oral] EscherNet: A Generative Model for Scalable View Synthesis
★ 378Marigold. [CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
★ 3.2kinter. The Inter font family
★ 20kMegatron-LM. Ongoing research training transformer models at scale
★ 17katlas. Code repository for supporting the paper "Atlas Few-shot Learning with Retrieval Augmented Language Models",(https//arxiv.org/abs/2208.03299)
★ 560ImageDream. The code releasing for https://image-dream.github.io/
★ 799LISA. Project Page for "LISA: Reasoning Segmentation via Large Language Model"
★ 2.7kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18klang-segment-anything. SAM with text prompt
★ 2.6kpixelsplat. [CVPR 2024 Oral, Best Paper Runner-Up] Code for "pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction" by David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann
★ 1.3kzerorf. ZeroRF: Fast Sparse View 360° Reconstruction with Zero Pretraining
★ 200mamba. Mamba SSM architecture
★ 19ktransformer_vq. Official implementation of 'Transformer-VQ: Linear-Time Transformers via Vector Quantization' (ICLR 2024)
★ 199AnimateAnyone. Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
★ 15kconsistencydecoder. Consistency Distilled Diff VAE
★ 2.2klatent-consistency-model. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
★ 4.6kWonder3D. Single Image to 3D using Cross-Domain Diffusion for 3D Generation
★ 5.4kCo-SLAM. [CVPR'23] Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM
★ 477mistral-inference. Official inference library for Mistral models
★ 11kswitch-scores. PDF Repository of switch score sheets.
★ 1.2kwaymo-open-dataset. Waymo Open Dataset
★ 3.4kthreestudio. A unified framework for 3D content generation.
★ 7kSyncDreamer. [ICLR 2024 Spotlight] SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
★ 1kMVDream. code page placeholder
★ 569OBELICS. Code used for the creation of OBELICS, an open, massive and curated collection of interleaved image-text web documents, containing 141M documents, 115B text tokens and 353M images.
★ 217visual-spatial-reasoning. [TACL'23] VSR: A probing benchmark for spatial undersranding of vision-language models.
★ 149scalingup. [CoRL 2023] This repository contains data generation and training code for Scaling Up & Distilling Down
★ 414zero123-hf. A diffuser implementation of Zero123. Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV23)
★ 154perceiver-io. A PyTorch implementation of Perceiver, Perceiver IO and Perceiver AR with PyTorch Lightning scripts for distributed training
★ 535stablediffusion-infinity. Outpainting with Stable Diffusion on an infinite canvas
★ 3.9kact3d-chained-diffuser. A unified architecture for multimodal multi-task robotic policy learning.
★ 186One-2-3-45. [NeurIPS 2023] Official code of "One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization"
★ 1.7khiveformer. Python
★ 33MVDiffusion. MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion, NeurIPS 2023 (spotlight)
★ 565perception_test. Jupyter Notebook
★ 255LightGlue. LightGlue: Local Feature Matching at Light Speed (ICCV 2023)
★ 4.7kDepthCov. [CVPR 2023] Learning a Depth Covariance Function
★ 102nope. [CVPR 2024] PyTorch implementation of NOPE: Novel Object Pose Estimation from a Single Image
★ 218OmniObject3D. [ CVPR 2023 Award Candidate ] OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation
★ 531rl. A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
★ 3.5kpoint2vec. [GCPR 2023 | CVPR 2023 Workshop] Self-Supervised Representation Learning on Point Clouds
★ 102FAMO. Official PyTorch Implementation for Fast Adaptive Multitask Optimization (FAMO)
★ 131TaleCrafter. [SIGGRAPH Asia 2023] An interactive story visualization tool that support multiple characters
★ 268qlora. QLoRA: Efficient Finetuning of Quantized LLMs
★ 11ksd-leap-booster. Fast finetuning using a booster model that puts the initial state to a local minimum
★ 113Everything-LLMs-And-Robotics. The world's largest GitHub Repository for LLMs + Robotics
★ 849hmm. Heightmap meshing utility.
★ 613llm-foundry. LLM training code for Databricks foundation models
★ 4.4kdatacomp. DataComp: In search of the next generation of multimodal datasets
★ 787Point-MAE. [ECCV2022] Masked Autoencoders for Point Cloud Self-supervised Learning
★ 638sdf. Parallelized triangle mesh --> continuous signed distance field on CPU
★ 523arnold. [ICCV 2023] ARNOLD: Language-Grounded Robot Manipulation with Continuous Object States in Realistic 3D Scenes
★ 188mmc4. MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
★ 954StableLM. StableLM: Stability AI Language Models
★ 16kinstruct-pix2pix. Python
★ 6.9kmesh-to-depth. Depth map generator for Python written in C++
★ 40shape-adaptor. The implementation of "Shape Adaptor: A Learnable Resizing Module" [ECCV 2020].
★ 71auto-lambda. The Implementation of "Auto-Lambda: Disentangling Dynamic Task Relationships" [TMLR 2022].
★ 140maxl. The implementation of "Self-Supervised Generalisation with Meta Auxiliary Learning" [NeurIPS 2019].
★ 174mtan. The implementation of "End-to-End Multi-Task Learning with Attention" [CVPR 2019].
★ 731reco. The implementation of "Bootstrapping Semantic Segmentation with Regional Contrast" [ICLR 2022].
★ 170genvs.
★ 640segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kMulti-Task-Transformer. Code of ICLR2023 paper "TaskPrompter: Spatial-Channel Multi-Task Prompting for Dense Scene Understanding" and ECCV2022 paper "Inverted Pyramid Multi-task Transformer for Dense Scene Understanding"
★ 326tomesd. Speed up Stable Diffusion with this one simple trick!
★ 1.4kmvp. Masked Visual Pre-training for Robotics
★ 246gpt4all. GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kMake-It-3D. [ICCV 2023] Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior
★ 1.9kpromptplusplus. Jupyter Notebook
★ 73Painter. Painter & SegGPT Series: Vision Foundation Models from BAAI
★ 2.6ktypst. A markup-based typesetting system that is powerful and easy to learn.
★ 55ktactile-dexterity. Official implementation of Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play.
★ 68zero123. Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)
★ 3.1kResizeRight. The correct way to resize images or tensors. For Numpy or Pytorch (differentiable).
★ 566r3m. Pre-training Reusable Representations for Robotic Manipulation Using Diverse Human Video Data
★ 377frankapy. Python interface to control Franka Emika Panda Research Robot Arms.
★ 250oscar. Data-Driven Operational Space Control for Adaptive and Robust Robot Manipulation
★ 140isaacgym-stubs. Isaac Gym Python Stubs for Code Completion
★ 126OpenChatKit. Python
★ 9kTaskMatrix. Python
★ 34kELITE. ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation (ICCV 2023, Oral)
★ 541diffusion_policy. [RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
★ 4.4kpizza. [3DV 2022 (Oral)] Pytorch implementation of "PIZZA: A Powerful Image-only Zero-Shot Zero-CAD Approach to 6 DoF Tracking" paper
★ 77prismer. The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
★ 1.3kgo-explore. Code for Go-Explore: a New Approach for Hard-Exploration Problems
★ 586apriltag-imgs. Pre-generated AprilTag images
★ 576apriltag. Extensions and tweaks to APRIL Robotics Laboratory apriltag C software
★ 193chatgpt_please_improve_my_paper_writing. a thin wrapper of chatgpt for improving paper writing.
★ 253realfusion. Official code for "RealFusion: 360° Reconstruction of Any Object from a Single Image" (CVPR 2023)
★ 564hiveformer-corl. PyTorch implementation of the Hiveformer research paper
★ 48PromptCraft-Robotics. Community for applying LLMs to robotics and a robot simulator with ChatGPT integration
★ 2.1kcalvin. CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 963MiDaS. Code for robust monocular depth estimation described in "Ranftl et. al., Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer, TPAMI 2022"
★ 5.4kFlexLLMGen. Running large language models on a single GPU for throughput-oriented scenarios.
★ 9.4kauxinash. Official implementation of Auxiliary learning as an Bargaining Game.
★ 34ov-seg. This is the official PyTorch implementation of the paper Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP.
★ 758ManiSkill. Manipulation Skill Framework, an open source GPU parallelized robotics simulator and benchmark
★ 3.2kControlNet. Let us control diffusion models!
★ 34komni3d. Code release for "Omni3D A Large Benchmark and Model for 3D Object Detection in the Wild"
★ 855X-Decoder. [CVPR 2023] Official Implementation of X-Decoder for generalized decoding for pixel, image and language
★ 1.3kOpen-Assistant. OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
★ 37kH3. Language Modeling with the H3 State Space Model
★ 525vMAP. [CVPR 2023] vMAP: Vectorised Object Mapping for Neural Field SLAM
★ 372mm-cot. Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
★ 4kSQA3D. [ICLR 2023] SQA3D for embodied scene understanding and reasoning
★ 170D4RL. A collection of reference environments for offline reinforcement learning
★ 1.7klanguage-table. Suite of human-collected datasets and a multi-task continuous control benchmark for open vocabulary visuolinguomotor learning.
★ 363mask-auto-labeler. Python
★ 164iris. Transformers are Sample-Efficient World Models. ICLR 2023, notable top 5%.
★ 898ECON. [CVPR'23, Highlight] ECON: Explicit Clothed humans Optimized via Normal integration
★ 1.2kSL-Decoder. This is the code for the application of our structured lighting system. We provide with structured lighting patterns and the decoder, for recovering the scene depth. In some cases, this depth is treated as depth ground truth.
★ 14M3Depth. Code for Self-Supervised Depth Estimation in Laparoscopic Image using 3D Geometric Consistency (MICCAI 2022)
★ 29taichi-ngp-renderer. An Instants-NGP renderer that has been implemented using Taichi
★ 373flash-attention. Fast and memory-efficient exact attention
★ 25kkernl. Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackable.
★ 1.6kPlayCover. Community fork of PlayCover
★ 12kDB. A PyTorch implementation of "Real-time Scene Text Detection with Differentiable Binarization".
★ 2.3kEasyOCR. Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
★ 30ksurface_normal_uncertainty. [ICCV 2021 Oral] Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal Estimation
★ 245DataAugmentationForObjectDetection. Data Augmentation For Object Detection
★ 1.2kstable-dreamfusion. Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
★ 8.8kopen_clip. An open source implementation of CLIP.
★ 14kPureCLIPNeRF. Python
★ 173UniDet. Object detection on multiple datasets with an automatically learned unified label space.
★ 517CuPL. Python
★ 203