This is your work, valued
verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kvggt-omega. [CVPR 2026 Oral] VGGT Omega
★ 3.8klingbot-map. A feed-forward 3D foundation model for reconstructing scenes from streaming data
★ 16klyra. Project Lyra: Open Generative 3D World Models
★ 2.2kHY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kVideoX-Fun. 📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
★ 2.2kdreamzero. Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
★ 2.5kclaw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195kReBalance. [ICLR 2026] Efficient Reasoning with Balanced Thinking
★ 333VIGA. VIGA: Vision-as-Inverse-Graphics Agent
★ 1.3kWan-Move. [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
★ 649RoboVerse. RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
★ 1.8knvitop. An interactive NVIDIA-GPU process viewer and beyond, the one-stop solution for GPU process management.
★ 7.1kopen-pi-zero. Re-implementation of pi0 vision-language-action (VLA) model from Physical Intelligence
★ 1.5ksam-3d-objects. SAM 3D Objects
★ 7.2kPartField. [ICCV 2025] PartField: Learning 3D Feature Fields for Part Segmentation and Beyond
★ 443Open3D. Open3D: A Modern Library for 3D Data Processing
★ 14kDepth-Anything-3. Depth Anything 3
★ 6klerobot. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
★ 26kHunyuanWorld-Mirror. [ICML 2026] WorldMirror: Fast and Universal 3D reconstruction model for versatile tasks
★ 1.2kFlashWorld. Code for "FlashWorld: High-quality 3D Scene Generation within Seconds" (ICLR 2026 Oral)
★ 834LongLive. Long Video Gen Infrastructure
★ 2.5kwall-x. Building General-Purpose Robots Based on Embodied Foundation Model
★ 1.2kReconViaGen. (ICLR2026) ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
★ 625OpenUSD. Universal Scene Description
★ 7.4kmujoco. Multi-Joint dynamics with Contact. A general purpose physics simulator.
★ 14kIsaac-GR00T. NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
★ 7.7kSTream3R. Dynamic 3D Foundation Model using Causal Transformer. [ICLR 2026]
★ 393octo. Octo is a transformer-based robot policy trained on a diverse mix of 800k robot trajectories.
★ 1.7kopenpi. Python
★ 13kmitsuba3. Mitsuba 3: A Retargetable Forward and Inverse Renderer
★ 2.9kBagel. Open-source unified multimodal model
★ 6.1kRLBench. A large-scale benchmark and learning environment.
★ 1.8kmobile-aloha. Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
★ 4.5kact-plus-plus. Imitation learning algorithms with Co-training for Mobile ALOHA: ACT, Diffusion Policy, VINN
★ 3.7kdiffusion_policy. [RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
★ 4.4kHunyuanWorld-1.0. Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
★ 2.9kprope. Cameras as Relative Positional Encoding
★ 744Long-RL. Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)
★ 727Shape-for-Motion. [Siggraph Asia 2025] Official code release of our paper "Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy"
★ 59FramePack. Lets make video diffusion practical!
★ 17kDNF-Avatar. [ICCV 2025 Findings Oral] DNF-Avatar: Distilling Neural Fields for Real-time Animatable Avatar Relighting
★ 39InstantMesh. InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
★ 4.5kMV-Adapter. [ICCV 2025] Official impl. of "MV-Adapter: Multi-view Consistent Image Generation Made Easy"
★ 1.3kGeomDist. Python
★ 273FlashVDM. Unleashing Vecset Diffusion Model for Fast Shape Generation / within 1 Second (ICCV'25 Highlight)
★ 333Hunyuan3D-2. High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
★ 14kstable-virtual-camera. Stable Virtual Camera: Generative View Synthesis with Diffusion Models
★ 1.6kvggsfm. VGGSfM: Visual Geometry Grounded Deep Structure From Motion
★ 1.4kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kDyT. Code release for DynamicTanh (DyT)
★ 1kDeepSeek-R1.
★ 92kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kEvaluation-Agent. [ACL2025 Oral & Award] Evaluate Image/Video Generation like Humans - Fast, Explainable, Flexible
★ 128Neural-LightRig. [CVPR2025] Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion
★ 192TRELLIS. Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
★ 13kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kSAR3D. Official repository for "SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE"
★ 200MaterialAnything. [CVPR 2025 Highlight] Material Anything: Generating Materials for Any 3D Object via Diffusion
★ 360zero123plus. Code repository for Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.
★ 2.1kAllegro. Allegro is a powerful text-to-video model that generates high-quality videos up to 6 seconds at 15 FPS and 720p resolution from simple text input.
★ 1.1kblender-plugin. Python
★ 1.3kMarigold. [CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
★ 3.2kml-depth-pro. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.
★ 5.6kPhidias-Diffusion. [ICLR 2025] Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
★ 295Step-DPO. Implementation for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"
★ 398DynamiCrafter. [ECCV 2024, Oral] DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
★ 3kX-Ray. The official source code for "X-Ray: A Sequential 3D Representation for Generation".
★ 116ThemeStation. [SIGGRAPH 2024] ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars
★ 219differentiable-sdf-rendering. Source code for "Differentiable Signed Distance Function Rendering" (Siggraph 2022)
★ 906cuda-gridsample-grad2. Cuda implementation for gridsample with second derivative support
★ 90ChamferDistancePytorch. Chamfer Distance in Pytorch with f-score
★ 375LWM. Large World Model -- Modeling Text and Video with Millions Context
★ 7.4kmamba. Mamba SSM architecture
★ 19krembg. Rembg is a tool to remove images background
★ 24kOpenLRM. An open-source impl. of Large Reconstruction Models
★ 1.2kopen_clip. An open source implementation of CLIP.
★ 1flash-attention. Fast and memory-efficient exact attention
★ 25ktrl. Train transformer language models with reinforcement learning.
★ 19kconsistency_models. Official repo for consistency models.
★ 6.5kARF-svox2. Artistic Radiance Fields
★ 579consistencydecoder. Consistency Distilled Diff VAE
★ 2.2kIF. Python
★ 7.8kdreamgaussian. [ICLR 2024 Oral] Generative Gaussian Splatting for Efficient 3D Content Creation
★ 4.3kobjaverse-xl. 🪐 Objaverse-XL is a Universe of 10M+ 3D Objects. Contains API Scripts for Downloading and Processing!
★ 1.3kdiff-gaussian-rasterization. Cuda
★ 488diff-gaussian-rasterization. Cuda
★ 1.5kLongLoRA. Code and documents of LongLoRA and LongAlpaca (ICLR 2024 Oral)
★ 2.7kDiffTF. Official PyTorch implementation of DiffTF (Accepted by ICLR2024)
★ 200infinigen. Infinite Photorealistic Worlds using Procedural Generation
★ 7.2kCityDreamer. The official implementation of "CityDreamer: Compositional Generative Model of Unbounded 3D Cities". (CVPR 2024)
★ 698OpenAgents. [COLM 2024] OpenAgents: An Open Platform for Language Agents in the Wild
★ 4.9ktilted. Canonical Factors for Hybrid Neural Fields @ ICCV 2023
★ 107AgentSims. AgentSims is an easy-to-use infrastructure for researchers from all disciplines to test the specific capacities they are interested in.
★ 959generative_agents. Generative Agents: Interactive Simulacra of Human Behavior
★ 22kAwesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
★ 5.7kGenBench. Benchmarking and Analyzing Generative Data for Visual Recognition
★ 26llama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
★ 19kCorrelational-Image-Modeling. Python
★ 31AnyDoor. Official implementations for paper: Anydoor: zero-shot object-level image customization
★ 4.2kDrag3D. DragGAN meets GET3D for interactive mesh generation and editing.
★ 466parti-pytorch. Implementation of Parti, Google's pure attention-based text-to-image neural network, in Pytorch
★ 537IST-Net. (ICCV2023) IST-Net: Prior-free Category-level Pose Estimation with Implicit Space Transformation
★ 120LLaVAR. Code/Data for the paper: "LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding"
★ 268Voyager. An Open-Ended Embodied Agent with Large Language Models
★ 7.1kSSDNeRF. [ICCV 2023] Single-Stage Diffusion NeRF
★ 447Magic123. [ICLR'24] Official PyTorch Implementation of Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors
★ 1.6koft. Official implementation of "Controlling Text-to-Image Diffusion by Orthogonal Finetuning".
★ 300DragGAN. Official Code for DragGAN (SIGGRAPH 2023)
★ 36klens. This is the official repository for the LENS (Large Language Models Enhanced to See) system.
★ 354ContextDET. Contextual Object Detection with Multimodal Large Language Models
★ 261unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kMineCLIP. Foundation Model for MineDojo
★ 301MineDojo. Building Open-Ended Embodied Agents with Internet-Scale Knowledge
★ 2.2ktree-of-thought-prompting. Using Tree-of-Thought Prompting to boost ChatGPT's reasoning
★ 821Cones-V2. [ICML 2023 Oral, NeurIPS 2023] Official implementations for paper: Customizable Image Synthesis with Multiple Subjects
★ 446Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kgenerative-models. Generative Models by Stability AI
★ 27kMobileSAM. This is the official code for MobileSAM project that makes SAM lightweight for mobile applications and beyond!
★ 5.8kAcademic-project-page-template. A project page template for academic papers. Demo at https://eliahuhorwitz.github.io/Academic-project-page-template/
★ 5.1kVersatile-Diffusion. Versatile Diffusion: Text, Images and Variations All in One Diffusion Model, arXiv 2022 / ICCV 2023
★ 1.3kmath_problems-step-by-step_solutions. Here we provide and collect many functions to generate math problem and step by step solutions for LLM training
★ 19LMFlow. An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
★ 8.5kLOMO. LOMO: LOw-Memory Optimization
★ 993dreamsim. DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data (NeurIPS 2023 Spotlight) / / / / When Does Perceptual Alignment Benefit Vision Representations? (NeurIPS 2024)
★ 616Mind2Web. [NeurIPS'23 Spotlight] "Mind2Web: Towards a Generalist Agent for the Web" -- the first LLM-based web agent and benchmark for generalist web agents
★ 1kPandaGPT. [TLLM'23] PandaGPT: One Model To Instruction-Follow Them All
★ 864lm-evaluation-harness. A framework for few-shot evaluation of language models.
★ 13kprm800k. 800,000 step-level correctness labels on LLM solutions to MATH problems
★ 2.2kmegfile. Megvii FILE Library - Working with Files in Python same as the standard library
★ 176ControlVideo. [ICLR 2024] Official pytorch implementation of "ControlVideo: Training-free Controllable Text-to-Video Generation"
★ 864Gen-L-Video. The official implementation for "Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising".
★ 308RaBit. This repository includes the related code of RaBit.
★ 95RenderMe-360. RenderMe-360: Large Digital Asset Library and Benchmark Towards High-fidelity Head Avatars
★ 254bloomchat. This repo contains the data preparation, tokenization, training and inference code for BLOOMChat. BLOOMChat is a 176 billion parameter multilingual chat model based on BLOOM.
★ 583CodeT5. Home of CodeT5: Open Code LLMs for Code Understanding and Generation
★ 3.1kLLMScore. LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation
★ 135tree-of-thought-llm. [NeurIPS 2023] Tree of Thoughts: Deliberate Problem Solving with Large Language Models
★ 6kdream-textures. Stable Diffusion built-in to Blender
★ 8.2kVisionLLM. VisionLLM Series
★ 1.2kdynibar. Implementation of DynIBaR Neural Dynamic Image-Based Rendering (CVPR 2023)
★ 817InternGPT. InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
★ 3.2kthreestudio. A unified framework for 3D content generation.
★ 7klgssl. [CVPR 2023] Learning Visual Representations via Language-Guided Sampling
★ 150DetGPT. Jupyter Notebook
★ 786ImageReward. [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7kImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kstarcoder. Home of StarCoder: fine-tuning & inference!
★ 7.5kshap-e. Generate 3D objects conditioned on text or images
★ 12kG.pt. Official PyTorch Implementation of "Learning to Learn with Generative Models of Neural Network Checkpoints"
★ 347mPLUG-Owl. mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
★ 2.5kGPT4Tools. GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.
★ 771Segment-Everything-Everywhere-All-At-Once. [NeurIPS 2023] Official implementation of the paper "Segment Everything Everywhere All at Once"
★ 4.8kskypilot. The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
★ 10klamini. The Official Python Client for Lamini's API
★ 2.5kdatacomp. DataComp: In search of the next generation of multimodal datasets
★ 787FastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kMultimodal-GPT. Multimodal-GPT
★ 1.5kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kgpt4free. The official gpt4free repository | various collection of powerful language models | opus 4.6 gpt 5.3 kimi 2.5 deepseek v3.2 gemini 3
★ 67kChatCaptioner. Official Repository of ChatCaptioner
★ 468Segment-Anything-NeRF. Segment-anything interactively in NeRF.
★ 288sparsegpt. Code for the ICML 2023 paper "SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot".
★ 892zipnerf-pytorch. Unofficial implementation of ZipNeRF
★ 853gradient-checkpointing. Make huge neural nets fit in memory
★ 2.8kMOSS. An open-source tool-augmented conversational language model from Fudan University
★ 12kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kllama.cpp. LLM inference in C/C++
★ 122kStableLM. StableLM: Stability AI Language Models
★ 16krich-text-to-image. Rich-Text-to-Image Generation
★ 800Inpaint-Anything. Inpaint anything using Segment Anything and inpainting models.
★ 7.7kODISE. Official PyTorch implementation of ODISE: Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models [CVPR 2023 Highlight]
★ 9453D-Box-Segment-Anything. We extend Segment Anything to 3D perception by combining it with VoxelNeXt.
★ 565AutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kMOOD. Official PyTorch implementation of MOOD series: (1) MOODv1: Rethinking Out-of-distributionDetection: Masked Image Modeling Is All You Need. (2) MOODv2: Masked Image Modeling for Out-of-Distribution Detection.
★ 137grounded-segment-any-parts. Grounded Segment Anything: From Objects to Parts
★ 416camel. 🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org
★ 18kEditAnything. Edit anything in images powered by segment-anything, ControlNet, StableDiffusion, etc. (ACM MM)
★ 3.4kVideoCrafter. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
★ 5.1kdolly. Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform
★ 11kChinese-LLaMA-Alpaca. 中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
★ 19kviper. Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning"
★ 1.7kgenvs.
★ 640segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kalpaca-lora. Instruct-tune LLaMA on consumer hardware
★ 19kinstruct-pix2pix. Python
★ 6.9kinstant-ngp. Instant neural graphics primitives: lightning fast NeRF and more
★ 18kAttend-and-Excite. Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
★ 770NeuralLift-360. [CVPR 2023, Highlight] "NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° Views", Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, Zhangyang Wang
★ 318zero123. Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)
★ 3.1kSet-the-Scene. Set-the-Scene
★ 83text2room. Text2Room generates textured 3D meshes from a given text prompt using 2D text-to-image models (ICCV2023).
★ 1.1kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kLLaMA-Adapter. [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5.9kgigagan-pytorch. Implementation of GigaGAN, new SOTA GAN out of Adobe. Culmination of nearly a decade of research into GANs
★ 1.9kMake-It-3D. [ICCV 2023] Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior
★ 1.9kRef-NPR. [CVPR 2023] Ref-NPR: Reference-Based Non-PhotoRealistic Radiance Fields
★ 126