This is your work, valued
Progressive3D. Official implementation of "Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts" [ICLR 2024]
★ 123threestudio-gaussiandreamer. GaussianDreamer extension of threestudio.
★ 49VTB. Official implementation of "A Simple Visual-Textual Baseline for Pedestrian Attribute Recognition" [TCSVT 2022]
★ 38cxh0519.
★ 1threestudio-lrm. Python
★ 1SPEED. [ICLR'26] SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
★ 41R-4B. The official repository of "R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Integration"
★ 141tinyworlds. A minimal implementation of DeepMind's Genie world model
★ 1.3kLLMs-from-scratch. Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
★ 100kllm_interview_note. 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
★ 15kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kMatrix-Game. Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
★ 2.3kOmniDrag. [IJCV 2025] OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
★ 16DanceGRPO. An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
★ 1.6kUniWorld. UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
★ 886Spark-Wan. Spark-Wan: Few-Step Video Diffusion Transformer via Hybrid Adversarial Post-Training
★ 16Q-Insight. Q-Insight Family: Q-Insight, VQ-Insight and RALI (NeurIPS 2025 Spotlight, AAAI 2026 Oral, and ICLR 2026 Oral)
★ 314ReCamMaster. [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
★ 1.8kVACE. [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
★ 3.9kPhantom. Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
★ 1.5kUniAnimate-DiT. UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
★ 854HoloTime. [ACM MM 2025] HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
★ 159genwarp. Python
★ 310SkyReels-V2. SkyReels-V2: Infinite-length Film Generative model
★ 7.3kVideoX-Fun. 📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
★ 2.2kDiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kCogKit. Finetuning and inference tools for the CogView4 and CogVideoX model series.
★ 127Depth-Anywhere. Python
★ 104pkuthss. LaTeX template for dissertations in Peking University
★ 619awesome-resume-for-chinese. :page_facing_up: 适合中文的简历模板收集(LaTeX,HTML/JS and so on)由 @hoochanlon 维护
★ 8.1kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kTrajectoryCrafter. [ICCV 2025, Oral] TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
★ 860Depth-Anything-V2. [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
★ 8.6kVideo-Depth-Anything. [CVPR 2025 Highlight] Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
★ 2kmetric_depth_video_toolbox. Python tools for rendering, viewing and generating metric 3D depth videos. Tools for recovering and exporting camera pose and 3D geometry to popular formats as well as tools for projecting depthvideo in to side by side stereo(sbs) 3D stereo video.
★ 137MoGe. [CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
★ 2.7kImagine360. [Neurips 2025] Imagine360: Immersive 360 Video Generation from Perspective Anchor
★ 187SpaTracker. [CVPR 2024 Highlight] Official PyTorch implementation of SpatialTracker: Tracking Any 2D Pixels in 3D Space
★ 1.1kDiffusionAsShader. [SIGGRAPH 2025] Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
★ 825TrajectoryAttention. [ICLR 2025] Trajectory Attention For Fine-grained Video Motion Control
★ 101360-1M. Python
★ 94LuminaBrush. Illumination Drawing Tools for Text-to-Image Diffusion Models
★ 776DiTCtrl. [CVPR 2025] Official code of "DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation"
★ 323Emu3. Next-Token Prediction is All You Need
★ 2.4kRuyi-Models. Python
★ 518CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13ktime_reversal. Official repo of "Time Reversal Fusion" (ECCV2024)
★ 50HunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kConsisID. [CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
★ 848DiffMorpher. Official Code for DiffMorpher: Unleashing the Capability of Diffusion Models for Image Morphing (CVPR 2024)
★ 506pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kkohya_ss. Python
★ 13kTVG. code for "TVG: A Training-free Transition Video Generation Method with Diffusion Models"
★ 51pidinet. Code for the ICCV 2021 paper "Pixel Difference Networks for Efficient Edge Detection" (Oral).
★ 617ResVR. [ACM MM 24 Best Paper Nomination] ResVR: Joint Rescaling and Viewport Rendering of Omnidirectional Images
★ 24ViewCrafter. [TPAMI 2025] ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
★ 1.6ksvd_keyframe_interpolation. Python
★ 297Kolors. Kolors Team
★ 4.6kSPO. [CVPR 2025] Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
★ 271AnyText. Official implementation code of the paper <AnyText: Multilingual Visual Text Generation And Editing>
★ 4.9kD3C2-Net. [TCSVT 2024] D3C2-Net: Dual-Domain Deep Convolutional Coding Network for Compressive Sensing
★ 45bidiff. [CVPR'24] Text-to-3D Generation with Bidirectional Diffusion using both 2D and 3D priors
★ 168interactive3d. [CVPR'24] Interactive3D: Create What You Want by Interactive 3D Generation
★ 205InstantMesh. InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
★ 4.5kAwesome-Image-Harmonization. A curated list of papers, code and resources pertaining to image harmonization.
★ 536spad. Code for SPAD : Spatially Aware Multiview Diffusers, CVPR 2024
★ 180LN3Diff. [ECCV-2024] LN3Diff creates high-quality 3D object mesh from text within 8 V100-SECONDS.
★ 228GRM. Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation
★ 641CRM. [ECCV 2024] Single Image to 3D Textured Mesh in 10 seconds with Convolutional Reconstruction Model.
★ 691Long-CLIP. [ECCV 2024] official code for "Long-CLIP: Unlocking the Long-Text Capability of CLIP"
★ 901generative-models. Generative Models by Stability AI
★ 27kMVControl. [3DV-2025] Official implementation of "Controllable Text-to-3D Generation via Surface-Aligned Gaussian Splatting"
★ 214PixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kMake-Your-3D. [ECCV 2024] Make-Your-3D: Fast and Consistent Subject-Driven 3D Content Generation
★ 128data-juicer. Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8kComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kOpen-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kFiT. [ICML 2024 Spotlight] FiT: Flexible Vision Transformer for Diffusion Model
★ 435ECDFormer. 【Nature Computational Science 2025🔥】Deep peak property learning for efficient chiral molecules ECD spectra prediction
★ 51DreamPropeller. Python
★ 89HeadStudio. [ECCV 2024] HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting.
★ 216GALA3D. [ICML 2024] GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting
★ 306MIGC. [CVPR 2024 Highlight] MIGC and [TPAMI 2024] MIGC++ (Official Implementation)
★ 612LGM. [ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
★ 2.1kDynamic3DGaussians. Python
★ 2.3kthreestudio-gaussiandreamer. GaussianDreamer extension of threestudio.
★ 49TriplaneGaussian. TriplaneGaussian: A new hybrid representation for single-view 3D reconstruction.
★ 925LODS. Official code for ECCV 2024 paper: Learn to Optimize Denoising Scores A Unified and Improved Diffusion Prior for 3D Generation
★ 72T2I-CompBench. [Neurips 2023 & TPAMI] T2I-CompBench (++) for Compositional Text-to-image Generation Evaluation
★ 345RPG-DiffusionMaster. [ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
★ 1.8k360DVD. [CVPR2024] 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
★ 181GPTEval3D. [ CVPR 2024 ] Implementation for "GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation"
★ 2883D-Gaussian-Splatting-Papers. 3D高斯论文,持续更新,欢迎交流讨论。
★ 3.1ksupersplat. 3D Gaussian Splat Editor
★ 9.8kthreestudio-3dgs. 3D Gaussian Splatting extension of threestudio.
★ 201richdreamer. [CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC
★ 478OpenLRM. An open-source impl. of Large Reconstruction Models
★ 1.2kMachine-Mindset. An MBTI Exploration of Large Language Models
★ 538repaint123. Official implementation of Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting (ECCV 2024)
★ 276GaussianDreamer. [CVPR 2024] GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models
★ 828HumanGaussian. [CVPR 2024 Highlight] Code for "HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting"
★ 492EditGuard. [CVPR 2024] EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright Protection
★ 264VTB. Official implementation of "A Simple Visual-Textual Baseline for Pedestrian Attribute Recognition" [TCSVT 2022]
★ 38GraphDreamer. [CVPR'24] GraphDreamer: a novel framework of generating compositional 3D scenes from scene graphs.
★ 201