This is your work, valued
CFG-Zero-star. Official repo for CFG-Zero*
★ 715UAE. Official repo for UAE
★ 208HierarchyFlow. Python
★ 41openpi. Python
★ 13kFastWAM. Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
★ 1.2kstable-worldmodel. A platform for reproducible world model research and evaluation
★ 2.1kminit2i-jax. Official JAX code of MiniT2I.
★ 125Spectral_Forcing. Code of "Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion"
★ 26tuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4kdino_wm. Python
★ 536HY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kSpectrumMatching. official code for the paper 'Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion'
★ 25NEO. NEO Series: Native Vision-Language Models from First Principles
★ 878EAR. Python
★ 42BiFlow. Official Implementation of BiFlow https://arxiv.org/abs/2512.10953
★ 64vjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.
★ 4.4kUniTok. [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
★ 529DynamicVLA. The official implementation of "DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation". (arXiv 2601.22153)
★ 314openclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 384kDCTdiff. [ICML 2025] Official code for the paper 'DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space'
★ 91CE7455. Jupyter Notebook
★ 46UAE. Official repo for UAE [ECCV 2026]
★ 208pico-banana-400k. Python
★ 1.8kawesome-nanobanana-pro. 🚀 An awesome list of curated Nano Banana pro prompts and examples. Your go-to resource for mastering prompt engineering and exploring the creative potential of the Nano banana pro(Nano banana 2) AI image model.
★ 10kJiT. PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
★ 2.5kmf-rae. Python
★ 40Diffusion-Explorer. Interactive visualizations of the geometric intuition behind diffusion models.
★ 1.2kUniFlow. Official Implementation of "UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation"
★ 143NFIG. The source code of NFIG
★ 9RAE. Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
★ 2kInfinity. [CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
★ 1.6kAwesome-Autoregressive-Visual-Generation. This is a repo to track the latest autoregressive visual generation papers.
★ 430TradingAgents. TradingAgents: Multi-Agents LLM Financial Trading Framework
★ 95kAwesome-From-Video-Generation-to-World-Model. A list of works on video generation towards world model
★ 504model-immunization-cond-num. [ICML 2025 Oral] Model Immunization from a Condition Number Perspective
★ 10perception_models. State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
★ 2.3kHarmon. [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
★ 192sdnext. SD.Next: All-in-one WebUI for AI generative image and video creation, captioning and processing
★ 7.2kEasyControl. Implementation of "EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer"(ICCV2025)
★ 1.7kComfyUI-KJNodes. Various custom nodes for ComfyUI
★ 2.9kWan2GP. A fast AI Video Generator for the GPU Poor. Supports Wan 2.1/2.2, LTX-2, Qwen Image, Hunyuan Video, LTX Video and Flux.
★ 6.8kCFG-Zero-star. Official repo for CFG-Zero*
★ 715VideoWorld. [CVPR 2025] VideoWorld is a simple generative model that learns purely from unlabeled videos—much like how babies learn by observing their environment.
★ 793Video-Depth-Anything. [CVPR 2025 Highlight] Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
★ 2kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kSkyReels-V1. SkyReels V1: The first and most advanced open-source human-centric video foundation model
★ 2.7kHRS_benchmark. Jupyter Notebook
★ 60ELLA. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
★ 1.3kCFGpp. Official repository for "CFG++: manifold-constrained classifier free guidance for diffusion models" (ICLR2025)
★ 243RepVideo. The official implementation of ”RepVideo: Rethinking Cross-Layer Representation for Video Generation“
★ 123flow_matching. A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
★ 4.7kFlowAR. “FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching” FlowAR employs a simplest scale design and is compatible with any VAE.
★ 171geneval. GenEval: An object-focused framework for evaluating text-to-image alignment
★ 472large_concept_model. Large Concept Models: Language modeling in a sentence representation space
★ 2.4kLaVi-Bridge. [ECCV 2024] Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
★ 300T2I-CompBench. [Neurips 2023 & TPAMI] T2I-CompBench (++) for Compositional Text-to-image Generation Evaluation
★ 345text2image-benchmark. Benchmark for generative image models
★ 111NOVA. [ICLR 2025] Autoregressive Video Generation without Vector Quantization
★ 657VIMA. Official Algorithm Implementation of ICML'23 Paper "VIMA: General Robot Manipulation with Multimodal Prompts"
★ 854HunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kLIBERO. Benchmarking Knowledge Transfer in Lifelong Robot Learning
★ 2.1kDancing2Music. Python
★ 538PantoMatrix. PantoMatrix: Generating Face and Body Animation from Speech
★ 1.3kVideoClaw. 🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞
★ 1.6kmvp. NeurIPS-2021: Direct Multi-view Multi-person 3D Human Pose Estimation
★ 339echomimic_v2. [CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
★ 4.6kGIMM-VFI. [NeurIPS 2024] Generalizable Implicit Motion Modeling for Video Frame Interpolation
★ 396Emu3. Next-Token Prediction is All You Need
★ 2.4kVchitect-2.0. Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
★ 922awesome-RLHF. A curated list of reinforcement learning with human feedback resources (continually updated)
★ 4.4kVAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kchameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kclvr_jaco_play_dataset. Official release of the CLVR Jaco Play Dataset, Dass et al. 2023
★ 17robotics_transformer. Python
★ 1.7kLumina-mGPT. Official Implementation of "Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining"
★ 646CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kflux. Official inference repo for FLUX.1 models
★ 26ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kSegment-and-Track-Anything. An open-source project dedicated to tracking and segmenting any objects in videos, either automatically or interactively. The primary algorithms utilized include the Segment Anything Model (SAM) for key-frame segmentation and Associating Objects with Transformers (AOT) for efficient tracking and propagation purposes.
★ 3.1kdata-juicer. Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8kIntuitive-Fine-Tuning. [ACL 2025, Main Conference, Oral] Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
★ 30openvla. OpenVLA: An open-source vision-language-action model for robotic manipulation.
★ 6.7kjynew. JinYongLegend-like RPG Game Framework with full Modding support and 10+ hours playable samples of game.
★ 8.9kVBench. [CVPR2024 Highlight] VBench - We Evaluate Video Generation
★ 1.7kVADER. Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as VideoCrafter, OpenSora, ModelScope and StableVideoDiffusion by finetuning them using various reward models such as HPS, PickScore, VideoMAE, VJEPA, YOLO, Aesthetics etc.
★ 316SEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1kVILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kSimpleTuner. A general fine-tuning kit geared toward image/video/audio diffusion models.
★ 2.9kLlamaGen. Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
★ 2kEasyAnimate. 📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion
★ 2.3ktaichi_dem. A minimal DEM simulation demo written in Taichi.
★ 46deep-fluids. Deep Fluids: A Generative Network for Parameterized Fluid Simulations
★ 4diffmpm. Differentiable Material Point Method
★ 48EasySynth. Unreal Engine plugin for easy creation of synthetic image datasets
★ 239CityDreamer. The official implementation of "CityDreamer: Compositional Generative Model of Unbounded 3D Cities". (CVPR 2024)
★ 698DiT-3D. 🔥🔥🔥Official Codebase of "DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation"
★ 321U-ViT. A PyTorch implementation of the paper "All are Worth Words: A ViT Backbone for Diffusion Models".
★ 1.1kvideo2game. Code release of Video2Game
★ 334Neural-Network-Diffusion. We introduce a novel approach for parameter generation, named neural network parameter diffusion (p-diff), which employs a standard latent diffusion model to synthesize a new set of parameters
★ 886hypernerf. Code for "HyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance Fields".
★ 963Dynamic3DGaussians. Python
★ 2.3kLumina-T2X. Lumina-T2X is a unified framework for Text to Any Modality Generation
★ 2.2k