This is your work, valued
Creative
TransPixeler. CVPR2025
★ 926dan_mmediting. Python
★ 17SG-Adapter. Python
★ 3wileewang.github.io. JavaScript
★ 2WorldDirector. Python
★ 68mira. Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
★ 468text-to-cad. A collection of agent skills for CAD, robotics and hardware design
★ 12kStreamMA. Official implementation of "Streaming Communication in Multi-Agent Reasoning"
★ 34LongLive-RAG. Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
★ 102Pointer-CAD. Python
★ 290HY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kStable-Video-Infinity. [ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
★ 2.5klyra. Project Lyra: Open Generative 3D World Models
★ 2.2kCC. TypeScript
★ 278AutoResearchClaw. Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
★ 14kIndexCache. IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
★ 130RealWonder. Real-Time Physical Action-Conditioned Video Generation
★ 218Motion-Forcing. Official implementation of Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics
★ 20CLI-Anything. "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
★ 46kPAP. Panoramic Affordance Prediction (PAP) (ECCV 2026)
★ 46autoresearch. AI agents running research on single-GPU nanochat training automatically
★ 92kworldfm. Python
★ 822Helios. Helios: Real Real-Time Long Video Generation Model
★ 2kCausal-Forcing. [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
★ 888Lean. Lean Algorithmic Trading Engine by QuantConnect (Python, C#)
★ 21kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 384kContext-Forcing. Context Forcing: Consistent Autoregressive Video Generation with Long Context [ICML26]
★ 106lingbot-world. Advancing Open-source World Models
★ 4.3kVideoMemory. VideoMemory's Homepage
★ 4Awesome-Video-World-Models. A Mechanistic View on Video Generation as World Models: State and Dynamics
★ 57YUME. The official code of Yume
★ 679open-p2p. Official Repo for paper: Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing
★ 172Wan-Alpha. [CVPR 2026 Highlight] High-Quality Text-to-Video Generation with Alpha Channel
★ 394TurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kStereoPilot. The official implementation of StereoPilot
★ 116titans-pytorch. Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
★ 2kMemFlow. Official Implementation of "MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives"
★ 216LongVie. Python
★ 334UnityVideo. [CVPR 2026]UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
★ 321DiT-Mem. Learning Plug-and-play Memory for Guiding Video Diffusion Models
★ 26ProphRL. Reinforcing Action Policies by Prophesying
★ 42DualCamCtrl. [ECCV 2026] DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
★ 75VANS. [CVPR 2026] Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
★ 119FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kTiViBench. [CVPR 2026] TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
★ 67WorldMem. [NeurIPS 2025] WorldMem: Long-term Consistent World Simulation with Memory
★ 381open-oasis. Inference script for Oasis 500M
★ 2.1kSTANCE. STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
★ 12UniCalli. Official implementation of UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
★ 215MTI. [ACL 2026] Official implementation of "Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention"
★ 41PhysToolBench. PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
★ 30LongLive. Long Video Gen Infrastructure
★ 2.5kT2I-CoReBench. [ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
★ 52FAR. Code for: "Long-Context Autoregressive Video Modeling with Next-Frame Prediction"
★ 311UniTEX. Official implementation of "UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes"
★ 5Matrix-3D. Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt.
★ 774Waver. Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
★ 950Wan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kLong-RL. Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)
★ 727ComfyMind. ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
★ 123CausVid. (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
★ 1.4kattention-gym. Helpful tools and examples for working with flex-attention
★ 1.2kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kUniTEX-FLUX. Flux training codes (lora) for UniTEX
★ 25FlexPainter. [Arxiv 2025] FlexPainter: Flexible and Multi-View Consistent Texture Generation
★ 44UniTEX. [CVPR 2026]Official implementation of "UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes"
★ 205diffusion-forcing-transformer. [ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
★ 705FractFlow. Python
★ 25codex. Lightweight coding agent that runs in your terminal
★ 102kFramePack. Lets make video diffusion practical!
★ 17kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kFusion360GalleryDataset. Data, tools, and documentation of the Fusion 360 Gallery Dataset
★ 695onshape-cad-parser. Python
★ 93hawk. 🔥[NeurIPS 2024] Official Implementation of Hawk: Learning to Understand Open-World Video Anomalies
★ 224MovieDreamer. [ICLR'25] MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequences
★ 323DeepCAD. code for our ICCV 2021 paper "DeepCAD: A Deep Generative Network for Computer-Aided Design Models"
★ 796MSI-NeRF. [WACV2025] Linking Omni-Depth with View Synthesis through Multi-Sphere Image aided Generalizable Neural Radiance Field
★ 14EINRUL. [ICRA2023] Efficient Implicit Neural Reconstruction Using LiDAR
★ 89RectifiedHR. [CVPR Findings 2026] Official implementation of "RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification"
★ 31Kiss3DGen. [CVPR 2025] Official implementation of "Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation"
★ 296Uni-Renderer. Official implementation of "Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion" [CVPR2025]
★ 56DiT-Extrapolation. Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
★ 820HunyuanVideo-Training. Python
★ 81diffusion-pipe. A pipeline parallel training script for diffusion models.
★ 2kGaussian-Property. Official implementation of GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs.
★ 79Awesome-Physics-aware-Generation. Physical laws underpin all existence, and harnessing them for generative modeling opens boundless possibilities for advancing science and shaping the future!
★ 296DeepSeek-R1.
★ 92kstar-history. The de facto GitHub star history graph.
★ 9.3kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kLeviTor. [CVPR'25 Highlight] Official implementation for paper - LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis
★ 161TransPixeler. CVPR2025
★ 926HunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kCoPRA. [AAAI 2025] CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction
★ 39common_metrics_on_video_quality. You can easily calculate FVD, PSNR, SSIM, LPIPS for evaluating the quality of generated or predicted videos.
★ 582VideoTuna. Let's finetune video generation models!
★ 551Cosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kHarmonizer. High-Resolution Image/Video Harmonization [ECCV 2022]
★ 407finetrainers. Scalable and memory-optimized training of diffusion models
★ 1.4kmochi. The best OSS video generation models, created by Genmo
★ 3.7kAllegro. Allegro is a powerful text-to-video model that generates high-quality videos up to 6 seconds at 15 FPS and 720p resolution from simple text input.
★ 1.1kLucidFusion. Official implementation of “LucidFusion: Reconstructing 3D Gaussians with Arbitrary Unposed Images”
★ 76cogvideox-loras. CogVideoX-LoRAs is a centralized repository for all LoRA models created for CogVideoX, filling the gap for a unified sharing space. With the rising demand for customized video generation, this hub enables users, developers, and researchers to easily access, contribute to, and collaborate on various tailored LoRA models.
★ 81DDBM. Python
★ 279ViewCrafter. [TPAMI 2025] ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
★ 1.6kVideoSys. VideoSys: An easy and efficient system for video generation
★ 2ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kflux. Official inference repo for FLUX.1 models
★ 26kAwesome-Video-Datasets. Video datasets
★ 1.7kdiffusion-forcing. code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"
★ 1.3kDefect_Spectrum. Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich Semantics (ECCV2024)
★ 135SEED-Story. SEED-Story: Multimodal Long Story Generation with Large Language Model
★ 884FilmRemoval. [CVPR 2024] Official Implementation of Learning to Remove Wrinkled Transparent Film with Polarized Prior
★ 43via-video.
★ 25HunyuanDiT. Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
★ 4.3kMGM. Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3.3kmvdream_diffusers. A unified diffusers implementation for MVDream and ImageDream
★ 113Open-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kVAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kdiffusion-motion-transfer. Official Pytorch Implementation for "Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer""
★ 197MotionDirector. [ECCV 2024 Oral] MotionDirector: Motion Customization of Text-to-Video Diffusion Models.
★ 1kDDSM. Denoising Diffusion Step-aware Models (ICLR2024)
★ 62MotionInversion. [SIGGRAPH 2025] Official implementation of 'Motion Inversion For Video Customization'
★ 153Moore-AnimateAnyone. Character Animation (AnimateAnyone, Face Reenactment)
★ 3.5kHumanML3D. HumanML3D: A large and diverse 3d human motion-language dataset.
★ 1.5kmagvit2-pytorch. Implementation of MagViT2 Tokenizer in Pytorch
★ 668Video-Swin-Transformer. This is an official implementation for "Video Swin Transformers".
★ 1.7kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kMotionCtrl. Official Code for MotionCtrl [SIGGRAPH 2024]
★ 1.5kLLaMA2-Accessory. An Open-source Toolkit for LLM Development
★ 2.8kmagic-animate. [CVPR 2024] Official repository for "MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model"
★ 11kTASC. Python
★ 27InstaFlow. :zap: InstaFlow! One-Step Stable Diffusion with Rectified Flow (ICLR 2024)
★ 1.4kLucidDreamer. Official implementation of "LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching"
★ 829gpt4free. The official gpt4free repository | various collection of powerful language models | opus 4.6 gpt 5.3 kimi 2.5 deepseek v3.2 gemini 3
★ 67klangchain-gpt4free. LangChain x gpt4free
★ 188Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kLLaMA-Adapter. [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5.9kReLA. [CVPR 2023 Highlight & IJCV 2026] GRES: Generalized Referring Expression Segmentation
★ 689Link-Context-Learning. Python
★ 101SUR-adapter. ACM MM'23 (oral), SUR-adapter for pre-trained diffusion models can acquire the powerful semantic understanding and reasoning capabilities from large language models to build a high-quality textual semantic representation for text-to-image generation.
★ 120SG-Adapter. Python
★ 3HyperThumbnail. [CVPR 2023] Real-time 6K Image Rescaling with Rate-distortion Optimization. Official implementation.
★ 74Selective-Diffusion-Distillation. Not All Steps are Created Equal: Selective Diffusion Distillation for Image Manipulation (ICCV 2023)
★ 64Emu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kunilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22krefer. Referring Expression Datasets API
★ 573detrex. detrex is a research platform for DETR-based object detection, segmentation, pose estimation and other visual recognition tasks.
★ 2.3kStableLM. StableLM: Stability AI Language Models
★ 16kconsistency_models. Official repo for consistency models.
★ 6.5ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kstanford_alpaca. Code and documentation to train Stanford's Alpaca models, and generate the data.
★ 30kControlNet. Let us control diffusion models!
★ 34kDifFace. DifFace: Blind Face Restoration with Diffused Error Contraction (TPAMI, 2024)
★ 707dream-textures. Stable Diffusion built-in to Blender
★ 8.2kprompt-to-prompt. Jupyter Notebook
★ 3.5kCrossAttentionControl. Unofficial implementation of "Prompt-to-Prompt Image Editing with Cross Attention Control" with Stable Diffusion
★ 1.3keinops. Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
★ 9.6kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kDAI2I. This is the repository for CVPR paper "Domain Adaptive Image-to-image Translation"
★ 9ResizeRight. The correct way to resize images or tensors. For Numpy or Pytorch (differentiable).
★ 566pytorch-fid. Compute FID scores with PyTorch.
★ 3.9klatent-diffusion. High-Resolution Image Synthesis with Latent Diffusion Models
★ 14kguided-diffusion. Python
★ 7.4kdenoising-diffusion-pytorch. Implementation of Denoising Diffusion Probabilistic Models in PyTorch
★ 405pytorch-tutorial. PyTorch Tutorial for Deep Learning Researchers
★ 32kDALLE2-pytorch. Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch
★ 11kopen_clip. An open source implementation of CLIP.
★ 14kAwesome-Diffusion-Models. A collection of resources and papers on Diffusion Models
★ 12kpytorch_diffusion. PyTorch reimplementation of Diffusion Models
★ 585ddim. Denoising Diffusion Implicit Models
★ 1.8kmmgeneration. MMGeneration is a powerful toolkit for generative models, based on PyTorch and MMCV.
★ 2kstylegan3. Official PyTorch implementation of StyleGAN3
★ 6.9klatent-transformer. Official implementation for paper: A Latent Transformer for Disentangled Face Editing in Images and Videos.
★ 149dan_mmediting. Python
★ 17mmagic. OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
★ 7.4klola. LoL (League of Legends) game data analysis / analytics
★ 142DeepLearning-500-questions. 深度学习500问,以问答形式对常用的概率知识、线性代数、机器学习、深度学习、计算机视觉等热点问题进行阐述,以帮助自己及有需要的读者。 全书分为18个章节,50余万字。由于水平有限,书中不妥之处恳请广大读者批评指正。 未完待续............ 如有意合作,联系scutjy2015@163.com 版权所有,违权必究 Tan 2018.06
★ 58kSongRecogn. A Song Recognition Program using Shazam Algorithm
★ 44dota2-win-rate-prediction-v1. This is our first version of vpesports dota2 win rate prediction model. Learning and Testing only now
★ 27