This is your work, valued
Awesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kSimDA. [CVPR 2024] SimDA: Simple Diffusion Adapter for Efficient Video Generation
★ 128SVFormer. Python
★ 90VIDiff.
★ 39YOLO-TrafficDetection. Python
★ 35Course-scheduling-system. Java
★ 24AID.
★ 19SSP3D. [ECCV 2022, Semi-Supervised Single-View 3D Reconstruction via Prototype Shape Priors]
★ 15WMSManager. Java
★ 8Compress-software. Java
★ 7Navigation-system. Java
★ 4Calculator. Java
★ 3ChenHsing.github.io. JavaScript
★ 1ChildrenProtectionApp.
★ 1software-test. softwareTest
★ 1MPCN. Python
★ 1FlashMotion. [CVPR 2026] FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
★ 65OpenRT. Open-source red teaming framework for MLLMs with 42+ attack methods
★ 258FlashPortrait. [CVPR2026]We present FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6$\times$ acceleration in inference speed.
★ 480Wan-Move. [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
★ 649ProLongVid. [EMNLP 2025 Oral] ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning
★ 5StableAvatar. We present StableAvatar, the first end-to-end video diffusion transformer, which synthesizes infinite-length high-quality audio-driven avatar videos without any post-processing, conditioned on a reference image and audio.
★ 1.3kWan2.2-TI2V-5B-Turbo. 4-steps distilled version of Wan2.2-TI2V-5B
★ 163Lumos. [ICLR 2026] Lumos Project: Frontier video unified model research by Alibaba DAMO Academy.
★ 161video-generation-survey. A reading list of video generation
★ 723ChenHsing.github.io. JavaScript
★ 1ReasonGen-R1. Official respository for ReasonGen-R1
★ 75ComfyUI_StableAnimator. ComfyUI nodes for StableAnimator
★ 17VideoX-Fun. 📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
★ 2.2kMAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.7kFramePack. Lets make video diffusion practical!
★ 17kSimpleAR. Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"
★ 431Awesome-Physics-Cognition-based-Video-Generation. A comprehensive list of papers investigating physical cognition in video generation, including papers, codes, and related websites.
★ 322CreatiLayout. [ICCV 2025] CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
★ 135MagicMotion. [ICCV 2025] MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
★ 184Wan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kRealisDance. The official implementation of RealisDance
★ 613StableAnimator. [CVPR2025] We present StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a sequence of poses.
★ 1.4kccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2kawesome-diffusion-v2v. Awesome diffusion Video-to-Video (V2V). A collection of paper on diffusion model-based video editing, aka. video-to-video (V2V) translation. And a video editing benchmark code.
★ 289HunyuanDiT. Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
★ 4.3kCtrl-Adapter. Official implementation of Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model (ICLR 2025 Oral)
★ 470fiftyone. Refine high-quality datasets and visual AI models
★ 11kMotionFollower. [ICCV2025] MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
★ 245SmartEdit. Official code of SmartEdit [CVPR-2024 Highlight]
★ 374generative-models. Generative Models by Stability AI
★ 27kNExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6kqformer. Implementation of Qformer from BLIP2 in Zeta Lego blocks.
★ 51PhotoMaker. PhotoMaker [CVPR 2024]
★ 10kPixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kSEINE. [ICLR 2024] SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
★ 967adaptive-diffusion. [CVPR'2024] Adaptive Teacher-Student Collaboration for Text-Conditional Diffusion Models
★ 33HD-VG-130M. The HD-VG-130M Dataset
★ 126MotionGPT. [NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
★ 1.9kSVD_Xtend. Stable Video Diffusion Training Code and Extensions.
★ 731latent-consistency-model. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
★ 4.6ksdxl-turbo-cog. SDXL-Turbo is a real-time synthesis model, derived from SDXL 1.0, and utilizes a training method called Adversarial Diffusion Distillation (ADD). It achieves high image quality within one to four sampling steps
★ 7AutoStory. [IJCV'24] AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort
★ 149freecontrol. Official implementation of CVPR 2024 paper: "FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition"
★ 480VideoBLIP. Supercharged BLIP-2 that can handle videos
★ 123ART.V.
★ 43VIDiff.
★ 39MotionEditor. [CVPR2024] MotionEditor is the first diffusion-based model capable of video motion editing.
★ 188consistencydecoder. Consistency Distilled Diff VAE
★ 2.2kNRVQA. no reference image/video quaity assessment(BRISQUE/NIQE/PIQE/DIQA/deepBIQ/VSFA
★ 318Open-VCLIP. Python
★ 119loveu-tgve-2023. Official GitHub repository for the Text-Guided Video Editing (TGVE) competition of LOVEU Workshop @ CVPR'23.
★ 77InstructDiffusion. PyTorch implementation of InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions.
★ 445Ground-A-Video. Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models (ICLR 2024)
★ 140AnimateDiff. Official implementation of AnimateDiff.
★ 12kPVDM. [CVPR'23] Video Probabilistic Diffusion Models in Projected Latent Space
★ 322diffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kProPainter. [ICCV 2023] ProPainter: Improving Propagation and Transformer for Video Inpainting
★ 6.8kMake-A-Protagonist. Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
★ 322Veri3d. Python
★ 38CrossLoc. [CVPR'22] CrossLoc localization: a cross-modal visual representation learning method for absolute localization
★ 113SwinGNN. SwinGNN: diffusion model for graph generation
★ 87PSG-biased-annotation. Panoptic Scene Graph Biased Annotation
★ 35WaterScenes. Official repository for WaterScenes dataset
★ 186Awesome-Radar-Camera-Fusion. Radar Camera Fusion in Autonomous Driving
★ 424ICL_SSL. [IEEE TNNLS 2022] An official source code for paper Interpolation-based Contrastive Learning for Few-Label Semi-Supervised Learning.
★ 40Awesome-Skeleton-based-Action-Recognition. A curated paper list of awesome skeleton-based action recognition.
★ 723Robot_Navigation_RL. Two paper About robot navigation in dynamic environment
★ 46DealMVC. [ACM MM 2023] An official source code for paper "DealMVC: Dual Contrastive Calibration for Multi-view Clustering"
★ 65CONVERT. [ACM MM 2023] An official source code for paper "CONVERT: Contrastive Graph Clustering with Reliable Augmentation".
★ 50CCGC. [AAAI 2023] An official source code for paper Cluster-guided Contrastive Graph Clustering Network.
★ 97ICCV23-IDPT. The code for the paper "Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud Models" (ICCV'23).
★ 111BadHash. The official implementation of BadHash
★ 58AdvCLIP. The implementation of our ACM MM 2023 paper "AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive Learning"
★ 100SigmaCCS. [Communications Chemistry 2023] Highly accurate and large-scale collision cross section prediction with graph neural network for compound identification
★ 61NeMF. [ICCV 2023] NeMF: Inverse Volume Rendering with Neural Microflake Field
★ 683D_Semantic_Subspace_Traverser. C++
★ 64ScaleVLN. [ICCV 2023 Oral]: Scaling Data Generation in Vision-and-Language Navigation
★ 225RFN. [TCSVT2022] Industria Scene Text Detection
★ 84SkeletonMAE. SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training
★ 86CMCIR. [IEEE T-PAMI 2023] Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering
★ 78DTP. [ICCV 2023] Official implement of <Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement>
★ 71FIMVC-VIA. [IEEE TNNLS 2022] Code of "Fast Incomplete Multi-view Clustering with View-independent Anchors"
★ 69EOMSC-CA. [AAAI 2022] Code of "Efficient One-pass Multi-view Subspace Clustering with Consensus Anchors"
★ 83SiT. Official implementation of "Self-slimmed Vision Transformer" (ECCV2022)
★ 72LoGoNet. [CVPR2023] LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross-Modal Fusion
★ 284Awesome-Nighttime-Enhancement. Collection of recent nighttime enhancement works, including papers, codes, datasets, and metrics.
★ 122FogRemoval. [ACCV22] Structure Representation Network and Uncertainty Feedback Learning for Dense Non-Uniform Fog Removal, https://arxiv.org/abs/2210.03061
★ 164S-Aware-network. [AAAI23] Estimating Reflectance Layer from A Single Image: Integrating Reflectance Guidance and Shadow/Specular Aware Learning, https://arxiv.org/abs/2211.14751
★ 89nighttime_dehaze. [ACMMM2023] "Enhancing Visibility in Nighttime Haze Images Using Guided APSF and Gradient Adaptive Convolution", https://arxiv.org/abs/2308.01738
★ 184night-enhancement. [ECCV2022] "Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression", https://arxiv.org/abs/2207.10564
★ 465DC-ShadowNet-Hard-and-Soft-Shadow-Removal. [ICCV2021]"DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided Network", https://arxiv.org/abs/2207.10434
★ 255NVDS. ICCV 2023 "Neural Video Depth Stabilizer" (NVDS) & TPAMI 2024 "NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation" (NVDS+)
★ 528UAL. The code for ECCV2022 paper: Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval
★ 55ISR_ICCV2023_Oral. The code for ICCV2023 Oral paper: Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-identification
★ 82SimDA. [CVPR 2024] SimDA: Simple Diffusion Adapter for Efficient Video Generation
★ 128VTSNN. Public code for VTSNN: A Virtual Temporal Spiking Neural Network (Fron. Neur.)
★ 53ICLR_TINY_SNN. Offical implementation of "When Spiking Neural Networks Meet Temporal Attention Image Decoding And AdaptiveE Spiking Neuron" (ICLR2023)
★ 67GAC. Offical implementation of "Gated Attention Coding for Training High-performance and Efficient Spiking Neural Networks" (AAAI2024)
★ 120ASDA. Python
★ 86MG-SCR. [IJCAI-2021] Official Codes for "Multi-Level Graph Encoding with Structural-Collaborative Relation Learning for Skeleton-Based Person Re-Identification"
★ 57SGE-LA. [IJCAI-2020] Official Codes for "Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification"
★ 69Locality-Awareness-SGE. [TPAMI-2022] Official Codes for "A Self-Supervised Gait Encoding Approach with Locality-Awareness for 3D Skeleton Based Person Re-Identification"
★ 78TranSG. [CVPR-2023] Official Codes for "TranSG: Transformer-Based Skeleton Graph Prototype Contrastive Learning with Structure-Trajectory Prompted Reconstruction for Person Re-Identification"
★ 98SM-SGE. [ACMMM-2021] Official Codes for "SM-SGE: A Self-Supervised Multi-Scale Skeleton Graph Encoding Framework for Person Re-Identification"
★ 57SimMC. [IJCAI-2022] Official Codes for "SimMC: Simple Masked Contrastive Learning of Skeleton Representations for Unsupervised Person Re-Identification"
★ 59Hi-MPC. [IJCV-2024] Official Codes for "Hierarchical Skeleton Meta-Prototype Contrastive Learning with Hard Skeleton Mining for Unsupervised Person Re-Identification"
★ 67NvEM. [ACM MM 2021 Oral] Official repo of "Neighbor-view Enhanced Model for Vision and Language Navigation"
★ 78DeRy. [NeurIPS2022] Deep Model Reassembly
★ 256ConsistentTeacher. [CVPR2023 Highlight] Consistent-Teacher: Towards Reducing Inconsistent Pseudo-targets in Semi-supervised Object Detection
★ 315Openvino_Yolov5_async. yolov5的openvino模型,带异步推理
★ 57AdvEncoder. The implementation of our ICCV 2023 paper "Downstream-agnostic Adversarial Examples"
★ 69RelightHuman.
★ 55C3BN. accepted by ICME2023 oral(CCF B)
★ 64AAAI2023-PVD. [IJCV] Official Implementation of PVD and PVDAL: http://sk-fun.fun/PVD-AL/
★ 180DiffusionTrack. [AAAI 2024] DiffusionTrack: Diffusion Model For Multi-Object Tracking. DiffusionTrack is the first work to employ the diffusion model for multi-object tracking by formulating it as a generative noise-to-tracking diffusion process.
★ 206LVOS.
★ 92PBRNet. The code is for PBRnet for action detection
★ 76CoDT. The implementaion of CoDT on the task of NTU-60+->PKUMMD
★ 76Variational-Degeneration-to-Structural-Refinement.
★ 54UniPC. [NeurIPS 2023] UniPC: A Unified Predictor-Corrector Framework for Fast Sampling of Diffusion Models
★ 357Math-Reasoning-With-PLMs. EMNLP22: Multi-View Reasoning: Consistent Contrastive Learning for Math Word Problem
★ 218UDR-S2Former_deraining. [ICCV'23] Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain Streaks
★ 149SnowFormer. This is the official implementation code repository of SnowFormer: Scale-aware Transformer via Context Interaction for Single Image Desnowing
★ 92DiffMIC. [MICCAI 2023] DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
★ 173ViWS-Net. [ICCV 2023] Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial Backpropagation
★ 66SIGA. [CVPR2023] Self-supervised Implicit Glyph Attention for Text Recognition
★ 110RecommendFlow. A Industrialization recommendation system which can easy to build pipeline between local/hdfs train data and tensorflow2.x model.
★ 66FaissSearcher. a common faiss searcher
★ 169VideoDesnowing. [ICCV 2023] Snow Removal in Video: A New Dataset and A Novel Method
★ 71MaskedDenoising. [CVPR 2023] Masked Image Training for Generalizable Deep Image Denoising https://arxiv.org/abs/2303.13132
★ 285Co-DETR. [ICCV 2023] DETRs with Collaborative Hybrid Assignments Training
★ 1.4kHoP. [ICCV 2023] Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction
★ 196MutexMatch4SSL. "MutexMatch: Semi-Supervised Learning with Mutex-Based Consistency Regularization" by Yue Duan (TNNLS)
★ 72VQ-Font. [ICCV 2023] Few shot font generation via transferring similarity guided global and quantization local styles
★ 154VPD. [ICCV 2023] VPD is a framework that leverages the high-level and low-level knowledge of a pre-trained text-to-image diffusion model to downstream visual perception tasks.
★ 540Exp-BLIP. Official implementation of BMVC2023 Oral paper: 《Describe Your Facial Expressions by Linking Image Encoders and Large Language Models》
★ 65Data-Copilot. Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
★ 1.5kParameterized-Cost-Volume-for-Stereo-Matching. PCVNet (ICCV2023)
★ 89XNet. [ICCV2023] XNet: Wavelet-Based Low and High Frequency Merging Networks for Semi- and Supervised Semantic Segmentation of Biomedical Images
★ 215SPC. [BMVC2023] Spatial and Planar Consistency for Semi-Supervised Volumetric Medical Image Segmentation
★ 77DS2DP. The official implement of DS2DP [TGRS 2022]
★ 63ProST. Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval --ICCV2023 Oral
★ 92Partial2Complete. [ICCV 2023] P2C: Self-Supervised Point Cloud Completion from Single Partial Clouds
★ 194BACL. Balanced Classification: A Unified Framework for Long-Tailed Object Detection (TMM 2023)
★ 101ETPNav. [TPAMI 2024] Official repo of "ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments"
★ 478DT-ST. Towards Better Stability and Adaptability: Improve Online Self-Training for Model Adaptation in Semantic Segmentation(CVPR-2023)
★ 81CCD. [ICCV2023] Self-supervised Character-to-Character Distillation for Text Recognition
★ 153IAGNet. The Pytorch implementation of Grounding 3D Object Affordance from 2D Interactios in Images.
★ 138PRG4SSL-MNAR. "Towards Semi-supervised Learning with Non-random Missing Labels" by Yue Duan (ICCV 2023)
★ 76CASE. Accepted by ICCV2023, Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach
★ 106DDS2M. The official implementation of DDS2M [ICCV 2023].
★ 126PCS. Boosting Single Image Super-Resolution via Partial Channel Shifting.
★ 77ViECap. Transferable Decoding with Visual Entities for Zero-Shot Image Captioning, ICCV 2023
★ 167CausalVLR. CausalVLR: A Toolbox and Benchmark for Vision-Language Causal Reasoning (多模态因果推理开源框架)
★ 1.1kTF-ICON. [ICCV 2023] "TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition" (Official Implementation)
★ 814VLN-BEVBert. [ICCV 2023] Official repo of "BEVBert: Multimodal Map Pre-training for Language-guided Navigation"
★ 260PanoSwinTransformerObjectDetection. Python
★ 18Awesome-Open-Vocabulary. (TPAMI 2024) A Survey on Open Vocabulary Learning
★ 999Awesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kmmagic. OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
★ 7.4kawesome-AI-system. paper and its code for AI System
★ 377svdiff-pytorch. Implementation of "SVDiff: Compact Parameter Space for Diffusion Fine-Tuning"
★ 386VideoCrafter. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
★ 5.1ksemi-vit. PyTorch implementation of Semi-supervised Vision Transformers
★ 63adamae. [CVPR'23] AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders
★ 84mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4kSVFormer. Python
★ 90SSP3D. [ECCV 2022, Semi-Supervised Single-View 3D Reconstruction via Prototype Shape Priors]
★ 15TALLFormer. Python
★ 53mmaction2. OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark
★ 5.1kVideoMAE. [NeurIPS 2022 Spotlight] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
★ 1.8kVideoTransformer-pytorch. PyTorch implementation of a collections of scalable Video Transformer Benchmarks.
★ 306jonbarron.github.io. HTML
★ 3.6kUP-TAL. [CVPR2022] Unsupervised Pre-training for Temporal Action Localization Tasks (UP-TAL)
★ 29unbiased-teacher. PyTorch code for ICLR 2021 paper Unbiased Teacher for Semi-Supervised Object Detection
★ 436FSL-Mate. FSL-Mate: A collection of resources for few-shot learning (FSL).
★ 1.8kEIL. Python
★ 9PanoBasic. Matlab Toolbox for Panorama Image Processing
★ 93A-Single-View-3D-Object-Point-Cloud-Reconstruction. 3D-ReConstnet: A Single-View 3D-Object Point Cloud Reconstruction Network for more details,please see this link:and please cite our paper: https://ieeexplore.ieee.org/document/9086481?source=authoralert B. Li, Y. Zhang, B. Zhao and H. Shao, "3D-ReConstnet: A Single-View 3D-Object Point Cloud Reconstruction Network," in IEEE Access, vol. 8, pp. 83782-83790, 2020, doi: 10.1109/ACCESS.2020.2992554.
★ 2A-Single-View-3D-Object-Point-Cloud-Reconstruction. 3D-ReConstnet: A Single-View 3D-Object Point Cloud Reconstruction Network for more details,please see this link:and please cite our paper: https://ieeexplore.ieee.org/document/9086481?source=authoralert B. Li, Y. Zhang, B. Zhao and H. Shao, "3D-ReConstnet: A Single-View 3D-Object Point Cloud Reconstruction Network," in IEEE Access, vol. 8, pp. 83782-83790, 2020, doi: 10.1109/ACCESS.2020.2992554.
★ 19pytorch_deephash. Pytorch implementation of Deep Learning of Binary Hash Codes for Fast Image Retrieval, CVPRW 2015
★ 187keras-ctpn. keras复现场景文本检测网络CPTN: 《Detecting Text in Natural Image with Connectionist Text Proposal Network》;欢迎试用,关注,并反馈问题...
★ 107keras-yolo3. A Keras implementation of YOLOv3 (Tensorflow backend)
★ 2Navigation-system. Java
★ 4Compress-software. Java
★ 7Course-scheduling-system. Java
★ 24YOLO-TrafficDetection. Python
★ 35caffe. Caffe: a fast open framework for deep learning.
★ 35kpytorch-vsumm-reinforce. Unsupervised video summarization with deep reinforcement learning (AAAI'18)
★ 505GitRepo. C#
★ 2996.ICU. Repo for counting stars and contributing. Press F to pay respect to glorious developers.
★ 277kWMSManager. Java
★ 8