This is your work, valued
Email: scz.wangxiao@gmail.com
Temporal-Language-Grounding-in-videos. Temporal Moment(Action) Localization via Language / Temporal Language Grounding / Video Moment Retrieval
★ 101video-ReTaKe. Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
★ 40ms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15khello-agents. 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
★ 70kVACE. [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing
★ 3.9kMegatron-LM. Ongoing research training transformer models at scale
★ 17kAI-Guide-and-Demos-zh_CN. 这是一份入门AI/LLM大模型的逐步指南,包含教程和演示代码,带你从API走进本地大模型部署和微调,代码文件会提供Kaggle或Colab在线版本,即便没有显卡也可以进行学习。项目中还开设了一个小型的代码游乐场🎡,你可以尝试在里面实验一些有意思的AI脚本。同时,包含李宏毅 (HUNG-YI LEE)2024生成式人工智能导论课程的完整中文镜像作业。
★ 4.4kEcho-Infinity. Official repo for paper "Echo-Infinity: Learnable Evolving Memory for Real-Time Infinite Video Generation"
★ 105live-vlm-webui. Real-time Vision Language Model interaction via webcam - WebRTC-based web interface
★ 408academic-research-skills. Academic Research Skills for Claude Code: research → write → review → revise → finalize
★ 40kgpu_vllm_tutorial. HTML
★ 13March7thAssistant. 崩坏:星穹铁道全自动 三月七小助手
★ 11khybrid-forcing. Python
★ 32Awesome-Multimodal-Modeling. Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 508Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18klearn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73klarge-performance-model.github.io. LPM 1.0: Video-based Character Performance Model
★ 361verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kPaddleOCR. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
★ 87kAwesome-VLM-Streaming-Video. 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for streaming video.
★ 190deepagents. The batteries-included agent harness.
★ 27kdeer-flow. An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
★ 78kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 92kHelios. Helios: Real Real-Time Long Video Generation Model
★ 2kAvatarForcing. [CVPR 2026] Official Pytorch implementation of Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
★ 340LiveTalk. Python
★ 329ProPainter. [ICCV 2023] ProPainter: Improving Propagation and Transformer for Video Inpainting
★ 6.8kTurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kMemFlow. Official Implementation of "MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives"
★ 216TransNetV2. TransNet V2: Shot Boundary Detection Neural Network
★ 1kUniReTaKe. [UniReTaKe] Unified KV Cache Compression for Long-Context Video Understanding and Generation (Extended version of AdaReTaKe, ACL 2025)
★ 1ComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kLongVie. Python
★ 334LTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kVTP. [ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation
★ 496sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kco-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5kWorldMM. [CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
★ 99CaptionQA. [CVPR '26] CaptionQA: Is Your Caption as Useful as the Image Itself?
★ 38LongVT. [CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
★ 258tuna.
★ 94Awesome-Video-Agent. A collection of awesome think with videos papers.
★ 100DataFlow. Easy Data Preparation with latest LLMs-based Operators and Pipelines.
★ 7.1khithesis. 嗨!thesis!哈尔滨工业大学毕业论文LaTeX模板
★ 2.4kSana. SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
★ 8.6kOneReward. Python
★ 348RAE. Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
★ 2kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5kLongLive. Long Video Gen Infrastructure
★ 2.5kAwesome-World-Models. A curated list of resources and papers on World Models — video world models, long video generation, unified generation-understanding, distillation and acceleration.
★ 3LLaVA-OneVision-2. Fully Open Framework for Democratized Multimodal Training
★ 1.2kMoBA. MoBA: Mixture of Block Attention for Long-Context LLMs
★ 2.2kMoChaBench. Official Repo for MoCha Towards Movie-Grade Talking Character Synthesis
★ 62Wan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kOpen-Qwen2VL. [COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
★ 314MME-Emotion. Official repository for the paper “MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models”
★ 47video-SALMONN-2. video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is developed by the Department of Electronic Engineering at Tsinghua University and ByteDance.
★ 204RAG-Retrieval. Unify Efficient Fine-tuning of RAG Retrieval, including Embedding, ColBERT, ReRanker.
★ 1.1kresume. An elegant \LaTeX\ résumé template. 大陆镜像 https://gods.coding.net/p/resume/git
★ 11kShotBench. ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
★ 102Keye. Python
★ 809FastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kgen-omnimatte-public. Generative Omnimatte (CVPR 2025)
★ 186Seed1.5-VL. Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
★ 1.6kVideo-XL. 🔥🔥First-ever hour scale video understanding models
★ 626TalkingMachines. TalkingMachines
★ 178nexrender. 📹 Data-driven render automation for After Effects
★ 1.8kMovieBench. [CVPR 2025] A Hierarchical Movie Level Dataset for Long Video Generation
★ 99aftereffects-aep-parser. ✨ An unofficial parser for Adobe After Effects *.aep project files
★ 128py-aep. .aep (After Effects Project) editing in Python
★ 53Awesome-Spatial-Intelligence-in-VLM. A paper list for spatial reasoning
★ 767MMLongBench. The official repo of the paper "MMLongBench Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly"
★ 175Qwen-Native-Sparse-Attention. qwen-nsa
★ 87native-sparse-attention-triton. Efficient triton implementation of Native Sparse Attention.
★ 284native-sparse-attention-pytorch. Implementation of the sparse attention pattern proposed by the Deepseek team in their "Native Sparse Attention" paper
★ 811native-sparse-attention. 🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
★ 1kNSA-pytorch. DeepSeek Native Sparse Attention pytorch implementation
★ 119gaze-estimation. Real-time gaze estimation with ResNet, MobileNet and MobileOne - PyTorch training, ONNX Runtime inference, pretrained weights.
★ 205Eagle. Eagle: Frontier Vision-Language Models with Data-Centric Strategies
★ 3.3kAwesome-Controllable-Video-Generation. [ArXiv 2025] A survey about controllable video generation: This repo is the official awesome of "Controllable video generation: A survey"
★ 760Qwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kHumanOmni. HumanOmni
★ 240mini-omni2. Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities。
★ 1.9kSeerAttention. SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
★ 213ACL25-AdaReTaKe. Official implementation of paper AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
★ 91GameGen-X.
★ 341Awesome-Embodied-AI-Job. Lumina Robotics Talent Call | Lumina社区具身智能招贤榜 | A list for Embodied AI / Robotics Jobs (PhD, RA, intern, etc
★ 1.5kVideoLLaMA3. Frontier Multimodal Foundation Models for Image and Video Understanding
★ 1.2kR1-V. Witness the aha moment of VLM with less than $3.
★ 4.1kCosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kkvpress. LLM KV cache compression made easy
★ 1.2kAwesome-LLM-Long-Context-Modeling. 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
★ 2.1kVideoChat-Flash. [ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
★ 526video-ReTaKe. Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
★ 40SparseVLMs. [ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".
★ 268omegalabs-bittensor-subnet. The World's Largest Decentralized AGI Multimodal Dataset
★ 59KVCache-Factory. Unified KV Cache Compression Methods for Auto-Regressive Models
★ 1.4kVILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kH2O. [NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.
★ 529LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kMInference. [NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
★ 1.2kMLVU. 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark
★ 266lmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kLongLLaVA. LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architecture
★ 211vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kGroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10ksapiens. High-resolution models for human tasks.
★ 5.4kLOOK-M. [EMNLP 2024 Findings🔥] Official implementation of ": LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference"
★ 103easy-rl. 强化学习中文教程(蘑菇书🍄),在线阅读地址:https://datawhalechina.github.io/easy-rl/
★ 15kVript. Python
★ 161Open-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kST-LLM. [ECCV 2024🔥] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"
★ 153unmasked_teacher. [ICCV2023 Oral] Unmasked Teacher: Towards Training-Efficient Video Foundation Models
★ 348Video-ChatGPT. [ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
★ 1.5ktarsier. Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.
★ 548LLaVA-NeXT. Python
★ 4.7kOpen-LLaVA-NeXT. An open-source implementation for training LLaVA-NeXT.
★ 439ShareGPT4V. [ECCV 2024] ShareGPT4V: Improving Large Multi-modal Models with Better Captions
★ 259InternLM-XComposer. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2.9kMAP-NEO. Python
★ 986PLLaVA. Official repository for the paper PLLaVA
★ 670MediaCrawler. 小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫
★ 59kPanda-70M. [CVPR 2024] Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
★ 702DiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kLangchain-Chatchat. Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
★ 38kmmengine. OpenMMLab Foundational Library for Training Deep Learning Models
★ 1.5kkatna. Tool for automating common video key-frame extraction, video compression and Image Auto-crop/Image-resize tasks
★ 398video-keyframe-detector. It is a simple python tool to extract key-frames from a video file using peak estimation from frame difference.
★ 215OmniDataComposer.
★ 16Omni-VideoAssistant. Video QA Assistant based on LLMs with frame convolution
★ 192Chat-UniVi. [CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
★ 943flash-attention. Fast and memory-efficient exact attention
★ 25kLanguageBind. 【ICLR 2024🔥】 Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
★ 883label-studio. Label Studio is a multi-type data labeling and annotation tool with standardized output format
★ 28ktext2vec. text2vec, text to vector. 文本向量表征工具,把文本转化为向量矩阵,实现了Word2Vec、RankBM25、Sentence-BERT、CoSENT等文本表征、文本相似度计算模型,开箱即用。
★ 5kVideo-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kemoji-cheat-sheet. A markdown version emoji cheat sheet
★ 14kjust-ask. [ICCV 2021 Oral + TPAMI] Just Ask: Learning to Answer Questions from Millions of Narrated Videos
★ 127Video-Bench. A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!
★ 140Semi-supervised-learning. A Unified Semi-Supervised Learning Codebase (NeurIPS'22)
★ 1.6kerror_norm_truncation. Code Repository for Error Norm Truncation [ICLR 2024 Spotlight]
★ 32021-NeurIPS-NCR. Python
★ 82LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kHL-Net. Jupyter Notebook
★ 13MiniGPT-5. Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"
★ 868AnimateDiff. Official implementation of AnimateDiff.
★ 12kChatReviewer. ChatReviewer: 使用ChatGPT分析论文优缺点,提出改进建议
★ 1.4kLLM-scientific-feedback. Can large language models provide useful feedback on research papers? A large-scale empirical analysis.
★ 534Cream. This is a collection of our NAS and Vision Transformer work.
★ 1.8kYouku-mPLUG. Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
★ 307stable-diffusion-webui. Stable Diffusion web UI
★ 164kDiST. ICCV2023: Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer Learning
★ 41VALOR. [TPAMI2024] Codes and Models for VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
★ 311MERL-RAV_dataset. [CVPR 2020] MERL-RAV Dataset contains over 19k faces annotated with 68 landmarks, with the additional information of whether each landmark is unoccluded, self-occluded or externally occluded.
★ 41Expression_Recognition. Tracking faces and multiple models detection from faces such as: Gender, Expressions, Illumination, Pose, Occlusion, Age, Makeup.
★ 11AI_power. AI toolbox and pretrain models.
★ 42clip-as-service. 🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
★ 13kAsk-Anything. [CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
★ 3.3kmPLUG-Owl. mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
★ 2.5kdata-centric-AI. A curated, but incomplete, list of data-centric AI resources.
★ 1.2kPySceneDetect. :movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
★ 5.1kMM23-RTQ. ACM Multimedia 2023 (Oral) - RTQ: Rethinking Video-language Understanding Based on Image-text Model
★ 15MM23-TSGVs. ACM Multimedia 2023 - Temporal Sentence in Streaming Videos
★ 10langchain. The agent engineering platform.
★ 143kPTI. Official Implementation for "Pivotal Tuning for Latent-based editing of Real Images" (ACM TOG 2022) https://arxiv.org/abs/2106.05744
★ 929Next3D. [CVPR 2023 Highlight] Next3D: Generative Neural Texture Rasterization for 3D-Aware Head Avatars
★ 501MM22-RADAR. ACM Multimedia 2022 - Micro-video Tagging via Jointly Modeling Social Influence and Tag Relation
★ 7AVSpeechDownloader. Simple python script for downloading AVSpeech Dataset
★ 47InternVideo. [ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2.3kface-alignment. :fire: 2D and 3D Face alignment library build using pytorch
★ 7.5kRobustVideoMatting. Robust Video Matting in PyTorch, TensorFlow, TensorFlow.js, ONNX, CoreML!
★ 9.5kGFPGAN. GFPGAN aims at developing Practical Algorithms for Real-world Face Restoration.
★ 38kDeep3DFaceRecon_pytorch. Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set (CVPRW 2019). A PyTorch implementation.
★ 1.9kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kwenlan-spider. 自用爬虫包,能爬取微博知乎和观察者网
★ 5ReferFormer. [CVPR2022] Official Implementation of ReferFormer
★ 356E2FGVI. Official code for "Towards An End-to-End Framework for Flow-Guided Video Inpainting" (CVPR2022)
★ 1.2kFollowYourPose. [AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text-to-Video Generation using Pose-Free Videos"
★ 1.4kmake-a-video-pytorch. Implementation of Make-A-Video, new SOTA text to video generator from Meta AI, in Pytorch
★ 2kInfiNet. Implementation of DiffusionOverDiffusion architecture presented in NUWA-XL in a form of ControlNet-like module on top of ModelScope text2video model for extremely long video generation.
★ 85AutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kCLIP4Clip. An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
★ 1kOFA. Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
★ 2.6kTaskMatrix. Python
★ 34kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kkinetics-dataset. Shell
★ 982AI_Tutorial. 大厂发布的AI落地实践、顶尖实验室的最新论文、工业界的真实踩坑记录
★ 3.7kCLiMB. The Continual Learning in Multimodality Benchmark
★ 68STAN. Official PyTorch implementation of the paper "Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring"
★ 107Awesome-ChatGPT. ChatGPT资料汇总学习,持续更新......
★ 4.2kscaling-laws-openclip. Reproducible scaling laws for contrastive language-image learning (https://arxiv.org/abs/2212.07143)
★ 201einops. Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
★ 9.6kopen_clip. An open source implementation of CLIP.
★ 14kpytorch_violet. A PyTorch implementation of VIOLET
★ 138TubeDETR. [CVPR 2022 Oral] TubeDETR: Spatio-Temporal Video Grounding with Transformers
★ 194X2-VLM. All-In-One VLM: Image + Video + Transfer to Other Languages / Domains (TPAMI 2023)
★ 170unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kvideo2dataset. Easily create large video dataset from video urls
★ 662img2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kDAMO-ConvAI. DAMO-ConvAI: The official repository which contains the codebase for Alibaba DAMO Conversational AI.
★ 1.6kHsuanTienLin_MachineLearning.
★ 2.8k