This is your work, valued
Pyramid-Flow. [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
★ 3.2kLaVIT. LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
★ 603STCAT. [NeurIPS 2022] Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding
★ 54Awesome-Text-to-Video-Generation. A curated list of Text-to-Video Generation papers and BibTeX entries
★ 3Compiler-for-C0-grammer. 一个可以编译C0文法的编译器
★ 1Scene-Graph-Benchmark.pytorch. A new codebase for popular Scene Graph Generation methods (2020). Visualization & Scene Graph Extraction on custom images/datasets are provided. It's also a PyTorch implementation of paper “Unbiased Scene Graph Generation from Biased Training CVPR 2020”
★ 1OpenHands. 🙌 OpenHands: AI-Driven Development
★ 83kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 384kNExT-Vid. Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
★ 22SoraWatermarkCleaner. The fastest and highest-quality deep learning powered Sora2 watermark cleaner.
★ 1.2kms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kAwesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
★ 1.5kVision-Language-Models-Overview. A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
★ 683Megatron-Bridge. Training library for Megatron-based models with bidirectional Hugging Face conversion capability
★ 834EasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kPai-Megatron-Patch. The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
★ 1.6kbig_vision. Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
★ 3.5kAwesome-RL-based-Reasoning-MLLMs. This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
★ 1.4kDMD2. (NeurIPS 2024 Oral 🔥) Improved Distribution Matching Distillation for Fast Image Synthesis
★ 1.4kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kIC-Light. More relighting!
★ 8.5kCosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kPyramid-Flow. [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
★ 3.2kflux. Official inference repo for FLUX.1 models
★ 26kDuix-Mobile. 🚀 The best real-time interactive AI avatar(digital human) with on-premise deployment and <1.5 s latency.
★ 8.2kMINT-1T. 🍃 MINT-1T: A one trillion token multimodal interleaved dataset.
★ 833flash-attention. Fast and memory-efficient exact attention
★ 25kopen_clip. An open source implementation of CLIP.
★ 14kchameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kReNO. [NeurIPS 2024] ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
★ 166rfpp. The codebase of our paper "Improving the Training of Rectified Flows", NeurIPS 2024
★ 133cog-consistent-character. Create images of a given character in different poses
★ 722LongLoRA. Code and documents of LongLoRA and LongAlpaca (ICLR 2024 Oral)
★ 2.7kRectifID. [NeurIPS 2024] RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance
★ 129HunyuanDiT. Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
★ 4.3kpiecewise-rectified-flow. PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator (NeurIPS 2024)
★ 538magvit2-pytorch. Implementation of MagViT2 Tokenizer in Pytorch
★ 668kornia. 🐍 Geometric Computer Vision Library for Spatial AI
★ 11kvideo-subtitle-remover. 基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.
★ 12kInstanceDiffusion. [CVPR 2024] Code release for "InstanceDiffusion: Instance-level Control for Image Generation"
★ 614Lumina-T2X. Lumina-T2X is a unified framework for Text to Any Modality Generation
★ 2.2kAwesome-Video-Datasets. Video datasets
★ 1.7kLatte. [TMLR 2025] Latte: Latent Diffusion Transformer for Video Generation.
★ 1.9kCompressAI. A PyTorch library and evaluation platform for end-to-end compression research
★ 1.6kLearned_Compression. This repository is a paper digest of DNN-based approaches in data compression tasks.
★ 27InstaFlow. :zap: InstaFlow! One-Step Stable Diffusion with Rectified Flow (ICLR 2024)
★ 1.4kRectifiedFlow. Official Implementation of Rectified Flow (ICLR2023 Spotlight)
★ 1.6kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kPixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kDiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kVideoSys. VideoSys: An easy and efficient system for video generation
★ 2kml-ferret. Python
★ 8.7kopen-prompts. Python
★ 790Awesome-Video-Diffusion-Models. [CSUR] A Survey on Video Diffusion Models
★ 2.3kVideo-ChatGPT. [ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
★ 1.5kgenerative-models. Generative Models by Stability AI
★ 27kVideo-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kText2Video-Zero. [ICCV 2023 Oral] Text-to-Image Diffusion Models are Zero-Shot Video Generators
★ 4.2kRewireNeuron. [NeurIPS 2023] Rewiring Neurons in Non-Stationary Environments
★ 9Awesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
★ 5.7kvideo2dataset. Easily create large video dataset from video urls
★ 662webvid. Large-scale text-video dataset. 10 million captioned short videos.
★ 685LaVIT. LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
★ 603vector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kpytorch-fid. Compute FID scores with PyTorch.
★ 3.9kWanJuan1.0. 万卷1.0多模态语料
★ 574CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kmmc4. MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
★ 953Megatron-LM. Ongoing research training transformer models at scale
★ 17kmPLUG-Owl. mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
★ 2.5kwebdataset. A high-performance Python-based I/O system for large (and small) deep learning problems, with strong support for PyTorch.
★ 3.1kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kdatacomp. DataComp: In search of the next generation of multimodal datasets
★ 787LAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kFengshenbang-LM. Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
★ 4.1kawesome-LLMs-In-China. 中国大模型
★ 6.5kChatGLM2-6B. ChatGLM2-6B: An Open Bilingual Chat LLM | 开源双语对话语言模型
★ 16kstanford_alpaca. Code and documentation to train Stanford's Alpaca models, and generate the data.
★ 30kChinese-LLaMA-Alpaca. 中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
★ 19kVideo-LLaMA. [EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
★ 3.1kInternVideo. [ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2.3kimg2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kCVinW_Readings. A collection of papers on the topic of ``Computer Vision in the Wild (CVinW)''
★ 1.4kAwesome-Anything. General AI methods for Anything: AnyObject, AnyGeneration, AnyModel, AnyTask, AnyX
★ 1.9kDreambooth-Stable-Diffusion. Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
★ 7.7kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1kPyTorch-VAE. A Collection of Variational Autoencoders (VAE) in PyTorch.
★ 7.7kjava-books-collections. :books:Java编程书籍收集分享。Java programming books collection to share.:rocket:
★ 6.7kAutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kUni-Perceiver. Python
★ 291bot-on-anything. A large model-based chatbot builder that can quickly integrate AI models (including ChatGPT, Claude, Gemini) into various software applications (such as Telegram, Gmail, Slack, and websites).
★ 4.2kControlNet. Let us control diffusion models!
★ 34ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kstable-diffusion-webui. Stable Diffusion web UI
★ 164kstable-diffusion. A latent text-to-image diffusion model
★ 73koptimate. A collection of libraries to optimise AI model performances
★ 8.3kOpen-Assistant. OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
★ 37kSceneGraphParser. A python toolkit for parsing captions (in natural language) into scene graphs (as symbolic representations).
★ 595stanza. Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages
★ 7.9kchatgpt-google-extension. This project is deprecated. Check my new project ChatHub:
★ 13kannotated_deep_learning_paper_implementations. 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠
★ 67kAwesome-ChatGPT. ChatGPT资料汇总学习,持续更新......
★ 4.2kmoment_detr. [NeurIPS 2021] Moment-DETR code and QVHighlights dataset
★ 349Cam2BEV. TensorFlow Implementation for Computing a Semantically Segmented Bird's Eye View (BEV) Image Given the Images of Multiple Vehicle-Mounted Cameras.
★ 790fairscale. PyTorch extensions for high performance and large scale training.
★ 3.4kdeit. Official DeiT repository
★ 4.4kSTCAT. [NeurIPS 2022] Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding
★ 54PKUAutoElective. 北大选课网补退选阶段自动选课小工具
★ 745OS-SGG. This is the repository for papr "One-Shot Scene Graph Generation"
★ 16pracmln. Markov Logic Networks in Python
★ 138daxigua. 最简单的魔改发布『 合成大西瓜 』,配套改图工具,不用改代码,修改配置即可!
★ 1.4kAdelaiDet. AdelaiDet is an open source toolbox for multiple instance-level detection and recognition tasks.
★ 3.5kawesome-computer-vision. A curated list of awesome computer vision resources
★ 23kawesome-scene-graph. A curated list of scene graph generation and related area resources. :-)
★ 90jieba. 结巴中文分词
★ 35klac. 百度NLP:分词,词性标注,命名实体识别,词重要性
★ 4kPyTorch-Networks. Pytorch implementation of cnn network
★ 2.1kChinese-NER. Chinese NER using BiLSTM/BERT + CRF
★ 63pytorch-crf. (Linear-chain) Conditional random field in PyTorch.
★ 980fairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
★ 32kfairseq-lua. Facebook AI Research Sequence-to-Sequence Toolkit
★ 3.7kconceptnet5. Code for building ConceptNet from raw data.
★ 3kneural-motifs. Code for Neural Motifs: Scene Graph Parsing with Global Context (CVPR 2018)
★ 545DeepLearningExamples. State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
★ 15kVC-R-CNN. [CVPR 2020] The official pytorch implementation of ``Visual Commonsense R-CNN''
★ 357CLUENER2020. CLUENER2020 中文细粒度命名实体识别 Fine Grained Named Entity Recognition
★ 1.5kCLUEPretrainedModels. 高质量中文预训练模型集合:最先进大模型、最快小模型、相似度专门模型
★ 810seamseg. Seamless Scene Segmentation
★ 301het-eccv20. Codes for ECCV paper: "Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation"
★ 16Long-Tailed-Recognition.pytorch. [NeurIPS 2020] This project provides a strong single-stage baseline for Long-Tailed Classification, Detection, and Instance Segmentation (LVIS). It is also a PyTorch implementation of the NeurIPS 2020 paper 'Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal Effect'.
★ 573classifier-balancing. This repository contains code for the paper "Decoupling Representation and Classifier for Long-Tailed Recognition", published at ICLR 2020
★ 981introRL. Intro to Reinforcement Learning (强化学习纲要)
★ 3.6kCS-Notes. :books: 技术面试必备基础知识、Leetcode、计算机操作系统、计算机网络、系统设计
★ 185kdetectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kScene-Graph-Benchmark.pytorch. A new codebase for popular Scene Graph Generation methods (2020). Visualization & Scene Graph Extraction on custom images/datasets are provided. It's also a PyTorch implementation of paper “Unbiased Scene Graph Generation from Biased Training CVPR 2020”
★ 1.2kPaddleDetection. Object Detection toolkit based on PaddlePaddle. It supports object detection, instance segmentation, multiple object tracking and real-time multi-person keypoint detection.
★ 14kFactorizableNet. Factorizable Net (Multi-GPU version): An Efficient Subgraph-based Framework for Scene Graph Generation
★ 220GraphNorm. [ICML 2021] GraphNorm: A Principled Approach to Accelerating Graph Neural Network Training (official implementation)
★ 106mmdetection. OpenMMLab Detection Toolbox and Benchmark
★ 33kC-HOI. Cascaded Human-Object Interaction Recognition (CVPR2020)
★ 84Transferable-Interactiveness-Network. Code for Transferable Interactiveness Knowledge for Human-Object Interaction Detection. (CVPR'19, TPAMI'21)
★ 239LIGHTEN-Learning-Interactions-with-Graphs-and-Hierarchical-TEmporal-Networks-for-HOI. Python
★ 16vRGV. Visual Relation Grounding in Videos (ECCV'20, Spotlight)
★ 57VidSTG-Dataset. This repository provides the dataset introduced by the paper "Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences"
★ 69deep_sort. Simple Online Realtime Tracking with a Deep Association Metric
★ 6.2kActivityNet. This repository is intended to host tools and demos for ActivityNet
★ 975google-images-download. Python Script to download hundreds of images from 'Google Images'. It is a ready-to-run code!
★ 8.7kplaces365. The Places365-CNNs for Scene Classification
★ 2.1kPreprocess-I3D-Deepmind. Preprocess data for deepmind's kinetics i3d network
★ 14clip-as-service. 🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
★ 13kkinetics-i3d. Convolutional neural network model for video classification trained on the Kinetics dataset.
★ 1.8kBUAA-1606-GradReq. 毕业要求自查,欢迎大家补充订正
★ 14deep-learning-with-python-notebooks. Jupyter notebooks for the code samples of the book "Deep Learning with Python"
★ 20kVidVRD-helper. To keep updates with VRU Grand Challenge, please use https://github.com/NExTplusplus/VidVRD-helper
★ 102Mask_RCNN. Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow
★ 26kmaskrcnn-benchmark. Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch.
★ 9.4kfaster_rcnn_pytorch. Faster RCNN with PyTorch
★ 1.8kpytorch-faster-rcnn. pytorch1.0 updated. Support cpu test and demo. (Use detectron2, it's a masterpiece)
★ 1.8ksimple-faster-rcnn-pytorch. A simplified implemention of Faster R-CNN that replicate performance from origin paper
★ 4kDive-into-DL-PyTorch. 本项目将《动手学深度学习》(Dive into Deep Learning)原书中的MXNet实现改为PyTorch实现。
★ 19kpytorch-handbook. pytorch handbook是一本开源的书籍,目标是帮助那些希望和使用PyTorch进行深度学习开发和研究的朋友快速入门,其中包含的Pytorch教程全部通过测试保证可以成功运行
★ 22kkeras_frcnn. Keras Implementation of faster-rcnn
★ 522Keras-frcnn. Keras Implementation of Faster R-CNN
★ 402barefoot. Map matching for the cities in China
★ 10deepgtt. DeepGTT: Learning Travel Time Distributions with Deep Generative Model
★ 36fmm. Fast map matching, an open source framework in C++
★ 1kSocial_School. 一个基于 React Native 的校园社交APP.
★ 24STGCN_IJCAI-18. [IJCAI'18] Spatio-Temporal Graph Convolutional Networks
★ 1.2kThe-Art-Of-Programming-By-July-2nd. 本项目曾冲到全球第一,干货集锦见本页面最底部,另完整精致的纸质版《编程之法:面试和算法心得》已在京东/当当上销售
★ 22kBaiduTraffic. This repo includes introduction, code and dataset of our paper Deep Sequence Learning with Auxiliary Information for Traffic Prediction (KDD 2018).
★ 246SimpleSSM. ssm所搭建的后台项目🦀🦀🦀
★ 2Compiler. C0 compiler 🎞🎞🎞
★ 4Guuog. NIOWeb服务器🚗🚂🚅✈🚀
★ 3