This is your work, valued
VLP. Vision-Language Pre-training for Image Captioning and Question Answering
★ 420e2e-gLSTM-sc. Code for paper "Image Caption Generation with Text-Conditional Semantic Attention"
★ 61YouCook2-Leaderboard. A one-stop shop for YouCook2 info such as leaderboard and recent advances on (cooking) video retrieval and captioning.
★ 41densecap. Dense video captioning in PyTorch
★ 41ProcNets-YouCook2. Source code for paper "Towards Automatic Learning of Procedures from Web Instructional Videos"
★ 34anet2016-cuhk-feature. Feature Extraction Toolbox from CUHKÐZ&SIAT submission to ActivityNet 2016
★ 32coco-caption. kdexd/coco-caption@de6f385
★ 26detectron-vlp. Detectron for image/video region feature extraction, inspired by Xinlei's repo
★ 22Negotiation-based-MARL. Source code for journal paper "Multiagent Reinforcement Learning With Sparse Interactions by Negotiation and Knowledge Transfer"
★ 13pytorch-pretrained-BERT. 📖The Big-&-Extending-Repository-of-Transformers: Pretrained PyTorch models for Google's BERT, OpenAI GPT & GPT-2, Google/CMU Transformer-XL.
★ 11densevid_eval_spice. Evaluation code for Dense-Captioning Events in Videos (with SPICE)
★ 7densevid_eval. Evaluation code for Dense-Captioning Events in Videos
★ 4faster-rcnn.pytorch. A faster pytorch implementation of faster r-cnn
★ 2Kimi-K3. Open Frontier Intelligence
★ 7.6kAgentENV. AgentENV (AENV) is a distributed platform for running agent environments at scale.
★ 2.6kMoonEP. MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
★ 938MetaClaw. 🦞 Just talk to your agent — it learns and EVOLVES 🧬.
★ 3.5kAgentsMesh. The AI Agent Workforce Platform. Run a hundred AI coding agents across your own machines — schedule, isolate, and steer them all from one console.
★ 2.3kedict. 🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
★ 16kMatrix-Game. Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
★ 2.3kHunyuanWorld-1.0. Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
★ 2.9kADIEE. Code for ICCV 2025 paper: "ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation"
★ 9cline. Autonomous coding agent as an SDK, IDE extension, or CLI assistant.
★ 65kFunClip. FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
★ 6.1k1d-tokenizer. This repo contains the code for 1D tokenizer and generator
★ 1.2kOpen-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kVLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kktransformers. A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
★ 19kgorilla. Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
★ 13kmultimodal_textbook. [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"
★ 196cosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11kxLAM. xLAM: A Family of Large Action Models to Empower AI Agent Systems
★ 636VILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8krho. Repo for Rho-1: Token-level Data Selection & Selective Pretraining of LLMs.
★ 471Cosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kmle-bench. MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
★ 1.7kAria. Codebase for Aria - an Open Multimodal Native MoE
★ 1.1kogx. Open GenAI Stack
★ 8.4kllama-stack-apps. Agentic components of the Llama Stack APIs
★ 4.3kAI-Scientist. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑🔬
★ 14kMiniCPM-V. A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kSpider2-V. [NeurIPS 2024] Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
★ 153Cradle. The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any computer task by enabling strong reasoning abilities, self-improvment, and skill curation, in a standardized general environment with minimal requirements.
★ 2.6kuniversal_manipulation_interface. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
★ 1.5kAwesome-Video-Diffusion. A curated list of recent diffusion models for video generation, editing, and various other applications.
★ 5.7kpinokio. AI Browser
★ 7.8kFooocus. Focus on prompting and generating
★ 52kSadTalker. [CVPR 2023] SadTalker:Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation
★ 14kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kVGen. Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models
★ 3.2kgpt-fast. Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.
★ 6.2kHeyGenClone. A simple and open-source analogue of the HeyGen system
★ 1kbasaran. Basaran is an open-source alternative to the OpenAI text completion API. It provides a compatible streaming API for your Hugging Face Transformers-based text generation models.
★ 1.3kCogVLM. a state-of-the-art-level open visual language model | 多模态预训练模型
★ 6.7kstreaming-llm. [ICLR 2024] Efficient Streaming Language Models with Attention Sinks
★ 7.3ktriton. Development repository for the Triton language and compiler
★ 20kChatDev. ChatDev 2.0: Dev All through LLM-powered Multi-Agent Collaboration
★ 34kmodular. The Modular Platform (includes MAX & Mojo)
★ 27kMulti-Modality-Arena. Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
★ 565VALL-E-X. An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
★ 7.9kAgentSims. AgentSims is an easy-to-use infrastructure for researchers from all disciplines to test the specific capacities they are interested in.
★ 959RandBox. [ICCV 2023] PyTorch implementation of RandBox
★ 58generative_agents. Generative Agents: Interactive Simulacra of Human Behavior
★ 22kAgentBench. A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
★ 3.6kMetaGPT. 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
★ 70kclippinator. AI programming assistant
★ 410Wav2Lip. This repository contains the codes of "A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild", published at ACM Multimedia 2020. For HD commercial model, please try out Sync Labs
★ 13kLLM-As-Chatbot. LLM as a Chatbot Service
★ 3.3kexllama. A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
★ 2.9kToolBench. [ICLR'24 spotlight] An open platform for training, serving, and evaluating large language model for tool learning.
★ 5.7kllama.cpp. LLM inference in C/C++
★ 122kawesome-langchain. 😎 Awesome list of tools and projects with the awesome LangChain framework
★ 9.5kso-vits-svc. SoftVC VITS Singing Voice Conversion
★ 28kBookGPT. Writes complete books with given paramters, using GPT-3.
★ 350mmc4. MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
★ 954openai-cookbook. Examples and guides for using the OpenAI API
★ 75kLLaMA-Adapter. [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5.9kwolverine. Python
★ 5.1kAgentGPT. 🤖 Assemble, configure, and deploy autonomous AI Agents in your browser.
★ 36kdolly. Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform
★ 11kGPT-4-LLM. Instruction Tuning with GPT-4
★ 4.3kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kopenplayground. An LLM playground you can run on your laptop
★ 6.4kbabyagi. Python
★ 22karxiv-bot. Python
★ 163chatgpt-well-known-plugin-finder. Checks Alexa's top 1M websites for the presence of OpenAI's new .well-known/ai-plugin.json files
★ 174MAD. MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions
★ 177EasyLM. Large language models (LLMs) made easy, EasyLM is a one stop solution for pre-training, finetuning, evaluating and serving LLMs in JAX/Flax.
★ 2.5kAutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kgpt4all. GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77kgpt_academic. 为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
★ 71kchatgpt-retrieval-plugin. The ChatGPT Retrieval Plugin lets you easily find personal or work documents by asking questions in natural language.
★ 21kllama_index. LlamaIndex is the leading document agent and OCR platform
★ 51kawesome-totally-open-chatgpt. A list of totally open alternatives to ChatGPT
★ 4.8kalpaca-lora. Instruct-tune LLaMA on consumer hardware
★ 19kOpenChatKit. Python
★ 9ktrl. Train transformer language models with reinforcement learning.
★ 19kxla. A machine learning compiler for GPUs, CPUs, and ML accelerators
★ 4.4koptimate. A collection of libraries to optimise AI model performances
★ 8.3khh-rlhf. Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"
★ 1.9kControlNet. Let us control diffusion models!
★ 34klangchain. The agent engineering platform.
★ 143kHIR. Python
★ 157summarize.site. Summarize web pages using OpenAI ChatGPT
★ 754EVA. EVA Series: Visual Representation Fantasies from BAAI
★ 2.7kOpen-Assistant. OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
★ 37knext.js. The React Framework
★ 141kbig_vision. Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
★ 3.5kprompts.chat. f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
★ 167kmind-vis. Code base for MinD-Vis
★ 795rl-teacher. Code for Deep RL from Human Preferences [Christiano et al]. Plus a webapp for collecting human feedback
★ 564DiffusionDet. [ICCV2023 Best Paper Finalist] PyTorch implementation of DiffusionDet (https://arxiv.org/abs/2211.09788)
★ 2.3kjax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kTokShift-Transformer. Python
★ 70OFA. Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
★ 2.6kcoyo-dataset. COYO-700M: Large-scale Image-Text Pair Dataset
★ 1.3kVideoGPT. Jupyter Notebook
★ 1.1ksygil-webui. Stable Diffusion web UI
★ 7.9kml-cvnets. CVNets: A library for training computer vision networks
★ 2kPrompt-align. [ICCV 2023] Prompt-aligned Gradient for Prompt Tuning
★ 170open_clip. An open source implementation of CLIP.
★ 14kgpt-neox. An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries
★ 7.4kstable-diffusion. A latent text-to-image diffusion model
★ 73kgget. 🧬 gget enables efficient querying of genomic reference databases
★ 1.2kMSCLIP. Official Code of ECCV 2022 paper MS-CLIP
★ 91CodeRL. This is the official code for the paper CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning (NeurIPS22).
★ 573GroupViT. Official PyTorch implementation of GroupViT: Semantic Segmentation Emerges from Text Supervision, CVPR 2022.
★ 788MineDojo. Building Open-Ended Embodied Agents with Internet-Scale Knowledge
★ 2.2kmaskvit.
★ 74parti.
★ 1.6kt5x. Python
★ 3kjonbarron.github.io. HTML
★ 3.6kBIG-bench. Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models
★ 3.2kCogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kFlagAI. FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.
★ 3.9kVidIL. Pytorch code for Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
★ 117E2FGVI. Official code for "Towards An End-to-End Framework for Flow-Guided Video Inpainting" (CVPR2022)
★ 1.2kmetaseq. Repo for external large-scale work
★ 6.6kort. Accelerate PyTorch models with ONNX Runtime
★ 369latent-diffusion. High-Resolution Image Synthesis with Latent Diffusion Models
★ 14kdalle-2-preview.
★ 1kVQ-Diffusion. Python
★ 487GLIP. Grounded Language-Image Pre-training
★ 2.6kclip-event. Python
★ 107how-do-vits-work. (ICLR 2022 Spotlight) Official PyTorch implementation of "How Do Vision Transformers Work?"
★ 822mediapipe. Cross-platform, customizable ML solutions for live and streaming media.
★ 36ks4. Structured state space sequence models
★ 2.9kcliport. CLIPort: What and Where Pathways for Robotic Manipulation
★ 547BEVT. PyTorch implementation of BEVT (CVPR 2022) https://arxiv.org/abs/2112.01529
★ 161svpc. Official implementation of state-aware video procedural captioning (ACM MM 2021)
★ 9moment_detr. [NeurIPS 2021] Moment-DETR code and QVHighlights dataset
★ 349mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4ktransformer-ls. Official PyTorch Implementation of Long-Short Transformer (NeurIPS 2021).
★ 228SSTAP. Code for our CVPR 2021 Paper "Self-Supervised Learning for Semi-Supervised Temporal Action Proposal".
★ 72NUWA. A unified 3D Transformer Pipeline for visual synthesis
★ 2.8kSPACH. Python
★ 205LOVEU-CVPR2021. Python
★ 27DPT. Dense Prediction Transformers
★ 2.3kGradCache. Run Effective Large Batch Contrastive Learning Beyond GPU/TPU Memory Constraint
★ 444VT-SSum.
★ 23FlatCLR. FlatNCE: A Novel Contrastive Representation Learning Objective
★ 90Awesome-Zero-Shot-Object-Detection.
★ 129decord. An efficient video loader for deep learning with smart shuffling that's super easy to digest
★ 2.5ktaming-transformers. Taming Transformers for High-Resolution Image Synthesis
★ 6.5kVideo-Swin-Transformer. This is an official implementation for "Video Swin Transformers".
★ 1.7kDALL-E. PyTorch package for the discrete VAE used for DALL·E.
★ 11kVL-T5. PyTorch code for "Unifying Vision-and-Language Tasks via Text Generation" (ICML 2021)
★ 372xcit. Official code Cross-Covariance Image Transformer (XCiT)
★ 681vimpac. Python
★ 73fastmoe. A fast MoE impl for PyTorch
★ 1.9kDataRelease. Data Release for VALUE Benchmark
★ 30StarterCode. Starter Code for VALUE benchmark
★ 79EvaluationTools. Evaluation code and codalab submission examples for the VALUE benchmark.
★ 10ml-contests-conf. ML and DL related contests, competitions and conference challenges.
★ 609vince. Video Noise Contrastive Estimation
★ 66CvT. This is an official implementation of CvT: Introducing Convolutions to Vision Transformers.
★ 609UC2. CVPR 2021 Official Pytorch Code for UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training
★ 34kinetics-dataset. Shell
★ 982simple-amt. A microframework for working with Amazon's Mechanical Turk
★ 303dino. PyTorch code for Vision Transformers training with the Self-Supervised learning method DINO
★ 7.6kpytorchvideo. A deep learning library for video understanding research.
★ 3.6ktesseract. Tesseract Open Source OCR Engine (main repository)
★ 76kPySceneDetect. :movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
★ 5.1kconceptual-12m. Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.
★ 426universal-computation. Official codebase for Pretrained Transformers as Universal Computation Engines.
★ 245Transformer-MM-Explainability. [ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers, a novel method to visualize any Transformer-based network. Including examples for DETR, VQA.
★ 911convit. Code for the Convolutional Vision Transformer (ConViT)
★ 474halonet-pytorch. Implementation of the 😇 Attention layer from the paper, Scaling Local Self-Attention For Parameter Efficient Visual Backbones
★ 199perceiver-pytorch. Implementation of Perceiver, General Perception with Iterative Attention, in Pytorch
★ 1.2kOWOD. (CVPR 2021 Oral) Open World Object Detection
★ 1.1kvissl. VISSL is FAIR's library of extensible, modular and scalable components for SOTA Self-Supervised Learning with images.
★ 3.3kClipBERT. [CVPR 2021 Best Student Paper Honorable Mention, Oral] Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks.
★ 730CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kmmaction2. OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark
★ 5.1kaudio. Data manipulation and transformation for audio signal processing, powered by PyTorch
★ 2.9kTransformer-Explainability. [CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization, a novel method to visualize classifications by Transformer based networks.
★ 2kpytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kDeformable-DETR. Deformable DETR: Deformable Transformers for End-to-End Object Detection.
★ 4kig65m-pytorch. PyTorch 3D video classification models pre-trained on 65 million Instagram videos
★ 265CV_LTH_Pre-training. [CVPR 2021] "The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models" Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Michael Carbin, Zhangyang Wang
★ 69vit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kdigital_video_introduction. A hands-on introduction to video technology: image, video, codec (av1, vp9, h265) and more (ffmpeg encoding). Translations: 🇺🇸 🇨🇳 🇯🇵 🇮🇹 🇰🇷 🇷🇺 🇧🇷 🇪🇸
★ 16kpytorch-coviar. Compressed Video Action Recognition
★ 523faiss. A library for efficient similarity search and clustering of dense vectors.
★ 41kSlowFast. PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.
★ 7.4kcoot-videotext. COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
★ 291vision_transformer. Jupyter Notebook
★ 13kkaldi. kaldi-asr/kaldi is the official location of the Kaldi project.
★ 15kYouCook2-Leaderboard. A one-stop shop for YouCook2 info such as leaderboard and recent advances on (cooking) video retrieval and captioning.
★ 41Mephisto. A suite of tools for managing crowdsourcing tasks from the inception through to data packaging for research use.
★ 310