This is your work, valued
ChatAnything. Official Repo for the Paper: CHATANYTHING: FACETIME CHAT WITH LLM-ENHANCED PERSONAS
★ 378dvit_repo. Python
★ 141rethinking_bottleneck_design. Python
★ 139Refiner_ViT. Python
★ 110OpenSORA. A public repository for reproducing a open source sora comparable video generation model
★ 9NES. code for the paper Neural Epitome Search for for architecture-agnostic compression
★ 8AutoSpace.
★ 4Dataset_Quantization. [ICCV2023] Dataset Quantization
★ 1homepage.io. SCSS
★ 1HOI-DETR. Improving and Evaluating Hand-Object Interaction Detection
★ 180box3d. Box3D is a 3D physics engine for games
★ 5.7kFast-dLLM. Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
★ 1.1kHumanNet. HumanNet: Scaling Human-centric Video Learning to One Million Hours
★ 278VideoRAG. [KDD'2026] "VideoRAG: Chat with Your Videos"
★ 3.2kReVidgen. [ICML 2026🔥]Rethinking Video Generation Model for the Embodied World
★ 89MHLA. [ICLR 2026🔥] MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
★ 150Agent-Memory-Paper-List. The paper list of "Memory in the Age of AI Agents: A Survey"
★ 2.3kawesome-humanoid-manipulation. A curated list of awesome papers and resources on humanoid manipulation, dexterous manipulation, bimanual dexterous manipulation, in-hand manipulation, and humanlike manipulation.
★ 152VeOmni. VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
★ 2.1kBitNet. Official inference framework for 1-bit LLMs
★ 40kAgent-S. Agent S: an open agentic framework that uses computers like a human
★ 12kTradingAgents. TradingAgents: Multi-Agents LLM Financial Trading Framework
★ 95kdeveloper-portfolios. A list of developer portfolios for your inspiration
★ 26kAwesome-GitHub-Repo. 收集整理 GitHub 上高质量、有趣的开源项目。
★ 17kMatrix-Game. Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
★ 2.3k1xgpt. world modeling challenge for humanoid robots
★ 564DeepSeek-R1.
★ 92kopenai-agents-python. A lightweight, powerful framework for multi-agent workflows
★ 28ktabby. Self-hosted AI coding assistant
★ 34kChronoEdit. [ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
★ 701Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kOmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kMeanFlow. Pytorch implementation of MeanFlow on ImageNet and CIFAR10
★ 504debugpy. An implementation of the Debug Adapter Protocol for Python
★ 2.4kEasyOCR. Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
★ 30kIMO25. An AI agent system for solving International Mathematical Olympiad (IMO) problems using Google's Gemini, OpenAI, and XAI APIs.
★ 935motionshop-2. Project page of Motionshop-2: An advanced version of the Motionshop; We add a new fancy feature: 3D Animation Engine; AI for computer graphics
★ 58coze-studio. An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before. Coze your way to AI Agent creation.
★ 21kBagel. Open-source unified multimodal model
★ 6.1kHunyuanCustom. HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
★ 1.2kSecond-Me. Train your AI self, amplify you, bridge the world
★ 16kME-rPPG. ME-rPPG is a memory-efficient realtime rPPG network
★ 42OpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58kVisualThinker-R1-Zero. Explore the Multimodal “Aha Moment” on 2B Model
★ 624Magma. [CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents
★ 1.9kR1-Onevision. R1-onevision, a visual language model capable of deep CoT reasoning.
★ 581Wan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kStep-Video-T2V. Python
★ 3.2kHunyuan3D-2. High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
★ 14kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kDeepSeek-V3. Python
★ 104kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26kTinyZero. Minimal reproduction of DeepSeek R1-Zero
★ 13kOpenHands. 🙌 OpenHands: AI-Driven Development
★ 83kMiniMax-01. The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
★ 3.4kGenesisEnvs. Genesis Reinforcement Learning Environments
★ 137droid. Distributed Robot Interaction Dataset.
★ 385VITA. ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2.5kcosmos. NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
★ 11ktransfusion-pytorch. Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
★ 1.4ktrendFinder. Stay on top of trending topics on social media and the web with AI
★ 4.1kDiffusion-LM. Diffusion-LM
★ 1.2kAgiBot-World. [IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
★ 3.1kNarratoAI. 利用 AI 大模型,一键解说并剪辑视频
★ 11kunsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek, GLM and other models.
★ 69kbrowser-use. 🌐 Make websites accessible for AI agents. Automate tasks online with ease.
★ 107kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kFastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kSlow_Thinking_with_LLMs. A series of technical report on Slow Thinking with LLM
★ 767DiffSensei. Implementation of [CVPR 2025] "DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation"
★ 923MovieGen. Unofficial implementation of Meta's MovieGen models
★ 16superclass. [NeurIPS 2024] Classification Done Right for Vision-Language Pre-Training
★ 223VideoSys. VideoSys: An easy and efficient system for video generation
★ 2kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kLLaMA-Mesh. Unifying 3D Mesh Generation with Language Models
★ 1.2knunchaku. [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
★ 3.9kdask. Parallel computing with task scheduling
★ 14kProtenix. Toward High-Accuracy Open-Source Biomolecular Structure Prediction.
★ 2kGrounded-Video-LLM. [EMNLP 2025 Findings] Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
★ 148bypy. Python client for Baidu Yun (Personal Cloud Storage) 百度云/百度网盘Python客户端
★ 8.6kMinerU. Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
★ 76kresource-stream. GPU programming related news and material links
★ 2.2kCosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7kVideo-XL. 🔥🔥First-ever hour scale video understanding models
★ 626kotaemon. An open-source RAG-based tool for chatting with your documents.
★ 26kswarm. Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
★ 22kMaskDiffusion. Python
★ 12infinigen. Infinite Photorealistic Worlds using Procedural Generation
★ 7.2kGR-1. Code for "Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation"
★ 308HPT. Heterogeneous Pre-trained Transformer (HPT) as Scalable Policy Learner.
★ 543Emu3. Next-Token Prediction is All You Need
★ 2.4kWeb. 千古前端图文教程,超详细的前端入门到进阶知识库。从零开始学前端,做一名精致优雅的前端工程师。
★ 29kaudio-diffusion-pytorch. Audio generation using diffusion models, in PyTorch.
★ 2.1kml-ferret. Python
★ 8.7kvllm-kvcompress. KV cache compression for high-throughput LLM inference
★ 159movie-gen. An open source community implementation of the model from the paper: "Movie Gen: A Cast of Media Foundation Models". Join our community to help implement this model!
★ 60exo. Run frontier AI locally.
★ 47kcrawl4ai. 🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
★ 76kxDiT. xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
★ 2.7klocalsend. An open-source cross-platform alternative to AirDrop
★ 86kLVD-2M. [NeurIPS 2024 D&B Track] Official Repo for "LVD-2M: A Long-take Video Dataset with Temporally Dense Captions"
★ 79ai-toolkit. The ultimate training toolkit for finetuning diffusion models
★ 11kclaude-cookbooks. A collection of notebooks/recipes showcasing some fun and effective ways of using Claude.
★ 51kx-flux. Python
★ 2.2kLongWriter. [ICLR 2025] LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
★ 1.9kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kMS-Diffusion. [ICLR 2025] Official implementation of MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
★ 311Megatron-LM. Ongoing research training transformer models at scale
★ 17kComfyUI_StoryDiffusion. You can using StoryDiffusion in ComfyUI
★ 512vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kmem0. Universal memory layer for AI Agents
★ 62kPaints-UNDO. Understand Human Behavior to Align True Needs
★ 4.1kKolors. Kolors Team
★ 4.6ktorchtitan. A PyTorch native platform for training generative AI models
★ 5.6klloco. The official repo for "LLoCo: Learning Long Contexts Offline"
★ 119ChronoMagic-Bench. [NeurIPS 2024 D&B Spotlight🔥] ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
★ 213cambrian. Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2khallo. Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
★ 8.7kLLM101n. LLM101n: Let's build a Storyteller
★ 38kAwesome-LLM-Compression. Awesome LLM compression research papers and tools.
★ 1.9kgpt4all. GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77kchamp. [ECCV 2024] Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
★ 4.3kS-LoRA. S-LoRA: Serving Thousands of Concurrent LoRA Adapters
★ 1.9kMoRF. Receptive field as experts
★ 5Unique3D. [NeurIPS 2024] Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
★ 3.6kllama3-from-scratch. llama3 implementation one matrix multiplication at a time
★ 15kRectifiedFlow. Official Implementation of Rectified Flow (ICLR2023 Spotlight)
★ 1.6kefficient-kan. An efficient pure-PyTorch implementation of Kolmogorov-Arnold Network (KAN).
★ 4.7kq-diffusion. [ICCV 2023] Q-Diffusion: Quantizing Diffusion Models.
★ 378SpeeD. SpeeD: A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training
★ 188StoryDiffusion. Accepted as [NeurIPS 2024] Spotlight Presentation Paper
★ 6.4kPLLaVA. Official repository for the paper PLLaVA
★ 670corenet. CoreNet: A library for training deep neural networks
★ 7kVAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kMLLM_Factory. A Dead Simple and Modularized Multi-Modal Training and Finetune Framework. Compatible to any LLaVA/Flamingo/QwenVL/MiniGemini etc series models.
★ 19motion-diffusion-model. The official PyTorch implementation of the paper "Human Motion Diffusion Model"
★ 4.1kco-tracker. CoTracker is a model for tracking any point (pixel) on a video.
★ 5kchug. Minimal sharded dataset loaders, decoders, and utils for multi-modal document, image, and text datasets.
★ 163MediaCrawler. 小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫
★ 59kgrok-1. Grok open release
★ 52kLLM-Viewer. Analyze the inference of Large Language Models (LLMs). Analyze aspects like computation, storage, transmission, and hardware roofline model in a user-friendly interface.
★ 666RankSortLoss. Official PyTorch Implementation of Rank & Sort Loss for Object Detection and Instance Segmentation [ICCV2021]
★ 245evolutionary-model-merge. Official repository of Evolutionary Optimization of Model Merging Recipes
★ 1.4kscaling_on_scales. When do we not need larger vision models?
★ 419autofd. Automatic Functional Differentiation in JAX
★ 88bias.
★ 115grok. Python
★ 4.3kmoviepy. Video editing with Python
★ 15kMedusa. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8kmamba. Mamba SSM architecture
★ 19kSora-Generates-Videos-with-Stunning-Geometrical-Consistency. Sora Generates Videos with Stunning Geometrical Consistency
★ 51Moore-AnimateAnyone. Character Animation (AnimateAnyone, Face Reenactment)
★ 3.5kOpen-AnimateAnyone. Unofficial Implementation of Animate Anyone
★ 2.9kEMO. Emote Portrait Alive: Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
★ 7.6kVILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kNeural-Network-Diffusion. We introduce a novel approach for parameter generation, named neural network parameter diffusion (p-diff), which employs a standard latent diffusion model to synthesize a new set of parameters
★ 886minbpe. Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
★ 11kOpenSORA. A public repository for reproducing a open source sora comparable video generation model
★ 9Magic-Me. Codes for ID-Specific Video Customized Diffusion
★ 458StableCascade. Official Code for Stable Cascade
★ 6.5kTinyLlama. The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.
★ 9kLVM. Python
★ 1.8kWeChatMsg.
★ 42kbig_vision. Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
★ 3.5kVideo-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kGPT-4V_Social_Media. GPT-4V(ision) as A Social Media Analysis Engine
★ 39MAgIC. Python
★ 43PixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kChatAnything. Official Repo for the Paper: CHATANYTHING: FACETIME CHAT WITH LLM-ENHANCED PERSONAS
★ 378fiftyone. Refine high-quality datasets and visual AI models
★ 11kFooocus. Focus on prompting and generating
★ 52kDiT-3D. 🔥🔥🔥Official Codebase of "DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation"
★ 321OpenMoE. A family of open-sourced Mixture-of-Experts (MoE) Large Language Models
★ 1.7kCoDeF. [CVPR'24 Highlight] Official PyTorch implementation of CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
★ 4.8ktext-to-motion. Official implementation for "Generating Diverse and Natural 3D Human Motions from Texts (CVPR2022)."
★ 710gigagan-pytorch. Implementation of GigaGAN, new SOTA GAN out of Adobe. Culmination of nearly a decade of research into GANs
★ 1.9kDataset_Quantization. [ICCV2023] Dataset Quantization
★ 261MovieChat. [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706DragDiffusion. [CVPR2024, Highlight] Official code for DragDiffusion
★ 1.3kpytorch_geometric. Graph Neural Network Library for PyTorch
★ 24kLISA. Project Page for "LISA: Reasoning Segmentation via Large Language Model"
★ 2.7kstarcoder. Home of StarCoder: fine-tuning & inference!
★ 7.5kAdaFace-dev. A Versatile Face Encoder for Zero-Shot Diffusion Model Personalization
★ 24RealChar. 🎙️🤖Create, Customize and Talk to your AI Character/Companion in Realtime (All in One Codebase!). Have a natural seamless conversation with AI everywhere (mobile, web and terminal) using LLM OpenAI GPT3.5/4, Anthropic Claude2, Chroma Vector DB, Whisper Speech2Text, ElevenLabs Text2Speech🎙️🤖
★ 6.2kbubogpt. BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
★ 510HumanML3D. HumanML3D: A large and diverse 3d human motion-language dataset.
★ 1.5kquivr. Opiniated RAG for integrating GenAI in your apps 🧠 Focus on your product rather than the RAG. Easy integration in existing products with customisation! Any LLM: GPT4, Groq, Llama. Any Vectorstore: PGVector, Faiss. Any Files. Anyway you want.
★ 39kmilvus. Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
★ 45kgenerative-models. Generative Models by Stability AI
★ 27kai-getting-started. A Javascript AI getting started stack for weekend projects, including image/text models, vector stores, auth, and deployment configs
★ 4.1kChatPaper. Use ChatGPT to summarize the arXiv papers. 全流程加速科研,利用chatgpt进行论文全文总结+专业翻译+润色+审稿+审稿回复
★ 20kchain-of-thought-hub. Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
★ 2.8kDragGAN. Official Code for DragGAN (SIGGRAPH 2023)
★ 36kVisualGLM-6B. Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型
★ 4.2kAudioGPT. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
★ 10kmissing_aware_prompts. Multimodal Prompting with Missing Modalities for Visual Recognition, CVPR'23
★ 232consistency_models. Official repo for consistency models.
★ 6.5kAutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kImageReward. [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7keai-vc. The repository for the largest and most comprehensive empirical study of visual foundation models for Embodied AI (EAI).
★ 508GLM-130B. GLM-130B: An Open Bilingual Pre-Trained Model (ICLR 2023)
★ 7.7kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kalpaca-lora. Instruct-tune LLaMA on consumer hardware
★ 19kLLM-As-Chatbot. LLM as a Chatbot Service
★ 3.3kTaskMatrix. Python
★ 34kVideo-P2P. Video-P2P: Video Editing with Cross-attention Control
★ 431make-a-video-pytorch. Implementation of Make-A-Video, new SOTA text to video generator from Meta AI, in Pytorch
★ 2kPrompt-Engineering-Guide. 🐙 Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.
★ 77kevals. Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
★ 19k