This is your work, valued
TRACE. [ICLR 2025] TRACE: Temporal Grounding Video LLM via Casual Event Modeling
★ 157VTG-LLM. [AAAI 2025] VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
★ 13025-fall-recruit. Try to summarize some information of 25 fall recruit
★ 5OpenCut. The open-source CapCut alternative
★ 80kskills. Skills for Design Engineers.
★ 23kAwesome-Credit-Assignment-in-LLM-RL. Curated papers, taxonomy, benchmarks, and decision guides for credit assignment in reasoning and agentic LLM reinforcement learning.
★ 121Hy3. Hy3 (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
★ 556DeepSpec. DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
★ 6.8kOpenMontage. World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
★ 44kGLM-5. GLM-5: From Vibe Coding to Agentic Engineering
★ 6.8kCodeIsland. Real-time AI coding agent status panel in your MacBook notch — live status, approvals & replies for 13 AI tools, with iPhone & Apple Watch companions
★ 2.2kAwesome-One-Step-Generation. A curated list of papers, code, and resources for one-step diffusion models that turn noise into high-quality samples in a single neural network forward pass.
★ 172figures4papers. My Python scripts to make high-quality figures for publications in top AI conferences and journals.
★ 2.9kimage-to-editable-ppt-skill. Codex skill for converting slide images, PDFs, and image-based PPTX files into editable PowerPoint decks.
★ 1.7kvime. An LLM post-training framework with vLLM for RL Scaling
★ 396skills. Original and practical skills for AI builders.
★ 425academic-research-skills. Academic Research Skills for Claude Code: research → write → review → revise → finalize
★ 40kLance. A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
★ 1.3kLLaDA2.0-Uni. LLaDA2.0-Uni: Understanding and Generation the World.
★ 771ERNIE-Image. ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT), with only 8B DiT parameters, it reaches state-of-the-art performance among open-weight text-to-image models.
★ 494Awesome-Multimodal-Modeling. Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 508ml-intern. 🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
★ 11klearn-harness-engineering. Harness engineering beginner tutorial, from 0 to 1
★ 11kTIR-Bench. [ECCV 2026] Official implementation of "TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning"
★ 25verl-agent. verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
★ 2.2kawesome-design-md. A collection of DESIGN.md files analysis by popular brand design systems. Drop one into your project and let coding agents generate a matching UI.
★ 106kawesome-harness-engineering. 🛠️ Awesome tools & guides for harness engineering.
★ 3.7kT3RL. Python
★ 48APEX. [Preprint] Self-Adversarial One Step Generation via Condition Shifting
★ 56andrej-karpathy-skills. A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
★ 198khermes-agent. The agent that grows with you
★ 223kJoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2kGen-Searcher. Gen-Searcher: Reinforcing Agentic Search for Image Generation
★ 377learn-coding-agent. Research on Coding Agents
★ 12kskills. Based on The Minimalist Entrepreneur by Sahil Lavingia
★ 9.7kskills. C#
★ 13kAgent-Reach. Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
★ 63kCapybara. Python
★ 203PersonaVLM. [CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs
★ 113WeEdit. A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
★ 20danghuangshang. AI 朝廷搭建完整教程 - 从零基础到进阶
★ 2.7kedict. 🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
★ 16kdeer-flow. An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
★ 78kG-OPD. Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"
★ 277autoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kOpenClaw-RL. OpenClaw-RL: Train any agent simply by talking
★ 5.6kpua. 你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
★ 19kSkillRL. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
★ 916SteptronOss. A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.
★ 579SupervisorAgent. [ICLR 2026] Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
★ 32dllm. dLLM: Simple Diffusion Language Modeling
★ 2.7kJavisDiT. [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
★ 377solaris. The first multiplayer video world model in Minecraft
★ 221openclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kQwen3.6. Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
★ 3.7kFireRed-Image-Edit. FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity consistency, and seamless multi-element fusion.
★ 1.3kQwen-Image-Layered. Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
★ 2kBagel. Open-source unified multimodal model
★ 6.1kNarratoAI. 利用 AI 大模型,一键解说并剪辑视频
★ 11kEdit-Banana. Edit Banana: A framework for converting statistical formats into editable.
★ 5.4kFastGen. NVIDIA FastGen: Fast Generation from Diffusion Models
★ 903DiffSynth-Engine. Python
★ 427Awesome-Image-Editing. A Survey of Image Editing
★ 469VideoX-Fun. 📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
★ 2.2kFlowEdit. Official implementation of the paper: "FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models"
★ 1kSDPO. Reinforcement Learning via Self-Distillation (SDPO)
★ 1kLanPaintBench. The benchmark code for LanPaint
★ 10LanPaint. High quality training free inpaint for every stable diffusion model. Supports ComfyUI
★ 1.3kAction100M. A Large-scale Video Action Dataset
★ 483GDPO. Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
★ 494xhs_ai_publisher. AI-powered Xiaohongshu/Rednote content creation and publishing tool with PyQt desktop UI, FastAPI service, login-state reuse, preview publish, and automated browser workflows.
★ 2kflux2. Official inference repo for FLUX.2 models
★ 2.6kDataFlow. Easy Data Preparation with latest LLMs-based Operators and Pipelines.
★ 7.1kpico-banana-400k. Python
★ 1.8kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34kWan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kSpotEdit. SpotEdit:Selective Region Editing in Diffusion Transformers
★ 196G2. 📊 The concise and progressive visualization grammar.
★ 13kInfographic. 🦋 An Infographic Generation and Rendering Framework, bring words to life with AI!
★ 5.7kdeepcompressor. Model Compression Toolbox for Large Language Models and Diffusion Models
★ 796awesome-discrete-diffusion-models. A curated list for awesome discrete diffusion models resources.
★ 572EditScore. [ICLR 2026] EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
★ 256LightX2V. Lightweight Image Video Action Generation Inference Framework
★ 2.5kagentsight. system-level profiler/monitor and skills for Agents in eBPF or Mac
★ 550dinov3. Reference PyTorch implementation and models for DINOv3
★ 11kSimpleTuner. A general fine-tuning kit geared toward image/video/audio diffusion models.
★ 2.9kMini-Agent. A minimal yet professional single agent demo project that showcases the core execution pipeline and production-grade features of agents.
★ 2.9kTwinFlow. [ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
★ 537TwinFlow.
★ 4dmdr. [ECCV 2026] Official Code of "Distribution Matching Distillation Meets Reinforcement Learning"
★ 288LongCat-Image. Python
★ 715cache-dit. A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
★ 1.2kvllm-omni. A framework for efficient model inference with omni-modality models
★ 5.7knunchaku. [ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
★ 3.9kinference. Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
★ 9.5kComfyUI. The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
★ 123kFastDM. Diffusion Model Inference Projects
★ 58dFactory. Easy and Efficient dLLM Fine-Tuning
★ 261Z-Image. Python
★ 12kyolov8-face-landmarks-opencv-dnn. 使用OpenCV部署yolov8检测人脸和关键点以及人脸质量评价,包含C++和Python两个版本的程序,只依赖opencv库就可以运行,彻底摆脱对任何深度学习框架的依赖。
★ 298DiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kAwesome-Diffusion-Model-Based-Image-Editing-Methods. Diffusion Model-Based Image Editing: A Survey (TPAMI 2025)
★ 712Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kControlNet. Let us control diffusion models!
★ 34kinpaint-anything. Inpaint Anything performs stable diffusion inpainting on a browser UI using masks from Segment Anything.
★ 274IOPaint. Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.
★ 23kvideo-subtitle-remover. 基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.
★ 12kDeblurGANv2. [ICCV 2019] "DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better" by Orest Kupyn, Tetiana Martyniuk, Junru Wu, Zhangyang Wang
★ 1.2kAwesome-Deblurring. A curated list of resources for Image and Video Deblurring
★ 2.9kHYPIR. Official implementation of HYPIR: Harnessing Diffusion-Yielded Score Priors for Image Restoration (SIGGRAPH 2025)
★ 1.2kunderstanding-attention-sinks. Understanding the emergence and role of attention sinks in LLMs. Benchmarking KV-caching attentinon implementations.
★ 2RCGM. [ICLR 2026] Any-step Generation via N-th Order Recursive Consistent Velocity Field Estimation
★ 41Pixel-Reasoner. Pixel-Level Reasoning Model trained with RL [NeuIPS25]
★ 301RecIS. A unified architecture deep learning framework designed specifically for ultra-large-scale sparse models.
★ 352gated_attention. The official implementation for [NeurIPS2025 Oral] Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
★ 973TriSense. [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
★ 27Implicit-Video-Reasoning. Python
★ 7reconstruction-alignment. [ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Learning.
★ 411MemVR. [ICML 2025] Official implementation of paper 'Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models'.
★ 171Real-ESRGAN. Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
★ 36kCodeFormer. [NeurIPS 2022] Towards Robust Blind Face Restoration with Codebook Lookup Transformer
★ 18kVision-SR1. Reinforcement Learning of Vision Language Models with Self Visual Perception Reward
★ 177Whisper-Finetune. Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
★ 1.2kwhisper. Robust Speech Recognition via Large-Scale Weak Supervision
★ 106kBeing-VL-0.5. Being-VL-0.5: Unified Multimodal Understanding via Byte-Pair Visual Encoding (ICCV 2025, Highlight)
★ 54audio-ai-hub. The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
★ 949agentscope. Build and run agents you can see, understand and trust.
★ 28kLightX2V-Qwen-Image-Lightning. Qwen-Image-Lightning: Speed up Qwen-Image model with distillation
★ 1.3kYouku-mPLUG. Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
★ 307Step-Audio2. Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
★ 1.5kQwen2-Audio. The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 2.1kQwen-Audio. The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 1.9ksmolagents. 🤗 smolagents: a barebones library for agents that think in code.
★ 29kVeOmni. VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
★ 2.1kByteVideoLLM. [ICCV 2025] Dynamic-VLM
★ 28Efficient_Attention_Survey. A Survey of Efficient Attention Methods: Hardware-efficient, Sparse, Compact, and Linear Attention
★ 305G2RPO-A. [ACL 2026] G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
★ 16Intern-S1. A Scientific Multimodal Foundation Model
★ 842verifiers. Our library for RL environments + evals
★ 4.4kEntropy-Mechanism-of-RL. The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
★ 446ARC-Hunyuan-Video-7B. Structured Video Comprehension of Real-World Shorts
★ 240video-SALMONN-2. video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is developed by the Department of Electronic Engineering at Tsinghua University and ByteDance.
★ 204LUFFY. Official Repository of "Learning to Reason under Off-Policy Guidance"
★ 461Thyme. ✨✨ [ICLR 2026] Think Beyond Images
★ 584R1-Searcher. R1-searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
★ 720m3-agent. Python
★ 1.4kROLL. An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.3kSearch-R1. Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL
★ 5.2kMooncake. Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.1kdlms-are-super-data-learners. The official github repo for "Diffusion Language Models are Super Data Learners".
★ 227aibrix. Cost-efficient and pluggable Infrastructure components for GenAI inference
★ 5ktorch_memory_saver. Allow torch tensor memory to be released and resumed later
★ 262slime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kMagiAttention. A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
★ 894qwen-code. An open-source AI coding agent that lives in your terminal.
★ 26kSageAttention. [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
★ 3.5ksiiRL. siiRL: Shanghai Innovation Institute RL Framework for Advanced LLMs and Multi-Agent Systems
★ 368PAPO. Official repo for "PAPO: Perception-Aware Policy Optimization for Multimodal Reasoning"
★ 153trae-agent. Trae Agent is an LLM-based agent for general purpose software engineering tasks.
★ 12kTrue-Story-of-Pangu. 诺亚盘古大模型研发背后的真正的心酸与黑暗的故事。
★ 12kGLM-V. GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2.4kKeye. Python
★ 809OmniBench. [ICML 2025 Oral] This is the official repository of the paper "What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities"
★ 23ILLUME_plus. [CVPR2025] Official Implementation of ILLUME+
★ 126gemini-cli. An open-source AI agent that brings the power of Gemini directly into your terminal.
★ 106kAwesome-ML-SYS-Tutorial. My learning notes for ML SYS.
★ 6.8kMASArena. A comprehensive framework for benchmarking single and multi-agent systems across a wide range of tasks—evaluating performance, accuracy, and efficiency with built-in visualization and tool integration.
★ 38DLLM-Survey. [Arxiv] Discrete Diffusion in Large Language and Multimodal Models: A Survey
★ 387OmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kParScale. Parallel Scaling Law for Language Model — Beyond Parameter and Inference Time Scaling
★ 480DeepEyes. Python
★ 1.3kMultiverse. Python
★ 119agent-zero. Agent Zero AI framework
★ 19kflow_matching. A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
★ 4.7kTokLIP. TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
★ 236Spurious_Rewards. Python
★ 361Paper2Poster. [NeurIPS 2025] Open-source Multi-agent Poster Generation from Papers
★ 3.9kMARTI. A Framework for LLM-based Multi-Agent Reinforced Training and Inference
★ 540discrete-diffusion-papers. A collection of papers on discrete diffusion models
★ 164LLaVA-ST. [CVPR 2025] LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
★ 84OpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58kMMR1. [CVPR 2026] MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
★ 217flux. A fast communication-overlapping library for tensor/expert parallelism on GPUs.
★ 1.4kowl. 🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
★ 20kAwesome-LLM-Post-training. Awesome Reasoning LLM Tutorial/Survey/Guide
★ 2.5kAwesome-Multimodal-Large-Language-Models. 🔥Awesome Multimodal Large Language Models Paper List
★ 154VisualThinker-R1-Zero. Explore the Multimodal “Aha Moment” on 2B Model
★ 624R1-Onevision. R1-onevision, a visual language model capable of deep CoT reasoning.
★ 581Muon. Muon is an optimizer for hidden layers in neural networks
★ 2.8kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kllm-reasoners. A library for advanced large language model reasoning
★ 2.3kDeepGEMM. DeepGEMM: clean and efficient BLAS kernel library on GPU
★ 7.6kFlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13kEasyR1. EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
★ 5.1kNamo-R1. A CPU Realtime VLM in 500M. Surpassed Moondream2 and SmolVLM. Training from scratch with ease.
★ 256VLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kMoE-plus-plus. [ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
★ 270open-infra-index. Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
★ 8kOpen-Reasoner-Zero. Official Repo for Open-Reasoner-Zero
★ 2.1kLIMR. Python
★ 221trl. Train transformer language models with reinforcement learning.
★ 19klmm-r1. Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.
★ 848verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23k