This is your work, valued
Working on computer vision, machine learning, deep learning, etc.
ShuffleNet_V2_pytorch_caffe. ShuffleNet-V2 for both PyTorch and Caffe.
★ 504SqueezeNet_v1.2. Top-1 Acc=61.0% on ImageNet, without any sacrificing compared with SqueezeNet v1.1.
★ 22dlfs. Deep Learning from Scratch
★ 2Megatron-Bridge. Training library for Megatron-based models with bidirectional Hugging Face conversion capability
★ 838xueersi-xiaomiao. 学而思ESP32掌机小喵开发
★ 116crosspoint-reader. Firmware for the Xteink X3 and X4 e-readers
★ 6.7kMiMo-Code. MiMo Code: Where Models and Agents Co-Evolve
★ 13kDataDesigner. 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
★ 2.1kQKeyMapper. [按键映射工具] QKeyMapper,Qt开发Win10&Win11可用,不修改注册表、不需重新启动系统,可立即生效和停止。支持游戏手柄映射到键鼠,手柄摇杆控制鼠标移动,键鼠映射到虚拟游戏手柄,鼠标控制虚拟手柄移动摇杆等功能。
★ 693skills. Skills for Real Engineers. Straight from my .agents directory.
★ 198kSpeech. A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
★ 18kbun. Incredibly fast JavaScript runtime, bundler, test runner, and package manager – all in one
★ 95kopenhuman. Your Personal AI super intelligence. A brain that builds a local-first memory of your life, a fantastic orchestrator of agent fleets and workflows, and a deep researcher.
★ 36kawesome. 😎 Awesome lists about all kinds of interesting topics
★ 491koptimizers. For optimization algorithm research and development.
★ 579heretic. Fully automatic censorship removal for language models
★ 27kml-intern. 🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
★ 11ktushu.
★ 21ECC. The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
★ 237kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kunsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek, GLM and other models.
★ 69knanoclaw. A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps,, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK
★ 30koh-my-claudecode. Teams-first Multi-agent orchestration for Claude Code
★ 38kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kclaude-code. Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
★ 140klangextract. A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
★ 38kimeanflow. Official Implementation of iMF https://arxiv.org/abs/2512.02012
★ 335vibetensor. Our first fully AI generated deep learning system
★ 635TurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kmini-sglang. A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
★ 4.7kInfiniteFusionPreloaded. Manages Configs for the Launcher and Releases
★ 34infinitefusion-e18. A heavily modified RPG Maker XP game project that makes the game play like a Pokémon game. Not a full project in itself; this repo is to be added into an existing RMXP game project.
★ 177creatiposter. This repository open-sources CreatiPoster, an AI-driven graphic design generation system for multi-layer and editable compositions with strong visual appeal.
★ 105DeepSeek-OCR. Contexts Optical Compression
★ 24knanochat. The best ChatGPT that $100 can buy.
★ 57kHunyuanImage-3.0. HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
★ 3.2kMegatron-LM. Ongoing research training transformer models at scale
★ 17kcopyparty. Portable file server with accelerated resumable uploads, dedup, WebDAV, SFTP, FTP, TFTP, zeroconf, media indexer, thumbnails++ all in one file
★ 46kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kUNO. [ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
★ 1.4kKimi-K2. Kimi K2 is the large language model series developed by Moonshot AI team
★ 11kOpenCut. The open-source CapCut alternative
★ 80kMuon. Muon is an optimizer for hidden layers in neural networks
★ 2.8kMonkeyOCR. A lightweight LMM-based Document Parsing Model
★ 6.6kOpenS2V-Nexus. [NeurIPS 2025 D&B🔥] OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
★ 225Megakernels. Kernels, of the mega variety :)
★ 788HunyuanVideo-I2V. HunyuanVideo-I2V: A Customizable Image-to-Video Model based on HunyuanVideo
★ 1.8kNCFM. Official PyTorch implementation of the paper "Dataset Distillation with Neural Characteristic Function: A Minmax Perspective" (NCFM) in CVPR 2025 (Full Score, Highlight).
★ 413Open-Qwen2VL. [COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
★ 315RedPajama-Data. The RedPajama-Data repository contains code for preparing large datasets for training large language models.
★ 5kClip-Anything. Clip any moment from any video with prompts
★ 284DualPipe. A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.
★ 3kFlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13kTMPI. Code for ICCV 2023 paper on tiled multiplane images for single-view 3D photography.
★ 62MultimodalOCR. On the Hidden Mystery of OCR in Large Multimodal Models (OCRBench)
★ 874open-r1. Fully open reproduction of DeepSeek-R1
★ 26kR1-V. Witness the aha moment of VLM with less than $3.
★ 4.1kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kminimind. 🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
★ 54kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kmarkitdown. Python tool for converting files and office documents to Markdown.
★ 171kRT-DETR. [CVPR 2024] Official RT-DETR (RTDETR paddle pytorch), Real-Time DEtection TRansformer, DETRs Beat YOLOs on Real-time Object Detection. 🔥 🔥 🔥
★ 5.4kclay. High performance UI layout library in C.
★ 18kOminiControl. [ICCV 2025 Highlight] OminiControl: Minimal and Universal Control for Diffusion Transformer
★ 1.9ksamurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
★ 7.1kechomimic_v2. [CVPR 2025] EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
★ 4.6kopen-oasis. Inference script for Oasis 500M
★ 2.1kcuda-python. CUDA Python: Performance meets Productivity
★ 3.3kOmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3kthorium-reader. A cross platform desktop reading app, based on the Readium Desktop toolkit
★ 2.8kMoVQGAN. MoVQGAN - model for the image encoding and reconstruction
★ 266hallo2. [ICLR 2025] Hallo2: Long-Duration and High-Resolution Audio-driven Portrait Image Animation
★ 3.7kmochi. The best OSS video generation models, created by Genmo
★ 3.7klingua. Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
★ 4.8ksd3.5. Python
★ 1.5kLanguageBreak. A kindle <=5.16.2.1.1 jailbreak
★ 1.2kStableDiffusionOnDevice. 本项目是一个通过文字生成图片的项目,基于开源模型Stable Diffusion V1.5生成可以在手机的CPU和NPU上运行的模型,包括其配套的模型运行框架。
★ 246ms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kesp-idf. Espressif IoT Development Framework. Official development framework for Espressif SoCs.
★ 19kLiClock. 一种兼具易用性与扩展性的多功能墨水屏天气时钟
★ 159Wav2Lip. This repository contains the codes of "A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild", published at ACM Multimedia 2020. For HD commercial model, please try out Sync Labs
★ 13kV-Express. V-Express aims to generate a talking head video under the control of a reference image, an audio, and a sequence of V-Kps images.
★ 2.4kAwesome-Talking-Head-Synthesis. 💬 An extensive collection of exceptional resources dedicated to the captivating world of talking face synthesis! ⭐ If you find this repo useful, please give it a star! 🤩
★ 1.5kultralytics. Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
★ 60kaccelerate. 🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
★ 9.8kpretzelai. The modern replacement for Jupyter Notebooks
★ 2.2kLiveTalking. Real time interactive streaming digital human
★ 8.6kSyncTalk. [CVPR 2024] This is the official source for our paper "SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis"
★ 1.6kSadTalker. [CVPR 2023] SadTalker:Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation
★ 14kGOT-OCR2.0. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
★ 8.2k3D-Photography-with-Image-Inpainting-and-Depth-Estimation. Sensing Depth from 2D Images and Inpainting Background behind the Foreground objects to create 3D Photos with Parallax Animation.
★ 28Depth-Anything-V2. [NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
★ 8.6kDepth-Anything. [CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation
★ 8.2ksapiens. High-resolution models for human tasks.
★ 5.4kVEnhancer. Official codes of VEnhancer: Generative Space-Time Enhancement for Video Generation
★ 577MPS. Python
★ 206MagicClothing. Official implementation of Magic Clothing: Controllable Garment-Driven Image Synthesis
★ 1.5kControlNeXt. Controllable video and image Generation, SVD, Animate Anyone, ControlNet, ControlNeXt, LoRA
★ 1.6kMonkey. Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)
★ 2kmar. PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
★ 1.9kCogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kCatVTON. [ICLR 2025] CatVTON is a simple and efficient virtual try-on diffusion model with 1) Lightweight Network (899.06M parameters totally), 2) Parameter-Efficient Training (49.57M parameters trainable) and 3) Simplified Inference (< 8G VRAM for 1024X768 resolution).
★ 1.8kIMAGDressing. [AAAI 2025]👔IMAGDressing👔: Interactive Modular Apparel Generation for Virtual Dressing. It enables customizable human image generation with flexible garment, pose, and scene control, ensuring high fidelity and garment consistency for virtual dressing.
★ 1.3ktrl. Train transformer language models with reinforcement learning.
★ 19kflux. Official inference repo for FLUX.1 models
★ 26kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88ksimple-computer. the scott CPU from "But How Do It Know?" by J. Clark Scott
★ 2ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kdoc2graph. Doc2Graph transforms documents into graphs and exploit a GNN to solve several tasks.
★ 139SimPO. [NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
★ 956graphrag. A modular graph-based Retrieval-Augmented Generation (RAG) system
★ 35kclaude-engineer. Claude Engineer is an interactive command-line interface (CLI) that leverages the power of Anthropic's Claude-3.5-Sonnet model to assist with software development tasks.This framework enables Claude to generate and manage its own tools, continuously expanding its capabilities through conversation. Available both as a CLI and a modern web interface
★ 11kMambaVision. [CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
★ 2.2kAutoShot. AutoShot: A Short Video Dataset and State-of-the-Art Shot Boundary Detection - CVPR NAS 2023
★ 249Kolors. Kolors Team
★ 4.6kcambrian. Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2kLivePortrait. Bring portraits to life!
★ 19kLLM101n. LLM101n: Let's build a Storyteller
★ 38kchameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kseed. Seed Machine Translation Data
★ 34mesop. Rapidly build AI apps in Python
★ 6.6kOmost. Your image is almost there!
★ 7.6kschedule_free. Schedule-Free Optimization in PyTorch
★ 2.3kgeektime-books. :books: 极客时间电子书
★ 13kMoMA. MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
★ 234Lumina-T2X. Lumina-T2X is a unified framework for Text to Any Modality Generation
★ 2.2ktext-generation-inference. Large Language Model Text Generation Inference
★ 11kDeepSpeed-MII. MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
★ 2.1kHunyuanDiT. Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
★ 4.3kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kIC-Light. More relighting!
★ 8.5ksglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kMS-DOS. The original sources of MS-DOS 1.25, 2.0, and 4.0 for reference purposes
★ 32kOpen-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12kBrushNet. [ECCV 2024] The official implementation of paper "BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion"
★ 1.7kdatasets. 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
★ 22kComfyUI_examples. Examples of ComfyUI workflows
★ 4.4klibzmq. ZeroMQ core engine in C++, implements ZMTP/3.1
★ 11knanomsg. nanomsg library
★ 6.3knng. nanomsg-next-generation -- light-weight brokerless messaging
★ 4.6kPixArt-alpha. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
★ 3.3kPixArt-sigma. PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
★ 1.9kmmfashion. Open-source toolbox for visual fashion analysis based on PyTorch
★ 1.4kmultimodal-garment-designer. This is the official repository for the paper "Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing". ICCV 2023
★ 445EfficientSAM. EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything
★ 2.5kTripoSR. TripoSR: Fast 3D Object Reconstruction from a Single Image
★ 6.8kgpt_academic. 为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
★ 71kTCD. Official Repository of the paper "Trajectory Consistency Distillation"
★ 360LayerDiffuse. Transparent Image Layer Diffusion using Latent Transparency
★ 2.2kInstaFlow. :zap: InstaFlow! One-Step Stable Diffusion with Rectified Flow (ICLR 2024)
★ 1.4kconditional-flow-matching. TorchCFM: a Conditional Flow Matching library
★ 2.6kcoyo-dataset. COYO-700M: Large-scale Image-Text Pair Dataset
★ 1.3khighway. Performance-portable, length-agnostic SIMD with runtime dispatch
★ 5.7kgemma.cpp. lightweight, standalone C++ inference engine for Google's Gemma models.
★ 7kOOTDiffusion. [AAAI 2025] Official implementation of "OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on"
★ 6.6kprofessional-programming. A collection of learning resources for curious software engineers
★ 51kLWM. Large World Model -- Modeling Text and Video with Millions Context
★ 7.4kmagika. Fast and accurate AI powered file content types detection
★ 18kminbpe. Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
★ 11kmicrograd. A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
★ 17kStableCascade. Official Code for Stable Cascade
★ 6.5kdonut. Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
★ 6.9kOLMo. Modeling, training, eval, and inference code for OLMo
★ 6.6ktampermonkey. Tampermonkey is the most popular userscript manager, with over 10 million users. It's available for Chrome, Microsoft Edge, Safari, Opera Next, and Firefox.
★ 5.6kexcelCPU. 16-bit CPU for Excel, and related files
★ 4.7kQwen-VL. The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6.7kAnyText. Official implementation code of the paper <AnyText: Multilingual Visual Text Generation And Editing>
★ 4.9kLLMs-from-scratch. Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
★ 100ksearch_with_lepton. Building a quick conversation-based search demo with Lepton AI.
★ 8.1kVim. [ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
★ 3.9kInpaint-Anything. Inpaint anything using Segment Anything and inpainting models.
★ 7.7kmanga-ocr. Optical character recognition for Japanese text, with the main focus being Japanese manga
★ 2.7kProPainter. [ICCV 2023] ProPainter: Improving Propagation and Transformer for Video Inpainting
★ 6.8kaudio2photoreal. Code and dataset for photorealistic Codec Avatars driven from audio
★ 2.9kgl-transitions. The open collection of GL Transitions
★ 2.1kFooocus. Focus on prompting and generating
★ 52kOpenVoice. Instant voice cloning by MIT and MyShell. Audio foundation model.
★ 37kdiff2lip. Python
★ 379StreamDiffusion. StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive Generation
★ 11kOPUS. The Open Parallel Corpus
★ 88kornia. 🐍 Geometric Computer Vision Library for Spatial AI
★ 11ksynthtiger. Official Implementation of SynthTIGER (Synthetic Text Image Generator), ICDAR 2021
★ 579gpt-fast. Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.
★ 6.2kDALI. A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
★ 5.7kMixtralKit. A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
★ 770LLaMA-VID. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
★ 861magic-animate. [CVPR 2024] Official repository for "MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model"
★ 11kAnimateAnyone. Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
★ 15kIPTV. M3U Playlist for free TV channels
★ 20kimage-compare-viewer. Compare before and after images, for grading and other retouching for instance. Vanilla JS, zero dependencies.
★ 591build-your-own-x. Master programming by recreating your favorite technologies from scratch.
★ 533kHAT. CVPR2023 - Activating More Pixels in Image Super-Resolution Transformer TPAMI - HAT: Hybrid Attention Transformer for Image Restoration
★ 1.6kStableSR. [IJCV2024] Exploiting Diffusion Prior for Real-World Image Super-Resolution
★ 2.7kfacefusion. Industry leading face manipulation platform
★ 29kT-Rex. [ECCV2024] API code for T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
★ 2.7kAI-For-Beginners. 12 Weeks, 24 Lessons, AI for All!
★ 55kData-Science-For-Beginners. 10 Weeks, 20 Lessons, Data Science for All!
★ 36kgenerative-ai-for-beginners. 21 Lessons, Get Started Building with Generative AI
★ 114khcaptcha-challenger. 🥂 Gracefully face hCaptcha challenge with multimodal large language model.
★ 2.4kCoCa-pytorch. Implementation of CoCa, Contrastive Captioners are Image-Text Foundation Models, in Pytorch
★ 1.2krobust-dynrf. An algorithm for reconstructing the radiance field of a dynamic scene from a casually-captured video.
★ 243godot. Godot Engine – Multi-platform 2D and 3D game engine
★ 115kaudio-fingerprint-identifying-python. The Shazam-similar app, that identify the song using audio fingerprints & spectrum analysis and Fast Fourier transform
★ 387