This is your work, valued
Researcher at Tongyi Lab.
siamfc-pytorch. A clean PyTorch implementation of SiamFC tracking/training, evaluated on 7 datasets.
★ 701GlobalTrack. Official PyTorch implementation of "GlobalTrack: A Simple and Strong Baseline for Long-term Tracking" @ AAAI2020.
★ 261siamrpn-pytorch. A clean PyTorch implementation of SiamRPN tracker, evaluated on 7 datasets.
★ 199mot-papers. A collection of Multiple Object Tracking (MOT) papers in recent years, with notes.
★ 195open-vot. Open source visual object tracking library in python
★ 84pay-attention-pytorch. PyTorch implementation of the ICLR 2018 paper Learning to Pay Attention.
★ 22video-detection-benchmark. Video object detection benchmark.
★ 19cortex. A minimal engine and a large benchmark for deep learning algorithms. Built upon PyTorch.
★ 8tensor-pooling. Tensor pooling for online visual tracking, in MATLAB (Oral and BEST paper candidate of ICME).
★ 4mdnet-pytorch. PyTorch implementation of the MDNet tracker.
★ 1part-tracking. Visual tracking by sampling in part space, in MATLAB.
★ 1Awesome-Video-World-Models-with-AR-Diffusion. A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrastructure, Aimed at Serving as a Comprehensive Resource for Researchers, Practitioners, and Enthusiasts.
★ 682MSA. Memory Sparse Attention - A scalable, end-to-end trainable latent-memory framework for 100M-token contexts.
★ 3.5kkraken. Triton-based Symmetric Memory operators and examples
★ 109Attention-Residuals.
★ 3.4kairi. 💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime voice chat, Minecraft, Factorio playing. Web / macOS / Windows supported.
★ 46kNitroGen. A Foundation Model for Generalist Gaming Agents
★ 2.1kWAVE. ICLR 2026 Oral: WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
★ 42open-p2p. Official Repo for paper: Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing
★ 172STEM. Python
★ 66pplx-kernels. Perplexity GPU Kernels
★ 595pplx-garden. Perplexity open source garden for inference technology
★ 611ReCamMaster. [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
★ 1.8kDeepEP. DeepEP: an efficient expert-parallel communication library
★ 9.9kcubvh. Mesh tools.
★ 298warp. A Python framework for GPU-accelerated simulation, robotics, and machine learning.
★ 6.9kGenie-Envisioner-V1. Python
★ 565awesome-embodied-vla-va-vln. A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
★ 3.4kvjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.
★ 4.4kSelf-Forcing-Endless. Make self forcing endless. Add cache purging. Add prompt controllability.
★ 71SkyReels-V2. SkyReels-V2: Infinite-length Film Generative model
★ 7.3kMAGI-1. MAGI-1: Autoregressive Video Generation at Scale
★ 3.7kharmony. Renderer for the harmony response format to be used with gpt-oss
★ 4.5kgpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20kFastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kGPT-SoVITS. 1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 60kms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kmodded-nanogpt. NanoGPT (124M) in 90 seconds
★ 5.6kfms-fsdp. 🚀 Efficiently (pre)training foundation models with native PyTorch features, including FSDP for training and SDPA implementation of Flash attention v2.
★ 288OmniAvatar. Python
★ 1.9kYUME. The official code of Yume
★ 679DeepGEMM. DeepGEMM: clean and efficient BLAS kernel library on GPU
★ 7.6khiggs-audio. Text-audio foundation model from Boson AI
★ 8.3kmixture_of_recursions. Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation (NeurIPS 2025)
★ 579index-tts. An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
★ 22kEgo4d. Ego4d dataset repository. Download the dataset, visualize, extract features & example usage of the dataset
★ 627MoBA. MoBA: Mixture of Block Attention for Long-Context LLMs
★ 2.2knative-sparse-attention. 🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
★ 1kSageAttention. [ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
★ 3.5kflashinfer. FlashInfer: Kernel Library for LLM Serving
★ 6.1kcarla. Open-source simulator for autonomous driving research.
★ 14kMOSS-TTSD. MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references.
★ 1.4kMagiAttention. A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
★ 894radial-attention. [NeurIPS 2025] Radial Attention: O(nlogn) Sparse Attention with Energy Decay for Long Video Generation
★ 605TalkingMachines. TalkingMachines
★ 178CoGenAV. Python
★ 64REPA-E. [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
★ 512BigVGAN. Official PyTorch implementation of BigVGAN (ICLR 2023)
★ 1.2kDolphin. The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
★ 9kencodec. State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
★ 4kdescript-audio-codec. State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.
★ 1.8kBagel. Open-source unified multimodal model
★ 6.1ktailwind-nextjs-starter-blog. This is a Next.js, Tailwind CSS blogging starter template. Comes out of the box configured with the latest technologies to make technical writing a breeze. Easily configurable and customizable. Perfect as a replacement to existing Jekyll and Hugo individual blogs.
★ 11ksmolvlm-realtime-webcam. Real-time webcam demo with SmolVLM and llama.cpp server
★ 5.6kVITA-Audio. ✨✨[NeurIPS 2025] VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
★ 683pyannote-audio. Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
★ 10kKimi-Audio. Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
★ 4.7kdia. A TTS model capable of generating ultra-realistic dialogue in one pass.
★ 19kLiquid. (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
★ 642FramePack. Lets make video diffusion practical!
★ 17kmetamorph. Code for MetaMorph Multimodal Understanding and Generation via Instruction Tuning
★ 235verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kSMPLest-X. [TPAMI 2025] Official Code for "SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation"
★ 306smplify-x. Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
★ 2.2kWonderWorld. Code release for https://kovenyu.com/WonderWorld/
★ 739Awesome-Talking-Head-Synthesis. 💬 An extensive collection of exceptional resources dedicated to the captivating world of talking face synthesis! ⭐ If you find this repo useful, please give it a star! 🤩
★ 1.5kMegaTTS3. Python
★ 6.1kgemma_pytorch. The official PyTorch implementation of Google's Gemma models
★ 5.7kinferno. 🔥🔥🔥 Set the world of 3D faces on fire with INFERNO 🔥🔥🔥
★ 320emoca. Official repository accompanying a CVPR 2022 paper EMOCA: Emotion Driven Monocular Face Capture And Animation. EMOCA takes a single image of a face as input and produces a 3D reconstruction. EMOCA sets the new standard on reconstructing highly emotional images in-the-wild
★ 852Arinar. Python
★ 43fsdp_qlora. Training LLMs with QLoRA + FSDP
★ 1.6kDyadic-Interaction-Modeling. [ECCV 2024] Dyadic Interaction Modeling for Social Behavior Generation
★ 65csm. A Conversational Speech Generation Model
★ 15kSpatialLM. [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
★ 4.7kAwesome-Human-Motion-Video-Generation. 【Accepted by TPAMI】Human Motion Video Generation: A Survey (https://ieeexplore.ieee.org/document/11106267)
★ 340AudioX. [ICLR 2026] Repository of AudioX
★ 1.5kSonic. Official implementation of "Sonic: Shifting Focus to Global Audio Perception in Portrait Animation"
★ 3.3kRectified-Diffusion. [ICLR 2025] Rectified Diffusion: Straightness Is Not Your Need
★ 250bd3lms. [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
★ 1kdiffused-heads. Official repository for Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
★ 489Synchformer. Source code for "Synchformer: Efficient Synchronization from Sparse Cues" (ICASSP 2024)
★ 130LightningDiT. [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
★ 1.5kSALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kfractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2kopen-oasis. Inference script for Oasis 500M
★ 2.1kPlayable-Game-Generation. An open-source lightweight game generation paradigm. It includes everything from data processing to model architecture design and playability-based evaluation methods. The game runs at 20 FPS on a single consumer-grade graphics card (RTX-2060) while maintaining high playability.
★ 120diffusion-forcing. code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"
★ 1.3kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kdiffusion-forcing-transformer. [ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
★ 705Gymnasium. A standard API for single-agent reinforcement learning environments, with popular reference environments and related utilities (formerly Gym)
★ 12kgym. A toolkit for developing and comparing reinforcement learning algorithms.
★ 37kVideo-Pre-Training. Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos
★ 1.7kmini-omni. open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
★ 3.6kwhisper. Robust Speech Recognition via Large-Scale Weak Supervision
★ 106kmoshi. Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
★ 11kAsk-Anything. [CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
★ 3.3kswarm. Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
★ 22kopenai-realtime-agents. This is a simple demonstration of more advanced, agentic patterns built on top of the Realtime API.
★ 6.9klexical. Lexical is an extensible text editor framework that provides excellent reliability, accessibility and performance.
★ 24k