This is your work, valued
AudioX. [ICLR 2026] Repository of AudioX
★ 1.5kAudio-Omni. [SIGGRAPH 2026] Repository of Audio-Omni
★ 399VidMuse. [CVPR 2025] Repository of VidMuse
★ 140ZeyueT.github.io. HTML
★ 3T2A-bench. [ICLR 2026] A two-stage evaluation benchmark for Text-to-Audio generation, assessing category, count, ordering, and timestamp accuracy with Gemini 2.5 Pro.
★ 2OPSD-V. On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
★ 458Awesome-MLLM-Reasoning-Collection. A collection of multimodal reasoning papers, codes, datasets, benchmarks and resources.
★ 36Boogu-Image. Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
★ 861AudioX-Turbo. 🚀 Fastest Anything-to-Audio Gen for conditioned sound and music creation.
★ 229MMAE. MMAE: A Massive Multitask Audio Editing Benchmark
★ 102sonoworld. Official implementation of the CVPR 2026 paper "SonoWorld: From One Image to a 3D Audio-Visual Scene."
★ 40Audio-Omni. [SIGGRAPH 2026] Repository of Audio-Omni
★ 399SAGE-GRPO. Official Implementation of SAGE-GRPO:Manifold-Aware Exploration for Reinforcement Learning in Video Generation
★ 127T2A-bench. [ICLR 2026] A two-stage evaluation benchmark for Text-to-Audio generation, assessing category, count, ordering, and timestamp accuracy with Gemini 2.5 Pro.
★ 3skills. Allow your 🦞 bot to Shout, Speak, with "human" vibe
★ 526Capybara. Python
★ 203Speech-Resources. 语音方向实验室/公司/资源/实习等,欢迎推荐或自荐
★ 609Awesome-World-Models. A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related websites.
★ 1.9klingbot-world. Advancing Open-source World Models
★ 4.3kAudioEditingCode. Python
★ 195LongVideoAgent. Python
★ 120Diff-Foley. Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models
★ 206Vision-to-Audio-and-Beyond. ICML 2024 "From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation"
★ 10Audeo. Python
★ 31xRIR_code. [CVPR 2025] Pytorch implementation of the paper "Hearing Anywhere in Any Environment"
★ 34video-to-audio-through-text. [NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”
★ 38ScaleCUA. [ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).
★ 1.1kWeTok. [ICLR2026] WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
★ 71Awesome-Vison2Audio. A curated list of Vision (video/image) to Audio Generation
★ 107Vlogger. [CVPR2024] Make Your Dream A Vlog
★ 435Video-GPT. [ICLR2026] Video-GPT via Next Clip Diffusion.
★ 45AudioX. [ICLR 2026] Repository of AudioX
★ 1.5kVideoGen-of-Thought. [Neurips 2025 NextVid Workshop Oral✨] Official Implementation of VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
★ 63Awesome-System2-Reasoning-LLM. Latest Advances on System-2 Reasoning
★ 1.4kAudio-FLAN. Audio-FLAN
★ 161VideoTuna. Let's finetune video generation models!
★ 551LibQPEP. TRO 2022 - QPEP: A C++/MATLAB library for solving generalized quadratic pose estimation problems and related uncertainty description
★ 185VidMuse. [CVPR 2025] Repository of VidMuse
★ 140Law_of_Vision_Representation_in_MLLMs. [COLM'25] Official implementation of the Law of Vision Representation in MLLMs
★ 176MultiTarget_WiFi_DFL. Jupyter Notebook
★ 1conditional-flow-matching. TorchCFM: a Conditional Flow Matching library
★ 2.6kstable-audio-tools. Generative models for conditional audio generation
★ 3.8kMMTrail. [Arxiv 2024] Official code for MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
★ 34suno-api. Use API to call the music generation AI of suno.ai, and easily integrate it into agents like GPTs.
★ 3.1klp-music-caps. LP-MusicCaps: LLM-Based Pseudo Music Captioning [ISMIR23]
★ 348audioset_tagging_cnn. Python
★ 1.8kDiffSHEG. [CVPR'24] DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
★ 207ai-audio-startups. Community list of startups working with AI in audio and music technology
★ 1.8kChatMusician. Python
★ 316Seeing-and-Hearing. [CVPR 2024] Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
★ 155Awesome-LLMs-meet-Multimodal-Generation. 🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 551wav2clip-changed. Python
★ 2Stable-Diffusion. FLUX, Stable Diffusion, SDXL, SD3, LoRA, Fine Tuning, DreamBooth, Training, Automatic1111, Forge WebUI, SwarmUI, DeepFake, TTS, Animation, Text To Video, Tutorials, Guides, Lectures, Courses, ComfyUI, Google Colab, RunPod, Kaggle, NoteBooks, ControlNet, TTS, Voice Cloning, AI, AI News, ML, ML News, News, Tech, Tech News, Kohya, Midjourney, RunPod
★ 2.8kbark. 🔊 Text-Prompted Generative Audio Model
★ 39kConference-Accepted-Paper-List. Some Conferences' accepted paper lists (including AI, ML, Robotic)
★ 1.3kGenAI_LLM_timeline. ChatGPT, GenerativeAI and LLMs Timeline
★ 953open-musiclm. Implementation of MusicLM, a text to music model published by Google Research, with a few modifications.
★ 560denoising-diffusion-pytorch. Implementation of Denoising Diffusion Probabilistic Model in Pytorch
★ 11kFollowYourPose. [AAAI 2024] Follow-Your-Pose: This repo is the official implementation of "Follow-Your-Pose : Pose-Guided Text-to-Video Generation using Pose-Free Videos"
★ 1.4kvideo-bgm-generation. [ACM MM 2021 Best Paper Award] Video Background Music Generation with Controllable Music Transformer
★ 327FateZero. [ICCV 2023 Oral] "FateZero: Fusing Attentions for Zero-shot Text-based Video Editing"
★ 1.2kHFGI3D. Jupyter Notebook
★ 206All-In-One-Deflicker. [CVPR2023] Blind Video Deflickering by Neural Filtering with a Flawed Atlas
★ 761