This is your work, valued
GSound-SIR. A Python Room Spatial Impulse Response Ray-Tracing Toolkit
★ 86SingFake. Official Repository for "SingFake: Singing Voice Deepfake Detection"
★ 64TrainingFreeMultiStepASR. Official Repository for "Training-Free Multi-Step Audio Source Separation"
★ 54SynthTab. Official Repository for ICASSP 2024 Paper "SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription"
★ 33music-source-restoration. Official Repository for "Music Source Restoration"
★ 31MSRKit. Model Implementations, Evaluation Scripts, etc. for Music Source Restoration Challenge 2025.
★ 23AreYouReallyListening. Official Repository for ISMIR 2025 paper "Are you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks"
★ 20ambisonizer. Official Repository for paper "Ambisonizer: Neural Upmixing as Spherical Harmonics Generation"
★ 19OpenGufeng. An Open-source Gufeng Melody and Chord Dataset
★ 15BachDuet-WebGUI. A Web Application for Baroque-style Human/Computer Musical Jamming.
★ 15Euterpe. A web-based template for hosting systems for real-time music HCI.
★ 14Real-time-Convolution-Reverb-on-Teensy. A Real-time Convolution Reverb system developed for Teensy 4.1
★ 10PhaseAntispoofing_INTERSPEECH. Official repository of the paper "Phase perturbation improves channel robustness for speech spoofing countermeasures"
★ 10International-Now. Internationalize any chinese lyric. Inspired By INTO1.
★ 9FakeSing. Enhance Your Live Vocal Performance with Real-Time Voice Tuning Against Pre-Recorded Track
★ 7AES_Microphone_Preamp_Data. Official Repository for AES NY 2023 Paper "Master Bus Coloring with Microphone Preamplifiers"
★ 2SSL_Anti-spoofing. This repository includes the code to reproduce our paper "Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation".
★ 1YuE. YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
★ 1math-wall-now. An LLM Wrapper Agentic Math Wall Generator
★ 1Open-Suno. trying to reproduce suno v3
★ 1Color-Latex-Table. Easily color your latex tables based on their corresponding values.
★ 1genblaze. Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
★ 504muscriptor. A multi-instrument music transcription model developed by Kyutai and Mirelo.
★ 938prose. A new kind of language for a new kind of computer
★ 1.7kskills. Skills for Real Engineers. Straight from my .agents directory.
★ 198kMOSS-TTS-Nano. MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run directly on CPU without a GPU, and keeps the deployment stack simple enough for local demos, web serving, and lightweight product integration.
★ 4kMOSS-TTS. MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
★ 3.9kdeep-representation-learning-book. Learning Deep Representations of Data Distributions
★ 1knnsight. The nnsight package enables interpreting and manipulating the internals of deep learned models.
★ 1kunipase. Official repository of UniPASE, a SOTA USE model
★ 54RCLI. Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
★ 1.5ktrilobyte-lossless-codec. Trilobyte Lossless Codec
★ 8slime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kheartlib. HeartMuLa Official Repo: The Most Powerful Open-Source Music Generation Model of 2026
★ 3.8ke2e. Official JAX implementation of End-to-End Test-Time Training for Long Context
★ 628PhysicsLM4. Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Reality
★ 356Awesome-ML-SYS-Tutorial. My learning notes for ML SYS.
★ 6.8kMiraTTS. A high quality and fast TTS repository
★ 518ml-sharp. Sharp Monocular View Synthesis in Less Than a Second
★ 8.8kPaCoRe. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
★ 338sim. Build, deploy, and orchestrate AI agents. Sim is the central intelligence layer for your AI workforce.
★ 29kwindowed-roformer. Official Repository for "Efficient Vocal Source Separation Through Windowed RoFormer"
★ 46smule-renaissance. Official Repository of Smule Renaissance, Smule's Vocal Restoration Models
★ 43seaweedfs. SeaweedFS is a distributed storage system for object storage (S3), file systems, and Iceberg tables, designed to handle billions of files with O(1) disk access and effortless horizontal scaling.
★ 34kmellow. small audio language model for reasoning
★ 88Syllabus. Synchronized Curriculum Learning for RL Agents
★ 123py-cpuinfo. A module for getting CPU info with pure Python
★ 343SonicMaster. SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
★ 191AISystem. AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术
★ 17kWan2.2. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kMuQ. Official repository of the paper "MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization".
★ 362CSR_Adaptive_Rep. Official Code for Paper: Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
★ 141AntiDeepfake. Project for training SSL-based deepfake speech detector
★ 56MultilingualALT. Repo of the paper "Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model""
★ 15genmusic_demo_list. a list of demo websites for automatic music generation research
★ 794AReaL. The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
★ 5.6khnet. H-Net: Hierarchical Network with Dynamic Chunking
★ 869mujoco. Multi-Joint dynamics with Contact. A general purpose physics simulator.
★ 14karia. Official repository for the paper: Scaling Self-Supervised Representation Learning for Symbolic Piano Performance (ISMIR 2025)
★ 109modded-nanogpt. NanoGPT (124M) in 90 seconds
★ 5.6ksymbolicai. A neurosymbolic perspective on LLMs
★ 1.7karia-midi. Official repository for Aria-MIDI: a MIDI dataset of 1,186,253 transcribed solo-piano recordings.
★ 98GDRetriever. Official implementation of the paper - GD-Retriever: Controllable generative text-music retrieval with diffusion models (Accepted at ISMIR25!)
★ 19LMCache. LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 11k1d-tokenizer. This repo contains the code for 1D tokenizer and generator
★ 1.2kefficient-spherical-harmonic-evaluation. http://jcgt.org/published/0002/02/06/
★ 16SongEval. A song aesthetic evaluation toolkit trained on SongEval.
★ 3SonicVerse. SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
★ 53frameless-eval. Lose the Frames: Event-based Metrics for Efficient Music Structure Analysis Evaluation
★ 11torch-audiomentations. Fast audio data augmentation in PyTorch. Inspired by audiomentations. Useful for deep learning.
★ 1.2kRSTnet. Real-time Speech-Text Foundation Model Toolkit (wip)
★ 255MeanFlow. PyTorch implementation of MeanFlow & iMF (one-step generative modeling).
★ 1.2kWildFX. Official implementation of WildFX Dataset Generating pipeline.
★ 21UCGM. [Preprint] UCGM: Unified Continuous Generative Models
★ 187REPA-E. [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
★ 512ZeroSep. [NeurIPS 2025] Separate Anything in Audio with Zero Training
★ 60Paper2Poster. [NeurIPS 2025] Open-source Multi-agent Poster Generation from Papers
★ 3.9kn8n-workflows. all of the workflows of n8n i could find (also from the site itself)
★ 56kminimp3py. Python bindings for minimp3
★ 17verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kDistributionEmbeddings. Jupyter Notebook
★ 42arxiv-latex-cleaner. arXiv LaTeX Cleaner: Easily clean the LaTeX code of your paper to submit to arXiv
★ 7kOpenRLHF-M. An Easy-to-use, Scalable and High-performance RLHF Framework designed for Multimodal Models.
★ 163visage. Official implementation of "ViSAGe: Video-to-Spatial AUdio Generation" (ICLR 2025)
★ 47LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kPINN-for-HRTF-upsampling. A physics-informed neural network method for head-related transfer function upsampling
★ 10ThunderKittens. Tile primitives for speedy kernels
★ 3.6kkleidiai. This repository is a read-only mirror of https://gitlab.arm.com/kleidi/kleidiai
★ 174matching-pursuit. This repository contains research and experiments aimed at producing sparse, interpretable representations of audio.
★ 8Systems-for-Foundation-Models.
★ 20Embodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15ktorchcrepe. Pytorch implementation of the CREPE pitch tracker
★ 523DCUnet. Phase-aware speech enchancement with Deep Complex U-Net
★ 139eko. Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai
★ 4.9ksound_field_nn. Python
★ 5Dolphin. Dolphin is a multilingual, multitask ASR model jointly trained by DataoceanAI and Tsinghua University.
★ 777music_source_separation. Python
★ 60r1-aqa. 🤗 R1-AQA Model: mispeech/r1-aqa
★ 325sylber. Sylber: Syllabic Embedding Representation of Speech from Raw Audio
★ 80LightningDiT. [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
★ 1.5ks1. s1: Simple test-time scaling
★ 6.7kaudio_caption_metrics. Metrics for evaluating audio caption
★ 3torchtitan. A PyTorch native platform for training generative AI models
★ 5.6kflash-linear-attention. 🚀 Efficient implementations for emerging model architectures
★ 5.5kYuE. YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
★ 6.4kMP-SENet. Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
★ 494stable-ts. Transcription, forced alignment, and audio indexing with OpenAI's Whisper
★ 2.3kR3GAN. Code for NeurIPS 2024 paper - The GAN is dead; long live the GAN! A Modern Baseline GAN - by Huang et al.
★ 871stable-audio-controlnet. Fine-tune Stable Audio Open with DiT ControlNet.
★ 256clarity-template. Clarity: A Minimalist Website Template for AI Research
★ 225awesome-lifelong-learning-methods-for-llm. [ACM Computing Surveys 2025] This repository collects awesome survey, resource, and paper for Lifelong Learning with Large Language Models. (Updated Regularly)
★ 164Awesome-Multimodal-Next-Token-Prediction. [Survey] Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
★ 477Audio-Denoiser-ONNX. Utilizes ONNX Runtime for audio denoising.
★ 134VAD_Benchmark. Benchmarking different VAD models on AVA-Speech dataset
★ 19ClearerVoice-Studio. An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
★ 4.4kgolf. A DDSP-based neural voice synthesiser.
★ 135BigCodec. Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"
★ 218ssl_speech_restoration_v2. Python
★ 17stream-vc. An unofficial PyTorch implementation of the StreamVC(Real-Time Low-Latency Voice Conversion)
★ 129pmqd. Perceived Music Quality Dataset
★ 12wechatbot-webhook. 轻量、可部署的微信机器人webhook服务,使用http接口收发微信消息, 用它作为个人通知、AIGC 应用或者 coze、n8n等自动化工作流的消息节点
★ 2.2kpg. PostgreSQL notes
★ 525semantic. Parsing, analyzing, and comparing source code across many languages
★ 9kBEHAVIOR-1K. BEHAVIOR-1K: a platform for accelerating Embodied AI research. Join our Discord for support: https://discord.gg/bccR5vGFEx
★ 1.6kREPA. [ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
★ 1.7kcgra4ml. An Open Workflow to Build Custom SoCs and run Deep Models at the Edge
★ 126SSR-Speech. SSR-Speech: Towards Stable, Safe and Robust Zero-shot Speech Editing and Synthesis
★ 155search_rec_ads_papers. Papers on Search, Recommendations, and Ads (搜广推)
★ 34ai-by-hand-excel.
★ 6.2kMIMO. Official implementation of "MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling"
★ 1.6kSOFA. SOFA: Singing-Oriented Forced Aligner
★ 228Pake. 🤱🏻 Turn any webpage into a desktop app with one command.
★ 60kEzAudio. High-quality Text-to-Audio Generation with Efficient Diffusion Transformer
★ 333MusicTI_AAAI2024. " Music Style Transfer with Time-Varying Inversion of Diffusion Models"
★ 59FoleyCrafter. [IJCV 2026] FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds. AI拟音大师,给你的无声视频添加生动而且同步的音效 😝
★ 659FunCodec. FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.
★ 445SEtrain. A training code template for DNN-based speech enhancement.
★ 207soundata. Python library for downloading, loading & working with sound datasets
★ 357stable-audio-tools. Generative models for conditional audio generation
★ 3.8kSimpleTuner. A general fine-tuning kit geared toward image/video/audio diffusion models.
★ 2.9krtfm. Research on Tabular Foundation Models
★ 71cambrian. Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2kttt-lm-pytorch. Official PyTorch implementation of Learning to (Learn at Test Time): RNNs with Expressive Hidden States
★ 1.4kHPSv2. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
★ 677voyager. 🛰️ An approximate nearest-neighbor search library for Python and Java with a focus on ease of use, simplicity, and deployability.
★ 1.6kCacophony. Inference codebase for "Cacophony: An Improved Contrastive Audio-Text Model". Preprint: https://arxiv.org/abs/2402.06986
★ 49WavCraft. Official repo for WavCraft, an AI agent for audio creation and editing
★ 523ltfatpy. This is the LTFATPY development repository.
★ 6ImHex. 🔍 A Hex Editor for Reverse Engineers, Programmers and people who value their retinas when working at 3 AM.
★ 54kminGPT. A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
★ 25kxtuner. A Next-Generation Training Engine Built for Ultra-Large MoE Models
★ 5.2kLLaMA-Adapter. [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5.9kVideo-LLaMA. [EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
★ 3.1kSingGraph. Official repository for the paper Singing Voice Graph Modeling for SingFake Detection (Interspeech 2024).
★ 24AudioSep. Official implementation of "Separate Anything You Describe"
★ 1.9kTRT-SE. An example of a speech enhancement model deployed with TensorRT.
★ 88soundctm. Pytorch implementation of SoundCTM
★ 101LLM-Codec. The open source code for LLM-Codec
★ 147DTTNet-Pytorch. An official implementation of the ICASSP 2024 paper: Dual-Path TFC-TDF UNet for Music Source Separation
★ 109ultimatevocalremovergui. GUI for a Vocal Remover that uses Deep Neural Networks.
★ 26kSimpleDiarization. Simple diarization model
★ 53Music-Source-Separation-Training. Repository for training models for music source separation.
★ 1.5kdreamgaussian4d. [arXiv 2023] DreamGaussian4D: Generative 4D Gaussian Splatting
★ 619AudioLDM. AudioLDM: Generate speech, sound effects, music and beyond, with text.
★ 2.9kaudio-diffusion. Apply diffusion models using the new Hugging Face diffusers package to synthesize music instead of images.
★ 793DiT. Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
★ 8.7kcroissant. Croissant is a high-level format for machine learning datasets that brings together four rich layers.
★ 882se-scaling. Model configurations for scaling SE models in the paper "Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement"
★ 42awesome-jax. JAX - A curated list of resources https://github.com/google/jax
★ 2.1kfish-speech. SOTA Open Source TTS
★ 32karia-amt. Efficient and robust implementation of seq-to-seq automatic piano transcription.
★ 71TaskMatrix. Python
★ 34kvector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4kddsp_pytorch. Implementation of Differentiable Digital Signal Processing (DDSP) in Pytorch
★ 518brouhaha-vad. Predicts the level of noise and reverberation on your audiofiles
★ 191ChatTTS. A generative speech model for daily dialogue.
★ 40kschedule_free. Schedule-Free Optimization in PyTorch
★ 2.3kAwesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
★ 3.3kAVSpatialAlignment. C++
★ 31DiM-DiffusionMamba. The official implementation of DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis
★ 240tango. A family of diffusion models for text-to-audio generation.
★ 1.2kAwesome-TimeSeries-SpatioTemporal-Diffusion-Model. A survey and paper list of current Diffusion Model for Time Series and SpatioTemporal Data with awesome resources (paper, application, review, survey, etc.).
★ 1kconditional-flow-matching. TorchCFM: a Conditional Flow Matching library
★ 2.6klabel-studio. Label Studio is a multi-type data labeling and annotation tool with standardized output format
★ 28kfrechet-audio-distance. A lightweight library for Frechet Audio Distance calculation.
★ 317MR-MT3. MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
★ 57basic-pitch-torch. PyTorch version of Spotify's Basic Pitch
★ 54timbre-trap. Code for the paper "Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription"
★ 43Dynamic3DGaussians. Python
★ 2.3kgentle. gentle forced aligner
★ 1.7kA-Convolutional-Recurrent-Neural-Network-for-Real-Time-Speech-Enhancement. A minimum unofficial implementation of the "A Convolutional Recurrent Neural Network for Real-Time Speech Enhancement" (CRN) using PyTorch
★ 350SceneFake. baseline for SMAD dataset
★ 8hilcodec. High fidelity, lightweight, end-to-end, streaming, convolution-based neural audio codec
★ 120SpatialScaper. Jupyter Notebook
★ 75buddy. BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models
★ 66minitorch. The full minitorch student suite.
★ 2.4kaudioseal. Localized watermarking for AI-generated speech audios, with SOTA on robustness and very fast detector
★ 762inspect_ai. Inspect: A framework for large language model evaluations
★ 2.4kzimtohrli. Jupyter Notebook
★ 215gtcrn. The official implementation of GTCRN, an ultra-lightweight SE model.
★ 702LLMs_interview_notes. 该仓库主要记录 大模型(LLMs) 算法工程师相关的面试题
★ 2.6kLLMs_interview_notes. LLMs interview notes and answers:该仓库主要记录大模型(LLMs)算法工程师相关的面试题和参考答案
★ 641demucs. Code for the paper Hybrid Spectrogram and Waveform Source Separation
★ 3kFreeVC. FreeVC: Towards High-Quality Text-Free One-Shot Voice Conversion
★ 714Diff-VC. Diffusion Model for Voice Conversion
★ 72DiffMorpher. Official Code for DiffMorpher: Unleashing the Capability of Diffusion Models for Image Morphing (CVPR 2024)
★ 506RAVE. Official implementation of the RAVE model: a Realtime Audio Variational autoEncoder
★ 1.8kAcademiCodec. AcademiCodec: An Open Source Audio Codec Model for Academic Research
★ 674EDMSound. Codebase and project page for EDMSound
★ 35mamba. Mamba SSM architecture
★ 19kpaper-reading. 深度学习经典、新论文逐段精读
★ 34kehrshot-benchmark. A benchmark for few-shot evaluation of foundation models for electronic health records (EHRs)
★ 229DawDreamer. Digital Audio Workstation with Python; VST instruments/effects, parameter automation, FAUST, JAX, Warp Markers, and JUCE processors
★ 1.3konnx2tf. A tool for converting ONNX files to LiteRT/TFLite/TensorFlow, PyTorch native code (nn.Module), TorchScript (.pt), state_dict (.pt), Exported Program (.pt2), and Dynamo ONNX. It also supports direct conversion from LiteRT to PyTorch.
★ 986StoryTTS. [ICASSP 2024] StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
★ 141streaming. A Data Streaming Library for Efficient Neural Network Training
★ 1.5kffcv. FFCV: Fast Forward Computer Vision (and other ML workloads!)
★ 3kgcn-tfilm. Modelling black-box audio effects with time-varying feature modulation
★ 45music-generation-research. A straightforward collection of Music Generation research resources.
★ 606Los-Angeles-MIDI-Dataset. SOTA kilo-scale MIDI dataset for MIR and Music AI purposes
★ 68FSD-Dataset. This repository presents FSD dataset for song deepfake detection.
★ 24