This is your work, valued
cocosda-SSL. pytorch code for sound event localization and classification
★ 13Audio-visual-sound-localization. Audio-visual sound localization
★ 11TASLP2022-AVRI. Python
★ 7SSL-Analytic-Class-Incremental-Learning-. Analytic Class Incremental Learning for Sound Source Localization with Privacy Protection
★ 3ICASSP2021-Audio-only. Python
★ 1SEANet. Code for Audio-Visual Target Speaker Extraction with Selective Auditory Attention (TASLP)
★ 32SSDQ. Python
★ 8SimDINO. [ICML 2025] Official Implementation for SimDINO/SimDINOv2
★ 206stable-audio-tools. Generative models for conditional audio generation
★ 3.8kVoice-Separation-and-Enhancement. A framework for quick testing and comparing multi-channel speech enhancement and separation methods, such as DSB, MVDR, LCMV, GEVD beamforming and ICA, FastICA, IVA, AuxIVA, OverIVA, ILRMA, FastMNMF.
★ 177basic-pitch. A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
★ 5.4kism. Python
★ 5hoa. Python
★ 8Kimi-Audio. Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
★ 4.7kOmniAudio. [ICML 2025] PyTorch Implementation of "OmniAudio: Generating Spatial Audio from 360-Degree Video"
★ 375MMAR. [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
★ 214ddsp. DDSP: Differentiable Digital Signal Processing
★ 3.3ksfs-python. SFS Toolbox for Python
★ 73LAVSS. Code for LAVSS: Location-Guided Audio-Visual Spatial Audio Separation
★ 19GWA. Geometric-Wave Acoustic dataset
★ 66ss-vq-vae. Self-supervised VQ-VAE for One-Shot Music Style Transfer
★ 99control-transfer-diffusion. Repository for the paper "Combining audio control and style transfer using latent diffusion", accepted at ISMIR 2024
★ 67Audio-FLAN. Audio-FLAN
★ 161seed-vc. zero-shot voice conversion & singing voice conversion, with real-time support
★ 3.9kImplementation-of-GelSight. An implementation of Gelsight Wedge.
★ 29BNM. code of Towards Discriminability and Diversity: Batch Nuclear-norm Maximization under Label Insufficient Situations (CVPR2020 oral)
★ 269Versatile-Domain-Adaptation. Code Release for "Minimum Class Confusion for Versatile Domain Adaptation"(ECCV2020)
★ 56deepsing. Generating Sentiment-aware Visual Stories using Cross-modal Music Translation
★ 71Spatial-AST. 🦇 Encoder of BAT (Learning to Reason about Spatial Sounds with Large Language Models)
★ 87ClearerVoice-Studio. An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
★ 4.3kTCSinger. PyTorch Implementation of TCSinger(EMNLP 2024): Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
★ 386SpatialScaper. Jupyter Notebook
★ 75KRT-MLCIL. Python
★ 12speechmetrics. A wrapper around speech quality metrics MOSNet, BSSEval, STOI, PESQ, SRMR, SISDR
★ 1.1kPanoAVQA. Official repository of PanoAVQA: Grounded Audio-Visual Question Answering in 360° Videos (ICCV 2021)
★ 16RealMAN. A description of "RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization" [NeurIPS 2024]
★ 175PianoMotion10M. Code release for PianoMotion10M
★ 117Datasets. Poetry-related datasets developed by THUAIPoet (Jiuge) group.
★ 240facechain. FaceChain is a deep-learning toolchain for generating your Digital-Twin.
★ 9.5kAnalytic-continual-learning. This repository will be posting analytic continual learning series, including Analytic Class-Incremental Learning (ACIL), Gaussian Kernel Embedded Analytic Learning (GKEAL), Dual-Stream Analytic Learning (DS-AL), etc.
★ 289AV-CIL_ICCV2023. [ICCV 2023] Audio-Visual Class-Incremental Learning
★ 36AV-Sepformer. Python
★ 65childrenize. Signal processing method to convert adult speech into child-like
★ 9SLfM. Official code for the paper: [ICCV2023] Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation
★ 43PyTouch. PyTouch is a machine learning library for tactile touch sensing.
★ 274TTS. 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 46kVoiceLDM. VoiceLDM: Text-to-Speech with Environmental Context
★ 194CREMA-D. Crowd Sourced Emotional Multimodal Actors Dataset (CREMA-D)
★ 543FGTD. Face Generation from Textual Description using GANs.
★ 14vits. VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
★ 7.9kAmphion. Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
★ 10kQwen-Audio. The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.
★ 1.9kReal-Time-Voice-Cloning. Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 60kVALL-E-X. An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
★ 7.9kencodec. State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
★ 4kvall-e. PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
★ 2.2kNExT-GPT. Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6kLLaSM. 第一个支持中英文双语语音-文本多模态对话的开源可商用对话模型。便捷的语音输入将大幅改善以文本为输入的大模型的使用体验,同时避免了基于 ASR 解决方案的繁琐流程以及可能引入的错误。
★ 561TVLT. PyTorch code for “TVLT: Textless Vision-Language Transformer” (NeurIPS 2022 Oral)
★ 127muavic. MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
★ 404SALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kACMR_demo. Python
★ 93stereocrw. Code for the Paper: [ECCV2022] Sound Localization by Self-Supervised Time-Delay Estimation
★ 28MNBVC. MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。
★ 4.2kWavCaps. This reporsitory contains metadata of WavCaps dataset and codes for downstream tasks.
★ 264ImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1kLLMZoo. ⚡LLM Zoo is a project that provides data, models, and evaluation benchmark for large language models.⚡
★ 2.9klangchain. The agent engineering platform.
★ 143kTalkLip. Python
★ 429GCNet. GCNet, official pytorch implementation of our paper "GCNet: Graph Completion Network for Incomplete Multimodal Learning in Conversation"
★ 104L3DAS23. Official repository supporting the L3DAS23 IEEE ICASSP Grand Challenge
★ 16VirtualConductor. 🎶 Music-Driven Conducting Motion Generation (IEEE ICME'21 Best Demo)
★ 98s3prl. Self-Supervised Speech Pre-training and Representation Learning Toolkit
★ 2.6kLibri-adhoc40. A dataset collected from synchronized ad-hoc microphone arrays
★ 19MIntRec. MIntRec: A New Dataset for Multimodal Intent Recognition (ACM MM 2022)
★ 138DSTC7-Audio-Visual-Scene-Aware-Dialog-AVSD-Challenge. Python
★ 54VoiceSplit. VoiceSplit: Targeted Voice Separation by Speaker-Conditioned Spectrogram
★ 271pyroomacoustics. Pyroomacoustics is a package for audio signal processing for indoor applications. It was developed as a fast prototyping platform for beamforming algorithms in indoor scenarios.
★ 1.9kvoicefilter. Unofficial PyTorch implementation of Google AI's VoiceFilter system
★ 1.2kspeakerbeam. Jupyter Notebook
★ 146move2hear-active-AV-separation. Code and datasets for 'Move2Hear: Active Audio-Visual Source Separation' (ICCV 2021)
★ 16learning-audio-visual-dereverberation. Code for paper Learning Audio-Visual Dereverberation
★ 322.5D-Visual-Sound. 2.5D visual sound
★ 121BinauralSpeechSynthesis. N/A
★ 190FAIR-Play. 2.5D visual sound dataset
★ 108waveglow. A Flow-based Generative Network for Speech Synthesis
★ 2.3kGACELA. Generative adversarial context encoder for audio inpainting
★ 25EasyComDataset. The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the *cocktail party effect* from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
★ 143SoundingEarth. Self-supervised Audiovisual Representation Learning for Remote Sensing Data
★ 34audioContextEncoder. A context encoder for audio inpainting
★ 26Catch-A-Waveform. Official pytorch implementation of the paper: "Catch-A-Waveform: Learning to Generate Audio from a Single Short Example" (NeurIPS 2021)
★ 189ast. Code for the Interspeech 2021 paper "AST: Audio Spectrogram Transformer".
★ 1.5kdemucs. Code for the paper Hybrid Spectrogram and Waveform Source Separation
★ 10kspeech2gesture. code for training the models from the paper "Learning Individual Styles of Conversational Gestures"
★ 394Dancing2Music. Python
★ 538TFood. [CVPRW22] Official Implementation of T-Food: "Transformer Decoders with MultiModal Regularization for Cross-Modal Food Retrieval". Accepted at CVPR22 's MULA Workshop.
★ 34regnet. Official PyTorch implementation of the TIP paper "Generating Visually Aligned Sound from Videos" and the corresponding Visually Aligned Sound (VAS) dataset.
★ 53Wnet. Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks
★ 24CMBS. cross modal background suppression for audio-visual event localization
★ 36meshtalk. Code for MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
★ 402ObjectFolder. ObjectFolder Dataset
★ 173AVCA-GZSL. This repository contains the code for our CVPR 2022 paper on "Audio-visual Generalised Zero-shot Learning with Cross-modal Attention and Language"
★ 43DSCMR. Deep Supervised Cross-modal Retrieval (CVPR 2019, PyTorch Code)
★ 143headnerf. This repository contains a pytorch implementation of "HeadNeRF: A Real-time NeRF-based Parametric Head Model (CVPR 2022)".
★ 446FaceFormer. [CVPR 2022] FaceFormer: Speech-Driven 3D Facial Animation with Transformers
★ 915vit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kSRMRpy. Python implementation of the SRMR toolbox
★ 132nara_wpe. Different implementations of "Weighted Prediction Error" for speech dereverberation
★ 570SpeechSplit. Unsupervised Speech Decomposition Via Triple Information Bottleneck
★ 697voicefixer_main. General Speech Restoration
★ 287SkipConvNet. Speech Dereverberation using Fully Convolutional Networks
★ 77panns_inference. Python
★ 266audio-pretrained-model. A collection of Audio and Speech pre-trained models.
★ 192Multimodal-Aerial-Scene-Recognition. Code for <Cross-Task Transfer for Geotagged Audiovisual Aerial Scene Recognition> (ECCV 2020)
★ 35pytorch-grad-cam. Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
★ 13kdeep-acoustic-analysis. Python
★ 26imgaug. Image augmentation for machine learning experiments.
★ 15kvidaug. Effective Video Augmentation Techniques for Training Convolutional Neural Networks
★ 416spatialaudiogen. Spatial Audio Generation
★ 118sound-spaces. A first-of-its-kind acoustic simulation platform for audio-visual embodied AI research. It supports training and evaluating multiple tasks and applications.
★ 468speech2image. Attempting speech-driven image synthesis.
★ 3SAAVN. SAAVN Code release for paper "Sound Adversarial Audio-Visual Navigation,ICLR2022" (In PyTorch)
★ 21speechbrain. A PyTorch-based Speech Toolkit
★ 12kawesome-embodied-vision. Reading list for research topics in embodied vision
★ 704image2reverb. [ICCV 2021] Image2Reverb: Cross-Modal Reverb Impulse Response Synthesis.
★ 91ACVAE-VC. Jupyter Notebook
★ 22autoencoders. implementations of various types of auto-encoders in tensorflow(in progress)
★ 9sentence-transformers. State-of-the-Art Embeddings, Retrieval, and Reranking
★ 19kpsla. Code for the TASLP paper "PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation".
★ 150animegan2-pytorch. PyTorch implementation of AnimeGANv2
★ 4.5kAnimeGANv2. [Open Source]. The improved version of AnimeGAN. Landscape photos/videos to anime
★ 5.4kAnimeGAN. A Tensorflow implementation of AnimeGAN for fast photo animation ! This is the Open source of the paper 「AnimeGAN: a novel lightweight GAN for photo animation」, which uses the GAN framwork to transform real-world photos into anime images.
★ 4.6kPatchVAE. PyTorch implementation of "PatchVAE: Learning Local Latent Codes for Recognition" to appear in CVPR 2020
★ 14av-se. Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
★ 222VAE-Torch. Implementation of Variational Auto-Encoder in Torch7
★ 266MAE-pytorch. Unofficial PyTorch implementation of Masked Autoencoders Are Scalable Vision Learners
★ 2.7klyrebird-wav2clip. Official implementation of the paper WAV2CLIP: LEARNING ROBUST AUDIO REPRESENTATIONS FROM CLIP
★ 360aisle. Official Repository for the paper "No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures", Findings at EMNLP 2020
★ 20annotated_deep_learning_paper_implementations. 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠
★ 67kECAPA-TDNN. Unofficial reimplementation of ECAPA-TDNN for speaker recognition (EER=0.86 for Vox1_O when train only in Vox2)
★ 823NvEM. [ACM MM 2021 Oral] Official repo of "Neighbor-view Enhanced Model for Vision and Language Navigation"
★ 78VAE-ResNet18-PyTorch. A Variational Autoencoder based on the ResNet18-architecture
★ 121TorchSSL. A PyTorch-based library for semi-supervised learning (NeurIPS'21)
★ 1.4kVisualEchoes. VisualEchoes Dataset (ECCV 2020)
★ 37beyond-image-to-depth. Python
★ 38ADENET. Code used in the paper "Towards End-to-End Acoustic Localization using Deep Learning: from Audio Signal to Source Position Coordinates"
★ 7PyTorch-VAE. A Collection of Variational Autoencoders (VAE) in PyTorch.
★ 7.7kInstanceLoc. [CVPR 2021] Instance Localization for Self-supervised Detection Pretraining
★ 145MM-DistillNet. PyTorch code for training MM-DistillNet for multimodal knowledge distillation. http://rl.uni-freiburg.de/research/multimodal-distill
★ 57KWS-Net. Seeing Wake Words: Audio-visual Keyword Spotting
★ 67nerf-pytorch. A PyTorch implementation of NeRF (Neural Radiance Fields) that reproduces the results.
★ 6.1kAD-NeRF. This repository contains a PyTorch implementation of "AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis".
★ 1.1kml-cpc. Jupyter Notebook
★ 37google-research. Google Research
★ 38kcocon. CoCon: Cooperative Contrastive Learning
★ 20General-network-architecture-for-sound-event-localization-and-detection. This repository consists of python code to train sound event localization and detection models.
★ 22OpenSound. Various Audio Process Baselines
★ 8EIN-SELD. An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
★ 79tal-hmo. Fusional approaches for temporal action localization in untrimmed videos
★ 35CVPR2026-Papers-with-Code. CVPR 2026 论文和开源项目合集
★ 23kseld-dcase2021. Baseline method for sound event localization task of DCASE 2021 challenge
★ 45DCASE2020_task1. Code for DCASE 2020 task 1a and task 1b.
★ 88AudioCLIP. Source code for models described in the paper "AudioCLIP: Extending CLIP to Image, Text and Audio" (https://arxiv.org/abs/2106.13043)
★ 872Awesome-Incremental-Learning. Awesome Incremental Learning
★ 4.5kFaceX-Zoo. A PyTorch Toolbox for Face Recognition
★ 2kReverberated_WSJ_2MIX. Code to simulate a reverberated, noisy version of the WSJ-2MIX dataset
★ 22Cross3D. Code repository for the paper Robust Sound Source Tracking Using SRP-PHAT and 3D Convolutional Neural Networks
★ 90speech2affective_gestures. This is the official implementation of the paper "Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning".
★ 54lyra. A Very Low-Bitrate Codec for Speech Compression
★ 4kaudiomentations. A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
★ 2.3kL3DAS21. Python
★ 37Speech-Emotion-Classification-with-PyTorch. This repository contains PyTorch implementation of 4 different models for classification of emotions of the speech.
★ 213mir_eval. Evaluation functions for music/audio information retrieval/signal processing algorithms.
★ 708TalkNet-ASD. ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'
★ 489fixmatch. A simple method to perform semi-supervised learning with limited data.
★ 1.2kMixMatch-pytorch. Code for "MixMatch - A Holistic Approach to Semi-Supervised Learning"
★ 652unsupervised-data-augmentation. Unofficial PyTorch Implementation of Unsupervised Data Augmentation.
★ 148uda. Unsupervised Data Augmentation (UDA)
★ 2.2kMLP-Mixer-pytorch. Unofficial implementation of MLP-Mixer: An all-MLP Architecture for Vision
★ 216audioset_tagging_cnn. Python
★ 1.8kExternal-Attention-pytorch. 🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
★ 12kKaleido-BERT. 💐Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
★ 230glow-pytorch. pytorch implementation of openai paper "Glow: Generative Flow with Invertible 1×1 Convolutions"
★ 514deep-clustering-1. deep clustering method for single-channel speech separation
★ 1DeeperForensics-1.0. [CVPR 2020] A Large-Scale Dataset for Real-World Face Forgery Detection
★ 586SpEx_Plus. SpEx+(tied) source code
★ 96routing-transformer. Fully featured implementation of Routing Transformer
★ 300ai-research-code. Python
★ 357DynamicViT. [NeurIPS 2021] [T-PAMI] DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
★ 668voxceleb_trainer. In defence of metric learning for speaker recognition
★ 1.2kSpeaker-Recognition-Demo. A ResNet Speaker Recognition&Verification Demo
★ 27CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kStyleCLIP. Official Implementation for "StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery" (ICCV 2021 Oral)
★ 4.1kgpuRIR. Python library for Room Impulse Response (RIR) simulation with GPU acceleration
★ 607scnn. Segment-CNN: A Framework for Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs
★ 233SupContrast. PyTorch implementation of "Supervised Contrastive Learning" (and SimCLR incidentally)
★ 3.4kSimCLR. PyTorch implementation of SimCLR: A Simple Framework for Contrastive Learning of Visual Representations
★ 2.5kpytorchvideo. A deep learning library for video understanding research.
★ 3.6kVMZ. VMZ: Model Zoo for Video Modeling
★ 1.1ktriplet-loss-mnist. Triplet Loss 损失函数
★ 209ViLT. Code for the ICML 2021 (long talk) paper: "ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision"
★ 1.5kPseudoBinaural_CVPR2021. Codebase for the paper "Visually Informed Binaural Audio Generation without Binaural Audios" (CVPR 2021)
★ 72DeepAFx. Third-party audio effects plugins as differentiable layers within deep neural networks.
★ 215AVID-CMA. Audio Visual Instance Discrimination with Cross-Modal Agreement
★ 133Conv-TasNet. Python
★ 337