This is your work, valued
Research Scientist at Google DeepMind
multimodal-ml-music. List of academic resources on Multimodal ML for Music
★ 300muscall. Official implementation of "Contrastive Audio-Language Learning for Music" (ISMIR 2022)
★ 123word2wave. Word2Wave: a framework for generating short audio samples from a text prompt using WaveGAN and COALA.
★ 119muscaps. Source code for "MusCaps: Generating Captions for Music Audio" (IJCNN 2021)
★ 86song-describer. Song Describer is a data collection platform for annotating music with textual descriptions.
★ 61mulap. Official implementation of "Learning Music Audio Representations Via Weak Language Supervision" (ICASSP 2022)
★ 47music-audio-tagging-pytorch. A PyTorch implementation of the musicnn model for music audio tagging
★ 38bela-sampler. MIDI controlled sampler/sequencer on the Bela platform
★ 6mir-datasets. Simple collection of MIR datasets with metadata and links
★ 2hidden-markov-models. Jupyter Notebook
★ 1torch-pitch-shift. Pitch-shift audio clips quickly with PyTorch (CUDA supported)! Additional utilities for searching efficient transformations are included.
★ 1astro-image-processing. Code for the implementation of a cluster detection algorithm to analyse an optical image and study the distribution of galaxies within it
★ 1oslo-model. Implementation, analysis and visualisation of the Oslo rice pile Model
★ 1audio-file-mcp-app. An MCP App for playing and inspecting local audio files in an MCP host.
★ 42pure-vibes. Pure Data 🤝 AI Agents
★ 20the-infinite-crate. The Infinite Crate is a DAW plugin built on JUCE, React, and the Lyria RealTime live music model
★ 102magenta-realtime. Magenta RealTime 2: An Open-Weights Live Music Model
★ 1.7kquantum-audio. A Python package for building Quantum Representations of Digital Audio. Developed by Moth.
★ 50AudioBench. AudioBench: A Universal Benchmark for Audio Large Language Models
★ 319audio-ai-hub. The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
★ 949lovely-llama. An implementation of the Llama architecture, to instruct and delight
★ 21gpu-optimization-workshop. Slides, notes, and materials for the workshop
★ 342memorization_generalization_in_diffusion_models. Jupyter Notebook
★ 60all-in-one. All-In-One Music Structure Analyzer
★ 808fadtk. A simple library for Fréchet Audio Distance (FAD) calculation
★ 266AQUA-Tk. AQUA-Tk = Audio QUality Assessment-Toolkit. (In development)
★ 105genmusic_demo_list. a list of demo websites for automatic music generation research
★ 794Academic-project-page-template. A project page template for academic papers. Demo at https://eliahuhorwitz.github.io/Academic-project-page-template/
★ 5.1kaudio-ai-timeline. A timeline of the latest AI models for audio generation, starting in 2023!
★ 1.9kaudio-diffusion-pytorch. Audio generation using diffusion models, in PyTorch.
★ 2.1kllark. Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.
★ 383research-templates. LaTeX templates I created for authoring research papers
★ 16quivr. Opiniated RAG for integrating GenAI in your apps 🧠 Focus on your product rather than the RAG. Easy integration in existing products with customisation! Any LLM: GPT4, Groq, Llama. Any Vectorstore: PGVector, Faiss. Any Files. Anyway you want.
★ 39kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kchatgpt-retrieval-plugin. The ChatGPT Retrieval Plugin lets you easily find personal or work documents by asking questions in natural language.
★ 21kmsu-benchmark. music semantic understanding evaluation benchmark
★ 24audio-dataset. Audio Dataset for training CLAP and other models
★ 748musiclm-pytorch. Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch
★ 3.3kCLAP. Contrastive Language-Audio Pretraining
★ 2.2kmusic_caps_dl. Unofficial download repository for MusicCaps
★ 47tuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
★ 30kmsd-subsets. million song dataset split for extended clean tag & artist-level stratified
★ 52ismir2022-datasets. list of MIR dataset papers presented at ISMIR 2022
★ 61music-audio-representations. Results and Models for Learning Audio Representations of Music Content
★ 107awesome-music-informatics. A curated list of awesome article, tutorial, library, webpage, etc.
★ 195DJ-Tools. DJ tools for downloading, processing, and sharing music / rekordbox data.
★ 146audio-algebra. alchemy with embeddings
★ 34VQ-Diffusion. Official implementation of VQ-Diffusion
★ 981NIME_workshop.
★ 14ai-audio-startups. Community list of startups working with AI in audio and music technology
★ 1.8ksimplified-jukemir. A minimum JukeMIR branch for feature extraction.
★ 32composer. Supercharge Your Model Training
★ 5.5kscaper. A library for soundscape synthesis and augmentation
★ 426hear-baseline. Simple baseline model for the HEAR benchmark
★ 23hear-eval-kit. Evaluation kit for the HEAR Benchmark
★ 65hear-preprocess. Dataset preprocessing code for the HEAR 2021 NeurIPS competition
★ 7s4. Structured state space sequence models
★ 2.9kRAVE. Official implementation of the RAVE model: a Realtime Audio Variational autoEncoder
★ 1.8krse-course. Materials for The Alan Turing Institute's Research Software Engineering course
★ 265perceiver-pytorch. Implementation of Perceiver, General Perception with Iterative Attention, in Pytorch
★ 1.2kmusic-cocreation-tutorial. Start-to-finish tutorial for interactive music co-creation in PyTorch and Tensorflow.js
★ 110tutorial. 2021 ISMIR tutorial - music classification
★ 151soundata. Python library for downloading, loading & working with sound datasets
★ 357torch-audiomentations. Fast audio data augmentation in PyTorch. Inspired by audiomentations. Useful for deep learning.
★ 1.2kneural-waveshaping-synthesis. efficient neural audio synthesis in the waveform domain
★ 191mathematics-of-ml-course. Jupyter Notebook
★ 44meetups. Meetups data and resources archive for the London Audio & Music AI Meetup
★ 28audiomentations. A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
★ 2.3kMelSpecVAE. Variational Autoencoder in the mel-spectrogram domain for one-shot audio synthesis
★ 146uclser20. Code for the paper "Unsupervised Contrastive Learning of Sound Event Representations", ICASSP 2021.
★ 93SciencePlots. Matplotlib styles for scientific plotting
★ 9.1kCLMR. Official PyTorch implementation of Contrastive Learning of Musical Representations
★ 338vissl. VISSL is FAIR's library of extensible, modular and scalable components for SOTA Self-Supervised Learning with images.
★ 3.3kMultimodal-Transformers. List of papers and resources for multimodal transformers
★ 7audio-captioning-resources. A list of resources that can help in research for automated audio captioning
★ 34auraloss. Collection of audio-focused loss functions in PyTorch
★ 874experiment-impact-tracker. Python
★ 293music4all_contrib. Jupyter Notebook
★ 32audioset_tagging_cnn. Python
★ 1.8kchecklist. Beyond Accuracy: Behavioral Testing of NLP models with CheckList
★ 2.1ktag-based-music-retrieval. Python
★ 58webdataset. A high-performance Python-based I/O system for large (and small) deep learning problems, with strong support for PyTorch.
★ 3.2ktorchkit. Research boilerplate for PyTorch.
★ 150listening-moods. Accompanying code for our ISMIR 2020 paper on mood estimation.
★ 34matplotlib_for_papers. Handout for the tutorial "Creating publication-quality figures with matplotlib"
★ 2.2kmirdata. Python library for working with Music Information Retrieval datasets
★ 412ismir2020-metric-learning. ISMIR 2020 Tutorial for Metric Learning in MIR
★ 130spotify-musixmatch-data-collector. A Python module to generate large scale Music datasets using both Spotify and MusixMatch API's.
★ 43meetups. Slides and resources of the Vienna Deep Learning Meetup
★ 124nnAudio. Audio processing by using pytorch 1D convolution network
★ 1.1kWavAugment. A library for speech data augmentation in time-domain
★ 689audio-captioning-papers. A list of papers about audio captioning
★ 79sota-music-tagging-models. Python
★ 439pytorch-metric-learning. The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.
★ 6.3kPaperNotes. Important notes on scientific papers
★ 21pytorch-playground. Base pretrained models and datasets in pytorch (MNIST, SVHN, CIFAR10, CIFAR100, STL10, AlexNet, VGG16, VGG19, ResNet, Inception, SqueezeNet)
★ 2.7kdifferential-privacy-library. Diffprivlib: The IBM Differential Privacy Library
★ 1pvse. Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval (CVPR 2019)
★ 135vsepp. PyTorch Code for the paper "VSE++: Improving Visual-Semantic Embeddings with Hard Negatives"
★ 523mmvae. Multimodal Mixture-of-Experts VAE
★ 225Visual-Semantic-Embeddings-an-incomplete-list. A paper list of visual semantic embeddings and text-image retrieval.
★ 41visualbert. Code for the paper "VisualBERT: A Simple and Performant Baseline for Vision and Language"
★ 542lxmert. PyTorch code for EMNLP 2019 paper "LXMERT: Learning Cross-Modality Encoder Representations from Transformers".
★ 965VLP. Vision-Language Pre-training for Image Captioning and Question Answering
★ 420grounded-video-description. Video Grounding and Captioning
★ 331self-critical.pytorch. Unofficial pytorch implementation for Self-critical Sequence Training for Image Captioning. and others.
★ 1kNeuralBabyTalk. Pytorch code of for our CVPR 2018 paper "Neural Baby Talk"
★ 525show_attend_and_tell_pytorch. Pytorch implement Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
★ 95ImageCaptioning.pytorch. I decide to sync up this repo and self-critical.pytorch. (The old master is in old master branch for archive)
★ 1.5ka-PyTorch-Tutorial-to-Image-Captioning. Show, Attend, and Tell | a PyTorch Tutorial to Image Captioning
★ 2.9kvisual-semantic-embedding. Implementation of the image-sentence embedding method described in "Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models"
★ 428pytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31kaudioset-vggish-tensorflow-to-pytorch. Script for converting the pretrained VGGish model provided with AudioSet from TensorFlow to PyTorch, along with a basic smoke test.
★ 85remixmatch. Python
★ 132releasing-research-code. Tips for releasing research code in Machine Learning (with official NeurIPS 2020 recommendations)
★ 3kopenl3. OpenL3: Open-source deep audio and image embeddings
★ 600pytorch-template. PyTorch deep learning projects made easy.
★ 5.1kbert_score. BERT score for text generation
★ 1.9kVL-BERT. Code for ICLR 2020 paper "VL-BERT: Pre-training of Generic Visual-Linguistic Representations".
★ 742mmf. A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
★ 5.6kpytorch-normalizing-flows. Normalizing flows in PyTorch. Current intended use is education not production.
★ 917attention-primer. A demonstration of the attention mechanism with some toy experiments and explanations.
★ 108move. PyTorch code for training and evaluating MOVE, musically-motivated version embeddings
★ 50SoundSeek. SoundSeek is a new interface to search through your sound file libraries
★ 22Interactive_Tools. Interactive Tools for Machine Learning, Deep Learning and Math
★ 2.9kconfugue. Hierarchical configuration framework for Python
★ 21ctrl. Conditional Transformer Language Model for Controllable Generation
★ 1.9kPPLM. Plug and Play Language Model implementation. Allows to steer topic and attributes of GPT-2 models.
★ 1.2kSampleVAE. Multi-purpose tool for sound design and music production implemented in TensorFlow.
★ 180