This is your work, valued
MIR / Deep Learning Researcher @RECON-Labs-Inc & @neutune / M.S from @MALerLab , Sogang University / Seoul, South Korea
cog-musicgen-remixer. Music remixer based on MusicGen-Chord
★ 104cog-musicgen-chord. Chord conditioning implemented MusicGen
★ 68cog-musicgen-fine-tuner. This is a cog implementation of the fine-tuner for Meta's MusicGen
★ 55demucs_batch-multigpu. [Batching/MultiGPU/DataLoader Implemented] Code for the paper Hybrid Spectrogram and Waveform Source Separation
★ 24musicgen-remixer. process breakdown of MusicGen Remixer, by calling separate Replicate API calls and processing the outputs of the API calls.
★ 13all-in-one. cog implementation of All-In-One Music Structure Analyzer
★ 6stable-audio-cvae. Generative models for conditional audio generation
★ 4SOME. SOME: Singing-Oriented MIDI Extractor.
★ 2youtube-parallel-crawl-nordvpn. Shell
★ 2cog-AudioSR. Versatile audio super resolution (any -> 48kHz) with AudioSR.
★ 1cog-MusiConGen. A cog implementation of MusiConGen
★ 1streamable-stable-audio-open. Streaming the autoencoder in Stable Audio Open 1.0 for realtime continuous inference in MaxMSP/PureData
★ 15audio-based-lyrics-matching. Official Implementation of the paper "Leveraging Whisper Embeddings for Audio-based Lyrics Matching"
★ 17contrastive-singing-voices. Implementation of "Self-Supervised Contrastive Learning for Singing Voices"
★ 21terrain-diffusion. Procedural generation with diffusion models (SIGGRAPH '26)
★ 1.3knn_terrain. Latent Terrain - Dissecting the Latent Space of Neural Audio Autoencoders
★ 63TMPO. Official implementation of "TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment"
★ 17lfm-graph-preflexor. Python
★ 2graph-preflexor-grpo. Jupyter Notebook
★ 10lexguard-mcp. 일반인들이 AI를 통해 법률 정보를 쉽게 조회할 수 있는 MCP 서버. 법령 검색, 조문 조회, 판례 검색 등 159개 API 지원.
★ 125PaperBanana. PaperBanana: Automating Academic Illustration For AI Scientists
★ 6.9kvqgan-clip. Jupyter Notebook
★ 354AI-Song-Cover-RVC. All-in-one RVC song cover toolkit for Google Colab: pull audio from YouTube, separate vocals, train a model, and run inference.
★ 1.2kmoebius-web. Run the Moebius inpainting model in the browser
★ 78fugu. Shell
★ 940MERIT. Python
★ 30CIR. Python
★ 4tribeV2_ViralAnalyser. Python
★ 200post--feature-visualization. Feature Visualization
★ 137tribev2. This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction
★ 3.1kbrainmagick. Training and evaluation pipeline for MEG and EEG brain signal encoding and decoding using deep learning. Code for our paper "Decoding speech perception from non-invasive brain recordings" published in Nature Machine Intelligence, 2023.
★ 475mira. MiRA (Music Replication Assessment) tool is a model-independent open evaluation method based on four diverse audio music similarity metrics to assess exact data replication of the training set.
★ 35localsend. An open-source cross-platform alternative to AirDrop
★ 87kponytail. Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
★ 93kchecklist.skill. no more HANDOFF documents
★ 1Omni-Embed-Audio. [ACL 2026 Oral] Official code, UIQ benchmark, and pretrained checkpoints for "Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval" (Yoo et al., Sogang University)
★ 11CoverHunterMPS. Fork of Liu Feng's CoverHunter to run on a single computer, plus more features and documentation.
★ 26odysseus. Self-hosted AI workspace.
★ 84kembracing-cacophony. Investigating data augmentation in music source separation
★ 4timbral_models. Python scripts for modelling timbral attributes
★ 93claw-hwp. Read, create & edit Korean Hangul Word Processor (.hwp / .hwpx) documents in Claude — Agent Skill built on rhwp WASM, with built-in browser preview. Runs locally, no Hancom Office, no cloud.
★ 38AI-Research-SKILLs. Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
★ 11kSongEcho_ICLR2026. Official code for SongEcho
★ 64ikp. IKP: Incompressive Knowledge Probes
★ 101MGPHot-audio. This repository enables the community to use MGPHot in further research without redistributing restricted files.
★ 14surge-python. This repo contains examples of how to use surgepy, Python bindings for the Surge synthesizer.
★ 36DeepFilterNet. Noise supression using deep filtering
★ 4.5kvoid-model. Python
★ 1.9kkeymyna. Official repository of "Myna-Style Contrastive Pre-Training Improves Music Audio Key Detection"
★ 8cssc. The implement of the Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
★ 3Amphion. Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
★ 10kcodicodec. Encode and decode audio samples to/from continuous and discrete compressed representations!
★ 121flowerjuce. hugo flowers' collection of (generative) digital musical instruments built with JUCE.
★ 10music-llm-lightning. Lightning version of Music LLM implementation
★ 12Neural-MRI. Model Resonance Imaging — visualize LLM internals like a brain MRI
★ 15newton. An open-source, GPU-accelerated physics simulation engine built upon NVIDIA Warp, specifically targeting roboticists and simulation researchers.
★ 5.3kjoint-apt-epr. Bridging Piano Transcription and Rendering via Disentangled Score Content and Style (ICLR 2026 accepted)
★ 10midi-function-alignment. Python
★ 14pkspell. Predict the correct pitch spelling and key signatures given a sequence of midi notes by using a deep-learning approach.
★ 18WarpFusion. WarpFusion
★ 1kopendataloader-pdf. PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
★ 28kseoul-world-model. Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
★ 622autoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kMEGAMI. Accompanying repository for the paper "Automatic Music Mixing Using a Generative Model of Effect Embeddings"
★ 41ComfyUI-Color-Matcher. ComfyUI-Color-Matcher
★ 1siljangnim. AI-powered real-time graphics creation tool.
★ 23sample_100. A dataset of Hip Hop samples for Music Information Retrieval research
★ 11sampleid. Code for the paper “Automatic Music Sample Identification with Multi-Track Contrastive Learning”.
★ 25clews. Python
★ 34tria. Python
★ 11Sight-Singing-Vocal-Data.
★ 4SOFA. SOFA: Singing-Oriented Forced Aligner
★ 228NeuralSampleID. An automatic sample identification (ASID) system using a contrastively trained GNN encoder.
★ 17OpenEWLD. A Public Domain Leadsheet Dataset
★ 38flash-foley. Python
★ 10Nightlight-Game-Launcher. This program is made to Bypass Games Launcher!
★ 194DragGAN. Official Code for DragGAN (SIGGRAPH 2023)
★ 36kmyna. Official repository of Myna: Masking-Based Contrastive Learning of Musical Representations
★ 17audio-flamingo. PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
★ 1.2knnAudio. Audio processing by using pytorch 1D convolution network
★ 1.1kmusic-text-representation-pp. Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval (TTMR++) [ICASSP24]
★ 43Retrieval-based-Voice-Conversion-WebUI. Easily train a good VC model with voice data <= 10 mins!
★ 37kcog-musicgen-chord-windows. A version of cog-musicgen-chord, heavily modified to work natively on Windows, with Cog, Linux and Docker dependencies removed
★ 1pogocache. Fast caching software with a focus on low latency and cpu efficiency.
★ 2.5kMGE-LDM. Official implementation of the paper MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
★ 20Paper2Code. Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
★ 4.8kFx-Encoder_PlusPlus. "Fx-Encoder++: Extracting Instrument-wise Audio Effect Representations from Mixtures"
★ 53FM4Music. The official GitHub page for the survey paper "Foundation Models for Music: A Survey".
★ 224Tonmeister-Grundlagen. Grundlagenskript fuer Tonmeisterstudenten (2000)
★ 6legato. Official codebase for paper "LEGATO: Large-Scale End-to-End Generalizable Approach to Typeset OMR".
★ 65magenta. YOLO'd Magenta for Colab
★ 6pillumi-khu. Pillumi: a mobile system for searching how to take the pill through image recognition (Korean only)
★ 6SOTA-Deep-Anomaly-Detection. List of implementation of SOTA deep anomaly detection methods
★ 110Music-Source-Separation-Training. Repository for training models for music source separation.
★ 1.5kfriendly-stable-audio-tools. Refactored / updated version of `stable-audio-tools` which is an open-source code for audio/music generative models originally by Stability AI.
★ 218cog-sdxl-lora. Python
★ 23oemer. End-to-end Optical Music Recognition (OMR) system. Transcribe phone-taken music sheet image into MusicXML, which can be edited and converted to MIDI.
★ 779cursor-talk-to-figma-mcp. TalkToFigma: MCP integration between AI Agent (Cursor, Claude Code, Codex) and Figma, allowing Agentic AI to communicate with Figma for reading designs and modifying them programmatically.
★ 6.9krewriting. Rewriting a Deep Generative Model, ECCV 2020 (oral). Interactive tool to directly edit the rules of a GAN to synthesize scenes with objects added, removed, or altered. Change StyleGANv2 to make extravagant eyebrows, or horses wearing hats.
★ 535network-bending. Manipulating the inner representations of StyleGAN2
★ 108three-geospatial. Geospatial Rendering in Three.js
★ 1.6kmagenta-js. Magenta.js: Music and Art Generation with Machine Learning in the browser
★ 2.1kVAE_jam. JavaScript
★ 1AudioMAE. This repo hosts the code and models of "Masked Autoencoders that Listen".
★ 673mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4kIC-Light. More relighting!
★ 8.5kOmost. Your image is almost there!
★ 7.6kControlNet. Let us control diffusion models!
★ 34kFramePack. Lets make video diffusion practical!
★ 17ktiny-audio-diffusion. A repository for generating and training short audio samples with unconditional waveform diffusion on accessible consumer hardware (<2GB VRAM GPU)
★ 189Image-Morphing. Automating Image Morphing using Structural Similarity on a Halfway Domain (Siggraph 2014)
★ 118IguanaTex. A PowerPoint add-in to insert LaTeX equations into PowerPoint presentations on Windows and Mac
★ 1.4k1Prompt1Story. 🔥ICLR 2025 (Spotlight) One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
★ 320latex_resume_template_kor. [이력서/경력기술서] Beautiful CV/resume template for Koreans, in LaTeX
★ 474taehv. Tiny AutoEncoder for Hunyuan Video (and other video models)
★ 451taesd. Tiny AutoEncoder for Stable Diffusion (and other image models)
★ 958hwp-mcp. mcp for handling hwp
★ 265thera. [TMLR 2025] Thera: Aliasing-Free Arbitrary-Scale Super-Resolution with Neural Heat Fields
★ 865labelme. Image annotation with Python. Supports polygon, rectangle, circle, line, point, and AI-assisted annotation.
★ 16kmelonix. [WIP] Pitch correction application written using ImGui and OpenGL 3
★ 68midi-pitch. Compare MIDI with Vocal Pitch
★ 26bindsnet. Simulation of spiking neural networks (SNNs) using PyTorch.
★ 1.7ktokensynth. The official implementation of TokenSynth (ICASSP 2025)
★ 93LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kWavetableCVAE. Python
★ 4onset_db. Onset data set which can be used to tune/evaluate onset detection algorithms.
★ 66SOME. SOME: Singing-Oriented MIDI Extractor.
★ 702rgbds. Rednex Game Boy Development System - An assembly toolchain for the Nintendo Game Boy and Game Boy Color
★ 1.6kPink-Trombone. A programmable version of Neil Thapen's Pink Trombone
★ 211YuE. YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
★ 6.4kmusic_translation_eval. Python
★ 1micro_diffusion. Official repository for our work on micro-budget training of large-scale diffusion models.
★ 1.6kDeepSeek-R1.
★ 92kmixbox. Mixbox is a library for natural color mixing based on real pigments.
★ 3.5kprodigy. The Prodigy optimizer and its variants for training neural networks.
★ 470gs-quant. Python toolkit for quantitative finance
★ 12kjltr-alignment. Audio-to-score alignment with human-labeled repeats
★ 7nmt. Python
★ 15pixel2style2pixel. Official Implementation for "Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation" (CVPR 2021) presenting the pixel2style2pixel (pSp) framework
★ 3.3kYourMT3. multi-task and multi-track music transcription for everyone
★ 242WavAugment. A library for speech data augmentation in time-domain
★ 689stable-audio-metrics. Metrics for evaluating music and audio generative models – with a focus on long-form, full-band, and stereo generations.
★ 300DeepSeek-Coder-V2. DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
★ 6.9kSejongMusic. Official Repository of Six Dragons Fly Again (ISMIR 2024)
★ 15llark. Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.
★ 383awesome-pretrained-stylegan2. A collection of pre-trained StyleGAN 2 models to download
★ 1.3kmaua-stylegan2. This is the repo for my experiments with StyleGAN2. There are many like it, but this one is mine. Contains code for the paper Audio-reactive Latent Interpolations with StyleGAN.
★ 179music2latent. Encode and decode audio samples to/from compressed latent representations!
★ 267ismir2024_tutorial_demo. Jupyter Notebook
★ 18mochi. The best OSS video generation models, created by Genmo
★ 3.7kRAVE-Latent-Diffusion. Generate new latent codes for RAVE with Denoising Diffusion models.
★ 186RAVE. Official implementation of the RAVE model: a Realtime Audio Variational autoEncoder
★ 1.8kAutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186klucid-sonic-dreams. Python
★ 774vqgan-training. Train VAE like a boss
★ 313TouchDiffusion. TouchDesigner implementation for real-time Stable Diffusion interactive generation with StreamDiffusion.
★ 291encodec. State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
★ 4krq-vae-transformer. The official implementation of Autoregressive Image Generation using Residual Quantization (CVPR '22)
★ 1kkiwipiepy. Python API for Kiwi
★ 397VQGAN-pytorch. Pytorch implementation of VQGAN (Taming Transformers for High-Resolution Image Synthesis) (https://arxiv.org/pdf/2012.09841.pdf)
★ 552taming-transformers-feat-LucidrainsVQ. Taming Transformers for High-Resolution Image Synthesis
★ 2vector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4kVQGAN-CLIP. Just playing with getting VQGAN+CLIP running locally, rather than having to use colab.
★ 2.6kstargan2. StarGAN2 for practice
★ 95graphmuse. A Graph Deep Learning Library for Music.
★ 113partitura. A python package for handling modern staff notation of music
★ 368fadtk. A simple library for Fréchet Audio Distance (FAD) calculation
★ 266ocr-vqgan. OCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers
★ 85MusiConGen. Python
★ 88genmusic_demo_list. a list of demo websites for automatic music generation research
★ 794taming-transformers. Taming Transformers for High-Resolution Image Synthesis
★ 6.5kEnCodec_Trainer. Python
★ 67FacePose_pytorch. 🔥🔥The pytorch implement of the head pose estimation(yaw,roll,pitch) and emotion detection with SOTA performance in real time.Easy to deploy, easy to use, and high accuracy.Solve all problems of face detection at one time.(极简,极快,高效是我们的宗旨)
★ 764AutoRAG. AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
★ 5knested-music-transformer-demo. SCSS
★ 1CrossViT. Official implementation of CrossViT. https://arxiv.org/abs/2103.14899
★ 417AIW. Alice in Wonderland code base for experiments and raw experiments data
★ 129human. Human: AI-powered 3D Face Detection & Rotation Tracking, Face Description & Recognition, Body Pose Tracking, 3D Hand & Finger Tracking, Iris Analysis, Age & Gender & Emotion Prediction, Gaze Tracking, Gesture Recognition
★ 3.2kControledAnimateDiff. Controlnet extension of AnimateDiff.
★ 54AI-Papers-of-the-Week. 🔥Highlighting the top ML papers every week.
★ 13kencodecmae. Codebase for the paper 'EncodecMAE: Leveraging neural codecs for universal audio representation learning'
★ 101AutoCrawler. Google, Naver multiprocess image web crawler (Selenium)
★ 1.7kLabanotation. Image Data Set
★ 5tuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
★ 30kmusicgen-dreamboothing. Fine-tune your own MusicGen with LoRA
★ 162direct-preference-optimization. Reference implementation for DPO (Direct Preference Optimization)
★ 2.9kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 59ISMIR-2023-Papers. ISMIR 2023 Papers: A complete collection of influential and exciting research papers from the ISMIR 2023 conference.
★ 105stable-audio-tools. Generative models for conditional audio generation
★ 3.8kmtg-jamendo-dataset. Metadata, scripts and baselines for the MTG-Jamendo dataset
★ 400cog-musicgen-remixer. Music remixer based on MusicGen-Chord
★ 104airgen. Official source codes of airsep
★ 39PokemonRedExperiments. Playing Pokemon Red with Reinforcement Learning
★ 7.9kmir_eval. Evaluation functions for music/audio information retrieval/signal processing algorithms.
★ 709gca. Game Boy Coding Adventure companion repository
★ 148demucs. Code for the paper Hybrid Spectrogram and Waveform Source Separation
★ 10kros_face. C++
★ 1.4klora. Using Low-rank adaptation to quickly fine-tune diffusion models.
★ 7.5kSALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kPyTSMod. An open-source Python library for audio time-scale modification.
★ 234all-in-one. All-In-One Music Structure Analyzer
★ 806cog-musicgen-fine-tuner. This is a cog implementation of the fine-tuner for Meta's MusicGen
★ 55musicgen-chord. MusicGen conditioned with chord progression.
★ 11KoAlpaca. KoAlpaca: 한국어 명령어를 이해하는 오픈소스 언어모델 (KoAlpaca: An open-source language model to understand Korean instructions)
★ 1.6ktempo-cnn. Framework for estimating temporal properties of music tracks.
★ 107cog. Containers for machine learning
★ 9.5kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kdemucs_batch-multigpu. [Batching/MultiGPU/DataLoader Implemented] Code for the paper Hybrid Spectrogram and Waveform Source Separation
★ 24