This is your work, valued
Main developer of ESPnet / COO @ Human Dataware Lab. Co., Ltd. / Postdoctoral researcher @ Nagoya University
ParallelWaveGAN. Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch
★ 1.6kPytorchWaveNetVocoder. WaveNet-Vocoder implementation with pytorch.
★ 301LibriTTSLabel. Alignment files of LibriTTS.
★ 70INTERSPEECH19_TUTORIAL. Interspeech 2019 tutorial materials
★ 49WaveNetVocoderSamples. WaveNet Vocoder Samples
★ 23Taco2withBERT. Tacotron2 with BERT examples
★ 10webMUSHRA. a MUSHRA compliant web audio API based experiment software
★ 10espnet. End-to-End Speech Processing Toolkit
★ 8NonARSeq2SeqVC. Non-autoregressive sequence-to-sequence voice conversion
★ 6dotfiles. My dotfiles (ghostty + tmux + neovim)
★ 6icassp2022-streaming-vc.
★ 4VCTKCorpusFullContextLabel. Full context label for VCTK Corpus.
★ 4asj-espnet2-tutorial. ESPnet2解説原稿付録
★ 2stiv. STIV - Simple Terminal Image Viewer
★ 2alpagym. AlpaGym is a reinforcement-learning framework for end-to-end autonomous-driving policies.
★ 125asrlance. 「あすらんす」は音声認識性能を比較評価するツールです.音声ファイルパスと正解文を実行時に入力することで,認識精度(マイクロCER),処理にかかった時間,CPU使用率を結果として出力します.
★ 4LongCat-AudioDiT. Python
★ 557opensmile-python. Python package for openSMILE
★ 329prompt-review. Python
★ 359paperclip. The open-source app everyone uses to manage agents at work
★ 75kjoserfc. Implementations of JOSE RFCs in Python
★ 171MeanVC. A Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
★ 298alpasim. AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies
★ 1.1kDetZero. [ICCV 2023] DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds
★ 437Sidon. Training code and dataset cleasing with Sidon
★ 1723D-Speaker. A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization
★ 3.1kvaluecell. ValueCell is a community-driven, multi-agent platform for financial applications.
★ 11kdcase-asd-toolkit. Python
★ 14uroman. Universal Romanizer that can convert any unicode script to roman (latin) script
★ 250claudecode.nvim. 🧩 Claude Code Neovim IDE Extension
★ 3kspandsp. SpanDSP is a low-level signal processing library that modulates and demodulates signals commonly used in telephony, such as the "noise" generated by a fax modem or DTMF touchpad.
★ 203dia. A TTS model capable of generating ultra-realistic dialogue in one pass.
★ 19kgptme. Your agent in your terminal, equipped with local tools: writes code, uses the terminal, browses the web. Make your own persistent autonomous agent on top!
★ 4.4kpoint-cloud-registration. Python
★ 84pyopenjtalk-plus. pyopenjtalk-plus: A Python wrapper for OpenJTalk with additional improvements
★ 58ollama_load_balancer. Autonomous Rust utility that load balances multiple https://ollama.com/ servers
★ 22OmniParser. A simple screen parsing tool towards pure vision based GUI agent
★ 25kparler-tts. Inference and training library for high-quality TTS models.
★ 5.6kYuE. YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
★ 6.4kneoconf.nvim. 💼 Neovim plugin to manage global and project-local settings
★ 960kickstart.nvim. A launch point for your personal nvim configuration
★ 31kawesome-mcp-servers. A collection of MCP servers.
★ 92kLotus. Official implementation of "Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction"
★ 819LightRAG. [EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation"
★ 38kvoice-changer. リアルタイムボイスチェンジャー Realtime Voice Changer
★ 21khubert. HuBERT content encoders for: A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion
★ 406insanely-fast-whisper. Jupyter Notebook
★ 13kwriter-framework. No-code in the front, Python in the back. An open-source framework for creating data apps.
★ 1.4klitellm. The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
★ 55kvits2_pytorch. unofficial vits2-TTS implementation in pytorch
★ 548aicommits. A CLI that writes your git commit messages for you with AI
★ 9.1kdelta. A syntax-highlighting pager for git, diff, grep, rg --json, and blame output
★ 32knaturalspeech2-pytorch. Implementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in Pytorch
★ 1.3kgpt-pilot. The first real AI developer
★ 34kmiipher. Unofficial implementation of miipher
★ 137mlx. MLX: An array framework for Apple silicon
★ 28kfunctionary. Chat language model that can use tools and interpret the results
★ 1.6kllama-cpp-python. Python bindings for llama.cpp
★ 11kstable-ts. Transcription, forced alignment, and audio indexing with OpenAI's Whisper
★ 2.3kSpeechMOS. Easy-to-Use Speech MOS predictors
★ 362xtreme1. Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
★ 1.2kJVSCorpusF0Range. JVSコーパスのF0探索範囲の詳細な調査結果
★ 4LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kSegment-Any-Anomaly. Official implementation of "Segment Any Anomaly without Training via Hybrid Prompt Regularization (SAA+)".
★ 843openinterpreter. A coding agent for open models like Kimi K3
★ 67kstack-chan. A JavaScript-driven M5Stack-embedded super-kawaii robot.
★ 1.6kfacefusion. Industry leading face manipulation platform
★ 29klocal-llm-function-calling. A tool for generating function arguments and choosing what function to call with local LLMs
★ 435roslibpy. Python ROS Bridge library
★ 333openai-cookbook. Examples and guides for using the OpenAI API
★ 75kXPhoneBERT. XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech (INTERSPEECH 2023)
★ 355ChatWaifu. Combined ChatGPT with Moegoe TTS to create a Chatting Waifu
★ 835ImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1ktslearn. The machine learning toolkit for time series analysis in Python
★ 3.2kspe-dss. Speech Parameter Estimation Using Differentiable Speech Synthesizer
★ 43octo.nvim. Edit and review GitHub issues and pull requests from the comfort of your favorite editor
★ 3.3ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kVTuber-Python-Unity. An Implementation of VTuber (Both 3D and Live2D) using Python and Unity. Providing face movement tracking, eye blinking detection, iris detection and tracking and mouth movement tracking using CPU only.
★ 554AutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kalt-ime-ahk. AutoHotkey
★ 647ChatGPT.nvim. ChatGPT Neovim Plugin: Effortless Natural Language Generation with OpenAI's ChatGPT API
★ 4kalqalign. multilingual speech aligner
★ 78supabase. The Postgres development platform. Supabase gives you a dedicated Postgres database to build your web, mobile, and AI applications.
★ 107kedgeyolo. an edge-real-time anchor-free object detector with decent performance
★ 527pystoi. Python implementation of the Short Term Objective Intelligibility measure
★ 359Weighted-Boxes-Fusion. Set of methods to ensemble boxes from different object detection models, including implementation of "Weighted boxes fusion (WBF)" method.
★ 1.8kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kwhisper.cpp. Port of OpenAI's Whisper model in C/C++
★ 52ktorch-yin. Yin pitch estimator in PyTorch
★ 119ReazonSpeech. Massive open Japanese speech corpus
★ 390nara_wpe. Different implementations of "Weighted Prediction Error" for speech dereverberation
★ 570nsxiv. Read-only mirror of Neo Simple X Image Viewer
★ 838cuml. cuML - RAPIDS Machine Learning Library
★ 5.2kLoFTR. Code for "LoFTR: Detector-Free Local Feature Matching with Transformers", CVPR 2021, T-PAMI 2022
★ 2.9kRoadDamageDetector. Jupyter Notebook
★ 977NCEPLIBS-nemsio. The NCEPLIBS-nemsio library and utilities perform I/O for the NCEP models using NOAA Environmental Modeling System (NEMS) format.
★ 3ProDiff. PyTorch Implementation of ProDiff (ACM-MM'22) with a Extremely-Fast diffusion speech synthesis pipeline
★ 432Real-ESRGAN. Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
★ 36kstable-diffusion. A latent text-to-image diffusion model
★ 73kflamingo-pytorch. Implementation of 🦩 Flamingo, state-of-the-art few-shot visual question answering attention net out of Deepmind, in Pytorch
★ 1.3kLAV. (CVPR 2022) A minimalist, mapless, end-to-end self-driving stack for joint perception, prediction, planning and control.
★ 442DeepPhonemizer. Grapheme to phoneme conversion with deep learning.
★ 433anomalib. An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference.
★ 6kswitch-scores. PDF Repository of switch score sheets.
★ 1.2kpydtk. A Python toolkit for managing, retrieving and processing data.
★ 14onnx-model-encrypt-sample. ONNXモデルをpyca/cryptographyを用いて暗号化/復号化するサンプル
★ 16cryptography. cryptography is a package designed to expose cryptographic primitives and recipes to Python developers.
★ 7.7kUTMOS22. UT-Sarulab MOS prediction system using SSL models
★ 309patchcore-inspection. Python
★ 1.4kginza. A Japanese NLP Library using spaCy as framework based on Universal Dependencies
★ 865dog-dataset. Python
★ 47pingu. 🐧ping command but with pingu
★ 2.1kwhylogs. An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
★ 2.8klm_build. Adapting your own Language Model for Kaldi
★ 63OSSGAN. Official implementation of OSSGAN [CVPR 2022]
★ 21diffsptk. A differentiable version of SPTK
★ 201PlotNeuralNet. Latex code for making neural networks diagrams
★ 25knvim-spectre. Find the enemy and replace them with dark power.
★ 2.4kdashboard-nvim. vim dashboard
★ 2.9kquickfix-reflector.vim. Change code right in the quickfix window
★ 366espnet_onnx. Onnx wrapper for espnet infrernce model
★ 169neural_prophet. NeuralProphet: A simple forecasting package
★ 4.3kadamatch. Python
★ 68kubric. A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
★ 2.8kpytorch_GAN_zoo. A mix of GAN implementations including progressive growing
★ 1.6kdate. Python
★ 29textlesslib. Library for Textless Spoken Language Processing
★ 559webvtt-py. Read, write, convert and segment WebVTT caption files in Python.
★ 2333detr. Code & Models for 3DETR - an End-to-end transformer model for 3D object detection
★ 712github-sponsors-tax. GitHub Sponsorsの確定申告手順
★ 176exkaldi-rt. An online speech recognition extension toolkit of Kaldi
★ 55searchbox.nvim. Start your search from a more comfortable place, say the upper right corner?
★ 358wenet. Production First and Production Ready End-to-End Speech Recognition Toolkit
★ 5.2ktorch-points3d. Pytorch framework for doing deep learning on point clouds.
★ 2.7kcupoch. Robotics with GPU computing
★ 1.1kprobreg. Python package for point cloud registration using probabilistic model (Coherent Point Drift, GMMReg, SVR, GMMTree, FilterReg, Bayesian CPD)
★ 987Point-BERT. [CVPR 2022] Pre-Training 3D Point Cloud Transformers with Masked Point Modeling
★ 695silero-vad. Silero VAD: pre-trained enterprise-grade Voice Activity Detector
★ 9.8kInformer2020. The GitHub repository for the paper "Informer" accepted by AAAI 2021.
★ 6.5kmdloader. Massdrop Firmware Loader - for CTRL / ALT / SHIFT / Rocketeer keyboards
★ 434Autoformer. About Code release for "Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting" (NeurIPS 2021), https://arxiv.org/abs/2106.13008
★ 2.5kGMA. Learning to Estimate Hidden Motions with Global Motion Aggregation (ICCV 2021)
★ 337RAFT. Python
★ 4.1kMultilingual_Text_to_Speech. An implementation of Tacotron 2 that supports multilingual experiments with parameter-sharing, code-switching, and voice cloning.
★ 844fast_bss_eval. A fast implementation of bss_eval metrics for blind source separation
★ 149Image-Super-Resolution-via-Iterative-Refinement. Unofficial implementation of Image Super-Resolution via Iterative Refinement by Pytorch
★ 3.9kVQMIVC. Official implementation of VQMIVC: One-shot (any-to-any) Voice Conversion @ Interspeech 2021 + Online playing demo!
★ 361textract. extract text from any document. no muss. no fuss.
★ 4.7kmahalanobis_3d_multi_object_tracking. [NeurIPS Workshop 2019] Official code of the paper "Probabilistic 3D Multi-Object Tracking for Autonomous Driving." First Place of the First NuScenes Tracking Challenge in the AI Driving Olympics Workshop of NeurIPS.
★ 400vim-sh-heredoc-highlighting. Syntax highlighting for code snippets in shell heredocs
★ 11pyannote-audio. Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
★ 10kgrammarly. Grammarly for VS Code
★ 1.6kwilder.nvim. A more adventurous wildmenu
★ 1.5kclean-text. 🧹 Python package for text cleaning
★ 1ksakai. Sakai is a freely available, feature-rich technology solution for learning, teaching, research and collaboration. Sakai is an open source software suite developed by a diverse and global adopter community.
★ 1.2ktdmelodic. A Japanese accent dictionary generator
★ 126css10. CSS10: A Collection of Single Speaker Speech Datasets for 10 Languages
★ 490mmdetection3d. OpenMMLab's next-generation platform for general 3D object detection.
★ 6.5kscenario_simulator_v2. A scenario-based simulation framework for Autoware
★ 154cvat. Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs.
★ 16kClassyVision. An end-to-end PyTorch framework for image and video classification
★ 1.6kawesome-diarization. A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.
★ 1.9kassem-vc. Official Code for Assem-VC @ICASSP2022
★ 269CenterFusion. CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection
★ 634ppg-vc. PPG-Based Voice Conversion
★ 348open-speech-corpora. 💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1.4kdeep_sort_pytorch. MOT using deepsort and yolov3 with pytorch
★ 3kpdfminer.six. Community maintained fork of pdfminer - we fathom PDF
★ 7kdvc. 🦉 Data Versioning and ML Experiments
★ 16kroc-star. Loss function which directly targets ROC-AUC
★ 251Kazakh_TTS. An expanded version of the previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In KazakhTTS2, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two to five (three females and two males), and the topic coverage has been diversified.
★ 157s3prl. Self-Supervised Speech Pre-training and Representation Learning Toolkit
★ 2.6kspeech-synthesis-paper. List of speech synthesis papers.
★ 1.1kthe-art-of-command-line. Master the command line, in one page
★ 162kpytorch-softdtw-cuda. Fast CUDA implementation of (differentiable) soft dynamic time warping for PyTorch
★ 734vim-pydocstring. Generate Python docstring to your Python source code.
★ 341ConformerSED. Python
★ 31TTS. 🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 46kbyol-pytorch. Usable Implementation of "Bootstrap Your Own Latent" self-supervised learning, from Deepmind, in Pytorch
★ 1.9kaudiomentations. A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
★ 2.3klhotse. Tools for handling multimodal data in machine learning projects.
★ 1.1kdiff-so-fancy. Make your diffs human readable for improved code quality and faster defect detection. :tada:
★ 18kquick-look-plugins. List of useful Quick Look plugins for developers
★ 19ktorchcrepe. Pytorch implementation of the CREPE pitch tracker
★ 523StreamingTransformer. Python
★ 277homebridge. HomeKit support for the impatient.
★ 25kwemux. Multi-User Tmux Made Easy
★ 3.7kvim-suda. 🥪 An alternative sudo.vim for Vim and Neovim, limited support sudo in Windows
★ 847scitldr. Python
★ 759LaboroTVSpeech. Shell
★ 90gluonts. Probabilistic time series modeling in Python
★ 5.2kCenterPoint. Python
★ 2.2kdetectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kSiamMask. [CVPR19/TPAMI23] SiamMask: A Framework for Fast Online Object Tracking and Segmentation
★ 3.6ktextrank. TextRank implementation for Python 3.
★ 1.3kgtn. Automatic differentiation with weighted finite-state transducers.
★ 453evaluation. A custom Rouge wrapper for summary evaluation.
★ 53darts. A python library for user-friendly forecasting and anomaly detection on time series.
★ 9.5kprophet. Tool for producing high quality forecasts for time series data that has multiple seasonality with linear or non-linear growth.
★ 20khugo-bearblog. 🧸 A Hugo theme based on »Bear Blog«. Free, no-nonsense, super-fast blogging. This theme now includes a dark color scheme to support dark mode 🦉 ⬛️!
★ 1.5kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kcontiguous_pytorch_params. Accelerate training by storing parameters in one contiguous chunk of memory.
★ 294fugashi. A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.
★ 533MatchSum. Code for ACL 2020 paper: "Extractive Summarization as Text Matching"
★ 520crank. A toolkit for non-parallel voice conversion based on vector-quantized variational autoencoder
★ 171AMFM_decompy. Package containing the tools necessary for decomposing a speech signal into its modulated components (also known as AM-FM decomposition). Includes the algorithms of the QHM family and the YAAPT pitch tracker.
★ 91sumeval. Well tested & Multi-language evaluation framework for text summarization.
★ 626blink. Blink Mobile Shell for iOS (Mosh based)
★ 6.9kdurian-pytorch. Implementation of "Duration Informed Attention Network for Multimodal Synthesis" paper in PyTorch.
★ 184pytorch-lightning. Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
★ 31kphonemizer. Simple text to phones converter for multiple languages
★ 1.6kTensorFlowTTS. :stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German and Easy to adapt for other languages)
★ 4kglow-tts. A Generative Flow for Text-to-Speech via Monotonic Alignment Search
★ 712onssen. An open-source speech separation and enhancement library
★ 214