This is your work, valued
a2s-transformer. A Transformer approach for polyphonic Audio-to-Score (A2S) transcription (ICASSP 2024)
★ 14ScoreGenerator. Music score generator for OMR systems
★ 4ismir-lbd-2021. Late-fusing OMR and A2S predictions using the Smith-Waterman local alignment algorithm (ISMIR 2021)
★ 3omr_a2s_multimodal_transformer. Python
★ 3late-fusion-music-transcription. Late-fusing OMR and A2S predictions using four different algorithms
★ 2real-a2s. On the use of synthetic music for transcribing real saxophone audio recordings (INTERSPEECH 2023)
★ 2smc-2022. Transfer learning between OMR and A2S models
★ 2early-vs-late-multimodal-music-transcription. Python
★ 1ssl-symbol-classification. Unlabelled self-supervised learning symbol classification
★ 1paperclip. The open-source app everyone uses to manage agents at work
★ 75kclaude-howto. A visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
★ 41kautoskills. One command. Your entire AI skill stack. Installed.
★ 6.6klegalize-pipeline. Legislación española como repositorio Git — cada ley es un fichero Markdown, cada reforma un commit
★ 48spec-kit. 💫 Toolkit to help you get started with Spec-Driven Development
★ 125kobsidian-skills. Agent skills for Obsidian. Teach your agent to use Obsidian CLI and open formats including Markdown, Bases, JSON Canvas.
★ 44kKanvas. Manage projects visually with humans and AI agents using Obsidian Canvas.
★ 179vietnamese-htr. Vietnamese handwritten text recognition system
★ 10CombLM. Python
★ 6paragraph_handwriting_imitation_ldm. Python
★ 4sscd-copy-detection. Open source implementation of "A Self-Supervised Descriptor for Image Copy Detection" (SSCD).
★ 419docext. An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
★ 2kISMIR-2024-SYNTHETIC2REAL-OMR. Source-Free Synthetic-to-Real Domain Adaptation for Optical Music Recognition (ISMIR 2024)
★ 2vision-transformers-cifar10. Let's train vision transformers (ViT) for cifar 10 / cifar 100!
★ 718power-efficient-nn.github.io. Python
★ 15blog. Data, code, and scripts for the analysis in the Mode blog.
★ 122audioflow. Python
★ 131face-recognition-liveness. Face detection and recognition + liveness detection and spoofing attack recognition using onnxruntime. Includes an easy-to-use Flask API and Dockerfile.
★ 68ocr-chunker-crud-vectordb. Python
★ 1Believable-Human-Writing-Conversor. A script that transforms any text to human-like writing.
★ 15Ollama-OCR. Jupyter Notebook
★ 2.7kebook2audiobook. Generate audiobooks from e-books, voice cloning & 1158+ languages!
★ 20kpdf-to-podcast. Transform PDFs into AI podcasts for engaging on-the-go audio content.
★ 861computer-vision-and-deep-learning-course. This repository contains both a collection of Jupyter Notebooks as well as other resources (e.g. presentations, links, ...) that are going to be used during the "Second quarter university extension courses" that the University of Oviedo is going to teach (online).
★ 26vlms-zero-to-hero. This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
★ 1.2krecommenders. Best Practices on Recommendation Systems
★ 22kannotated-s4. Implementation of https://srush.github.io/annotated-s4
★ 519annotated-mamba. Annotated version of the Mamba paper
★ 502CLAP. Contrastive Language-Audio Pretraining
★ 2.2kmarkitdown. Python tool for converting files and office documents to Markdown.
★ 171kpdftext. Extract structured text from pdfs quickly
★ 713super-gradients. Easily train or fine-tune SOTA computer vision models with one open source training library. The home of Yolo-NAS.
★ 5.1kml-retreat. Machine Learning Journal for Intermediate to Advanced Topics.
★ 2.4kOCRDatasets. A collection of OCR-related datasets
★ 222void. TypeScript
★ 29kattentions. PyTorch implementation of some attentions for Deep Learning Researchers.
★ 548xlstm. Official repository of the xLSTM.
★ 2.2kLEAD. [CVPR 2024] The official repository of our paper "LEAD: Learning Decomposition for Source-free Universal Domain Adaptation"
★ 43awesome-test-time-adaptation. Collection of awesome test-time (domain/batch/instance) adaptation methods
★ 1.3kzero_to_gpt. Go from no deep learning knowledge to implementing GPT.
★ 1.3kGemini. The open source implementation of Gemini, the model that will "eclipse ChatGPT" by Google
★ 466LLMs-from-scratch. Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
★ 100kTextRecognitionDataGenerator. A synthetic data generator for text recognition
★ 3.7kbehavioral-cloning. Behavioral cloning: end-to-end learning for self-driving cars.
★ 122professional-programming. A collection of learning resources for curious software engineers
★ 51kdistil-whisper. Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
★ 4.1kDeslantImg. The deslanting algorithm sets text upright in images. Python, C++ and OpenCL implementations provided.
★ 155surface-crack-detection. Deep Learning Model for Crack Detection and Segmentation
★ 163text-segmentation. Document Scanner and Word Segmentation
★ 123SFDA-OMR. Source-Free Domain Adaptation for Optical Music Recognition (ICDAR 2024)
★ 1surya. OCR, layout analysis, reading order, table recognition in 90+ languages
★ 21kTransNorm. Code release for "Transferable Normalization: Towards Improving Transferability of Deep Neural Networks" (NeurIPS 2019)
★ 79Awesome-CV. :page_facing_up: Awesome CV is LaTeX template for your outstanding job application
★ 28ka2s-transformer. A Transformer approach for polyphonic Audio-to-Score (A2S) transcription (ICASSP 2024)
★ 15sakura. :cherry_blossom: a minimal css framework/theme.
★ 4.4kinsanely-fast-whisper. Jupyter Notebook
★ 13kawesome-github-profile-readme-templates. This repository contains best profile readme's for your reference.
★ 5.3ktortoise-tts. A multi-voice TTS system trained with an emphasis on quality
★ 15kml-engineering. Machine Learning Engineering Open Book
★ 19kNeural-Net-Zero-to-Hero-with-Andrej. This repository contains the collection of explorative notebooks pure in python and in the language that we, humans can read. Have tried to compile all lectures from the Andrej Karpathy's 💎 playlist on Neural Networks - which we will end up with building GPT.
★ 144awesome-domain-adaptation. A collection of AWESOME things about domain adaptation
★ 5.5kopeninterpreter. A coding agent for open models like Kimi K3
★ 67kML-Papers-Explained. Explanation to key concepts in ML
★ 8.6kVATr. Python
★ 89audiolm-pytorch. Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
★ 2.6kreal-a2s. On the use of synthetic music for transcribing real saxophone audio recordings (INTERSPEECH 2023)
★ 2research-WriterAdaptation-HTR. Source code for WACV20 paper "Unsupervised Adaptation for Synthetic-to-Real Handwritten Word Recognition".
★ 16AudioGPT. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
★ 10kSTR-Fewer-Labels. Scene Text Recognition (STR) methods trained with fewer real labels (CVPR 2021)
★ 185Case-Sensitive-Scene-Text-Recognition-Datasets. This dataset contains re-annotations of 4 popular Latin/English scene text recognition datasets.
★ 53Scene-Text-Recognition-Recommendations. Papers, Datasets, Algorithms, SOTA for STR. Long-time Maintaining
★ 353Scene-Text-Recognition.
★ 619early-vs-late-multimodal-music-transcription. Python
★ 1ssl-symbol-classification. Unlabelled self-supervised learning symbol classification
★ 1continual-learning-binarization. Continual Learning for Document Binarization
★ 2SparK. [ICLR'23 Spotlight🔥] The first successful BERT/MAE-style pretraining on any convolutional network; Pytorch impl. of "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling"
★ 1.4klate-fusion-music-transcription. Late-fusing OMR and A2S predictions using four different algorithms
★ 2ismir-lbd-2021. Late-fusing OMR and A2S predictions using the Smith-Waterman local alignment algorithm (ISMIR 2021)
★ 3smc-2022. Transfer learning between OMR and A2S models
★ 2ScoreGenerator. Music score generator for OMR systems
★ 4audiomentations. A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
★ 2.3kstable-diffusion-tensorflow. Stable Diffusion in TensorFlow / Keras
★ 1.6kfast-stable-diffusion. fast-stable-diffusion + DreamBooth
★ 7.9kpytorch-CycleGAN-and-pix2pix. Image-to-Image Translation in PyTorch
★ 1Person_remover. People removal in images using Pix2Pix and YOLO.
★ 162EffectivePyTorch. PyTorch tutorials and best practices.
★ 1.7kpytorch-softdtw-cuda. Fast CUDA implementation of (differentiable) soft dynamic time warping for PyTorch
★ 735pysdtw. Torch implementation of Soft-DTW, supports CUDA.
★ 49crnn_seq2seq_ocr_pytorch. Extremely simple implement for Chinese OCR by PyTorch.
★ 346ICASSP2021-A2S. accompanying code for my ICASSP2021 paper
★ 19PlotNeuralNet. Latex code for making neural networks diagrams
★ 25kiterative-stratification. scikit-learn cross validators for iterative stratification of multilabel data
★ 896annotated_research_papers. This repo contains annotated research papers that I found really good and useful
★ 2.8kPython. All Algorithms implemented in Python
★ 223keat_tensorflow2_in_30_days. Tensorflow2.0 🍎🍊 is delicious, just eat it! 😋😋
★ 9.9kdeepmind-research. This repository contains implementations and illustrative code to accompany DeepMind publications
★ 15kScoreGenerator. Python Scores Generator
★ 2