This is your work, valued
ocrd_detectron2. OCR-D wrapper for detectron2 based segmentation models
★ 16workflow-configuration. a makefilization for OCR-D workflows, with configuration examples
★ 10ocrd_publaynet. convert PubLayNet data into METS/PAGE-XML
★ 10nmalign. forced alignment of lists of string by fuzzy string matching
★ 9page_dewarp. Text page dewarping using a "cubic sheet" model
★ 8docstruct. Document structure detection from PAGE-XML to METS-XML
★ 6ocrd_wrap. OCR-D wrapper for arbitrary coords-preserving image operations
★ 4ocrd_page2tei. OCR-D wrapper for page2tei
★ 3wrap_opencv-python-headless. rebrand opencv-python-headless
★ 3Mask_RCNN. Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow
★ 3mkn-kurrent-gt. Kurrent GT from the Moravian Knowledge Network handwritten periodicals
★ 2alto-tools. Python tools for performing various operations on ALTO XML files
★ 2hsbcala. train Calamari models for Upper Sorbian (Fraktur and Antiqua) prints on HPC
★ 2dta-lexdb-applications. formatting and integrating the Deutches Textarchiv dictionary into various applications
★ 2ocrd_doxa. OCR-D wrapper for DoxaPy image binarization via locally adaptive thresholding
★ 1Adabelief-Optimizer. Repository for NeurIPS 2020 Spotlight "AdaBelief Optimizer: Adapting stepsizes by the belief in observed gradients"
★ 1ocrd-demo-2021-05-12. Demos for OCR-D presentation at OCR@vDHd
★ 1OCR. OCR Models and Training Data Sorbian languages in Latin and Fraktur
★ 1pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kpage-dewarp. Document image dewarping library using a cubic sheet model
★ 236sbb_ocr_conf_eval. Python
★ 2pagexml-mets-viewer. Web app to upload and display multiple PageXML files
★ 2awesome-ocr.
★ 1kDocLayout-YOLO. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
★ 2.2kDocTr. The official code for “DocTr: Document Image Transformer for Geometric Unwarping and Illumination Correction”, ACM MM, Oral Paper, 2021.
★ 438ocrd_paddleocr. OCR-D integration for PaddleOCR
★ 2nanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.
★ 62kocrd_party. OCR-D processor for the party text recognizer
★ 3ocrd_yolo. OCR-D wrapper for yolo based on the ocrd_detectron2 wrapper
★ 2sbb_binarizer_pytorch_converter. Python
★ 5mlm-scoring. Python library & examples for Masked Language Model Scoring (ACL 2020)
★ 350bytellama. A llama with octet tokenization
★ 3HTRMoPo. Schemas for repositories of HTR/OCR models
★ 12party. Page-wise text recognition with lower-supervision line data models
★ 54DeepSeek-V3. Python
★ 104kDewarpNet. Code for the paper "DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression Networks" (ICCV '19)
★ 624tei-publisher-app. The main TEI Publisher app
★ 81Projects. Projects of Mannheim University Library
★ 3ocr-model-catalogue. This repository contains a collection of layout analysis and text recogntion models.
★ 1ocapy. OCR Confidence Analysis script written in python
★ 6pdftotree. :evergreen_tree: A tool for converting PDF into hOCR with text, tables, and figures being recognized and preserved.
★ 461dach-gt. Ground truth and full text for selected prints of German libraries
★ 2DHd2024-demo. Shell
★ 3Kitodo-OCR-Docker. Dockerfile
★ 2loghi. Loghi is a comprehensive toolkit designed for Handwritten Text Recognition (HTR) and Optical Character Recognition (OCR), offering an accessible approach to transcribing historical documents and training models for specialized needs.
★ 148olahd_backend. Project repository for the backend module of OCR-D Implementation Project OLA-HD
★ 6olahd_user_frontend. Project repository for the frontend module of OCR-D Implementation Project OLA-HD
★ 6gt_structure_all. Python
★ 4awesome-ocr. Links to awesome OCR projects
★ 3.1kdta-tools. Tools used in the project "Deutsches Textarchiv"
★ 4llama.cpp. LLM inference in C/C++
★ 122ktesstrain. Train Tesseract LSTM with make
★ 722github-cheat-sheet. A list of cool features of Git and GitHub.
★ 58kgt-management-dhd-2024. Slides and materials for the GT workshop at DHd 2024
★ 2GT-commentaries-layout. Ground truth data for Page Layout Analysis of Historical Classical Commentaries.
★ 1gt_structure_text. The OCR-D Ground Truth text and structure corpus was created between 2015 -2017. In the years since 2017, this corpus has been further curated and supplemented with metadata where appropriate. The corpus includes page XML files within annotations of the text and structure include.
★ 5frat. Fast Rectangle Annotation Tool
★ 9ulb-groundtruth-eval-odem-other. OCR Groundtruth ULB VD18 - OCR-D Phase III
★ 4ulb-groundtruth-eval-odem-ger. OCR Grountruth ULB VD18 German Fraktur - OCR-D Phase III
★ 5ulb-groundtruth-eval-odem-lat. OCR Groundtruth ULB VD18 Latin - OCR-D Phase III
★ 4gt_structure_dtaText. Batchfile
★ 2ddblabs-ometha. A fast and resilient OAI PMH Harvester with TUI and CLI and support for automation.
★ 6slub_digitalcollections. Templates, Styles and Configuration for Kitodo.Presentation based Digital Collections by SLUB Dresden
★ 9ocrd-demo-mets-server. Makefile
★ 4calamari. Line based ATR Engine based on OCRopy
★ 1.2kocrd-workflows. Shell
★ 2YALTAi. You Actually Look Twice At it
★ 42sampo. A shell script API server for running your shell scripts.
★ 57jdeskew. ICIP 2022: Adaptive Radial Projection on Fourier Magnitude Spectrum for Document Image Skew Estimation
★ 168imgtabdet. Tool that tries to find line grids (i.e. a table) in an image
★ 2buildg. Interactive debugger for Dockerfile, with support for IDEs (VS Code, Emacs, Neovim, etc.)
★ 1.5kgt-guideline-examples.
★ 1gt-repo-template. A template for creating a ground truth repo with the various functions and features: such as metadata creation, data analysis and presentation.
★ 8tei2hocr. XSLT Stylesheet to convert TEI OCR data to HOCR
★ 2ocr-conversion. Conversions between various OCR formats
★ 84ocrd-webapi-implementation. Python
★ 4Docker-layout-OHG. Shell
★ 1Docker-htr-OHG. Docker image that isolates and automates the use of htr tools for the OHG dataset.
★ 1laypa. Layout analysis to find layout elements in documents (similar to P2PaLA)
★ 22textract2page. Convert AWS Textract JSON to PRImA PAGE XML
★ 6hpc-rocket. Python
★ 30quipucamayoc. dev repo for article
★ 33Removing-Shadows-from-Images-of-Documents-by-Steve-Bako. This is the source code from "Removing Shadows from Images of Documents" by Steve Bako , Soheil Darabi, Eli Shechtman, Jue Wang, Kalyan Sunkavalli, and Pradeep Sen
★ 8codon. A high-performance, zero-overhead, extensible Python compiler with built-in NumPy support
★ 17kLayout2Graph. An official implementation of paper "Paragraph2Graph: A Language-independent GNN-based framework for layout analysis"
★ 82deepdoctection. A Repo For Document AI
★ 3.2kdoctr. docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
★ 6.2kDocLayNet. DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis
★ 452PaddleDetection. Object Detection toolkit based on PaddlePaddle. It supports object detection, instance segmentation, multiple object tracking and real-time multi-person keypoint detection.
★ 14kHistorical-document-layout-analysis. Python
★ 6gt_corpus_benchmark. This repo provides a collection of ground truth data. The collection was compiled under different aspects (complexity of the layouts and use of the fonts). The individual data are also characterized by metadata. The metadata is based on the labeling scheme of OCR-D/PrimaLab.
★ 2paperless-ngx. A community-supported supercharged document management system: scan, index and archive all your documents
★ 44kwf_server_nf_script. Nextflow
★ 1ocrd_trocr. OCR-D processor for TrOCR
★ 2ocrd_froc. Python
★ 8CITlabErrorRate. tool to calculate quality of HTR/ATR, KWS, and Text2Image
★ 3ocreval. Update of the ISRI Analytic Tools for OCR Evaluation with UTF-8 support
★ 60METS-schema. METS 1.x and METS 2 schemas
★ 26digital-derivans. Derive new digitals from existing ones
★ 7NodeFlow. An Editor for creating simple or complex OCR workflows
★ 17newspaper-navigator. Jupyter Notebook
★ 272ICA_for_demixing_images. Python
★ 7bombers_baedeker. Weiterführung der ersten, prototypischen Vorgehensweise bei der XMLisierung und Extraktion von Entitäten aus dem zweibändigen Werk „Bomber's Baedeker“
★ 2Awesome-Image-Inpainting. A curated list of image inpainting and video inpainting papers and resources
★ 2.2kCascadeTabNet. This repository contains the code and implementation details of the CascadeTabNet paper "CascadeTabNet: An approach for end to end table detection and structure recognition from image-based documents"
★ 1.5kreichsanzeiger-nlp. Reichsanzeiger-NLP: NER/NEL corpus for the German historical newspaper "Deutscher Reichsanzeiger und Preußischer Staatsanzeiger" (1819–1939)
★ 16scorep_binding_python. Allows tracing of python code using Score-P
★ 37docstruct. Document structure detection from PAGE-XML to METS-XML
★ 5download-gitter.im-chat. tiny tool to download gitter.im chat
★ 4soldatenbriefe. Corpus data for collection “Soldatenbriefe des 18. und 19. Jahrhunderts”
★ 3gt-MufiLevelRules. OCR-D-Level-Rules can be created automatically with gt-MufiLevelRules from the encodings published by MUFI: The Medieval Unicode Font Initiative.
★ 2eynollah. Document Layout Analysis
★ 409OtoN_Converter. Converter from basic OCRD process workflow to Nextflow workflow script
★ 4hyperlearn. 2-2000x faster ML algos, 50% less memory usage, works on all hardware - new and old.
★ 2.5kreichsanzeiger-gt. Ground truth for German newspaper "Deutscher Reichsanzeiger und Preußischer Staatsanzeiger" (1819–1945)
★ 11ocr-util. Evaluate data from mass digitalization workflows
★ 7ddb-metadata-schematron-validation. Schematron-Validierungen der Fachstelle Bibliothek der Deutsche Digitalen Bibliothek
★ 6ddev-dfgviewer. DDEV Development System for the DFG-Viewer
★ 3snakeviz. An in-browser Python profile viewer
★ 2.6ktiler. N-dimensional NumPy array tiling and merging with overlapping, padding and tapering
★ 76quiver-back-end. The back end of the OCR-D quality dashboard webapp.
★ 1pero-enhance. Python
★ 20hsp-erfassung. TypeScript
★ 5hsbcala. train Calamari models for Upper Sorbian (Fraktur and Antiqua) prints on HPC
★ 2pero-ocr. Python
★ 72fix-perspective. C++
★ 4blitzDrt. Tool to correct perspective distortion. Does not correct verticals (column separators, table lines) yet. Uses Blitz++ library.
★ 7docformer. Implementation of DocFormer: End-to-End Transformer for Document Understanding, a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU)
★ 290pytesstrain. Python tools for Tesseract OCR training
★ 26citlab-article-separation-new. Modules used for separating articles in (historical) newspapers and similar documents. This repository is part of the European Union's Horizon 2020 project NewsEye. For more information about the project see https://www.newseye.eu/.
★ 22ddblabs-iiimets. IIIF to METS/MODS conversion script
★ 3edx-lint. Custom tooling for pylint and other repo management tools
★ 52pd3f. 🏭 PDF text extraction pipeline: self-hosted, local-first, Docker-based
★ 335ocrd_manager. frontend for ocrd_controller and adapter towards ocrd_kitodo
★ 10alphashape. Toolbox for constructing alpha shapes.
★ 311OPERANDI_TestRepo. This is the OPERANDI project's test repository.
★ 3yt-dlp. A feature-rich command-line audio/video downloader
★ 181kocrd_kitodo. Docker integration of Kitodo.Production and OCR-D
★ 9pyhunspell. (Official repo for pypi package) Python bindings for the Hunspell spellchecker engine
★ 190ocrd_controller. Path to network implementation of OCR-D
★ 6docker-openssh-server. Dockerfile
★ 610ocrd_page2tei. OCR-D wrapper for page2tei
★ 3page2tei. XSLT
★ 2layout-model-training. The scripts for training Detectron2-based Layout Models on popular layout analysis datasets
★ 220magnitude. A fast, efficient universal vector embedding utility package.
★ 1.7kocrmultieval. Extensible evaluation of (intermediate) results of an OCR workflow
★ 4dfg-viewer. The DFG Viewer is a free web service for browsing digitized books from remote library repositories in a rich and dynamic environment.
★ 34PaddleOCR2Pytorch. PaddleOCR inference in PyTorch. Converted from [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
★ 1.2kocrd_detectron2. OCR-D wrapper for detectron2 based segmentation models
★ 16simple-docstrum. A step-by-step C# implementation of the Docstrum algorithm
★ 24carolineminuscule-groundtruth. OCR ground truth for Caroline Miniscule
★ 11BoundaryNet. BoundaryNet - A Semi-Automatic Layout Annotation Tool
★ 24glosat_table_dataset. GloSAT Historical Measurement Table Dataset
★ 11Order_Relation_Operator. Learning to Sort Handwritten Text Lines in Reading Order through Estimated Binary Order Relations
★ 5ReadingBank. ReadingBank: A Benchmark Dataset for Reading Order Detection
★ 117awesome-document-understanding. A curated list of resources for Document Understanding (DU) topic
★ 1.5kDocBank. DocBank: A Benchmark Dataset for Document Layout Analysis
★ 653docsa. SLUB Document Classification and Similarity Analysis
★ 10LuigiNLP. A workflow system for Natural Language Processing.
★ 21openaudiosearch. Open Audio Search
★ 120vosk-api. Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
★ 15kocrd_vandalize. Demo processor to illustrate OCR-D Python API
★ 5jc. CLI tool and python library that converts the output of popular command-line tools, file-types, and common strings to JSON, YAML, or Dictionaries. This allows piping of output to tools like jq and simplifying automation scripts.
★ 8.7kworkflow-configuration. a makefilization for OCR-D workflows, with configuration examples
★ 10transkribus-to-prima. Convert Transkribus PAGE-XML to standard PAGE-XML
★ 12teihub. Automated listing of repos in GitHub with XML files containing teiHeader. Find a project using TEI today!
★ 17dtabf. DTA Base Format (DTABf)
★ 19AletheiaTools. AletheiaTools is a collection of tools for transforming file formats (PAGE XML) and metadata formats (METS). It is a kind of Ground Truth Swiss Knife ;-)
★ 2gt-fraktur. Shell
★ 7maskrcnn_tf2. Mask R-CNN for object detection and instance segmentation with Keras and TensorFlow V2 and ONNX and TensorRT optimization support.
★ 39geokelone. integrates spatial and textual data processing tools into a modular software package which features preprocessing, geocoding, disambiguation and visualization
★ 5page-to-alto. Convert PAGE (v. 2019) to ALTO (v. 2.0 - 4.2)
★ 17ocrd_fileformat. OCR-D wrapper for ocr-fileformat
★ 6unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kbowkin. A tool for patching binaries to use specific versions of glibc
★ 22tensorflow-wheels. A repo to store custom tensorflow wheels in releases
★ 5mets-mods2tei. Convert bibliographic meta data in MODS format to TEI headers
★ 10Fibeln. Transkriptionen von Fibeln (19. Jahrhundert)
★ 11awesome-pipeline. A curated list of awesome pipeline toolkits inspired by Awesome Sysadmin
★ 6.6kmets2iiif. An implementation of the IIIF Presentation API v2 based on XSLT
★ 7Multi-Type-TD-TSR. Extracting Tables from Document Images using a Multi-stage Pipeline for Table Detection and Table Structure Recognition
★ 289CTCDecoder. Connectionist Temporal Classification (CTC) decoding algorithms: best path, beam search, lexicon search, prefix search, and token passing. Implemented in Python.
★ 837ocr-greek_cursive. Training files for Greek cursive script (in early print)
★ 15German-NLP. Curated list of open-access/open-source/off-the-shelf resources and tools developed with a particular focus on German
★ 527timur. Finite-state morphology for German
★ 10DeslantImg. The deslanting algorithm sets text upright in images. Python, C++ and OpenCL implementations provided.
★ 155ALBEF. Code for ALBEF: a new vision-language pre-training method
★ 1.8kcookiecutter. A cross-platform command-line utility that creates projects from cookiecutters (project templates), e.g. Python package projects, C projects.
★ 25kocrd-gbn. OCR-D compliant toolset for optical layout recognition on historical german-language documents published in Brazil
★ 11table-transformer. Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS evaluation metric.
★ 2.9kPyLaia. A deep learning toolkit specialized for handwritten document analysis
★ 259PrintedBookLayout.
★ 4Doxa. A Local Adaptive Thresholding framework for image binarization written in C++, with JS, Python and MATLAB bindings. Implementing: Otsu, Bernsen, Niblack, Bradley, Sauvola, Wolf, Gatos, NICK, Su, T.R. Singh, WAN, ISauvola, Feng, Phansalkar, AdOtsu, along with DRDM and all Pseudo Merics.
★ 194hist_thresh. Jupyter Notebook
★ 105dask. Parallel computing with task scheduling
★ 14kwarp-ctc. Fast parallel CTC.
★ 4.1knautilusocr. METS/ALTO OCR enhancing tool by the National Library of Luxembourg (BnL)
★ 56TextRecognitionDataGenerator. A synthetic data generator for text recognition
★ 3.7kEasyOCR. Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
★ 30ktess5train-fonts. Files and Scripts to run Tesseract 5 LSTM Training using fonts
★ 78tractjs. Run ONNX and TensorFlow inference in the browser.
★ 76wtpsplit. Toolkit to segment text into sentences or other semantic units in a robust, efficient and adaptable way.
★ 1.3kmodern-unix. A collection of modern/faster/saner alternatives to common unix commands.
★ 33ksolr-ocrhighlighting. Highlighting various OCR formats directly in Solr
★ 88DocCreator. DIAR software for synthetic document image and groundtruth generation, with various degradation models for data augmentation
★ 138Android-OTA-update-by-USB-disk. Roff
★ 4bibliothekartag-2021. Slides presenting the collaborative OCR GT collection at SLUB
★ 1neuspell. NeuSpell: A Neural Spelling Correction Toolkit
★ 713ACL2018_Multi_Input_OCR. Python
★ 13text-matcher. A simple text reuse detection CLI tool.
★ 140ocr4all_models. Pretrained mixed models to be used in OCR4all
★ 8langid.py. Stand-alone language identification system
★ 2.5kcibuildwheel. 🎡 Build Python wheels for all the platforms with minimal configuration.
★ 2.3kpylsd. python bindings for LSD - Line Segment Detector.
★ 169DocumentLayoutAnalysis. Document Layout Analysis resources repos for development with PdfPig.
★ 637ocr-text-extraction. A simple program to extract the text from an image before performing OCR
★ 221htr-gt. Sources for training Handwritten Text Recognition models
★ 6