This is your work, valued
Ⓐ ಥ_ಥ (╯°□°)╯︵ ┻━┻ ★。・:*¯\_(ツ)_/¯*:・゚★
awesome-ocr. Links to awesome OCR projects
★ 3.1khocrjs. Working with hOCR in Javascript
★ 134hocr-spec. The hOCR Embedded OCR Workflow and Output Format
★ 74jcurses. Java Curses implementation
★ 42canvas-editor. Vue component for editing shapes in a canvas
★ 24makefile-parser. Parser and documentation generator for Makefiles
★ 23anno-common. Node.JS/Browser Web Annotation Framework
★ 13anno-frontend. Vue application for displaying/editing annotations
★ 12transkribus-to-prima. Convert Transkribus PAGE-XML to standard PAGE-XML
★ 12hocr-spec-python. Validation of hOCR close to the specs
★ 11vfs. Virtual File Systems with a node fs-like API
★ 10docker-ocropy. OCR recognition with Ocropus in a Docker container
★ 9hocr-dom. Extend DOM to handle hOCR
★ 8ocror-detector. Detect errors in OCR
★ 7tsht. A tiny shell-script based testing framework
★ 7winston-timer. Extend winston to measure time intervals
★ 6mpv-config. mpv config
★ 6ocr-fileformat-samples. Samples for common OCR file formats (hOCR, ABBYY, ALTO)
★ 6kraken-docker. Docker container for the kraken OCR engine
★ 6zts-in-a-box. Zotero Translation Server + Simple Query API + Swagger in Docker
★ 6turtleson. Concise, permissive, TURTLE-like dialect of JSON
★ 6grip-docker. Run grip markdown renderer in a docker container
★ 5jsonld-rapper. Create RDF from JSON-LD with rapper
★ 5ocr-xsl. XSLT 2.0 functions for transforming between hOCR, ALTO and ABBYY
★ 5gdxai-btree.vim. Vim Syntax highlighting for gdx-ai behavior tree files
★ 4shinclude. Include directives for code/markup comments
★ 4object-prune. JavaScript
★ 2tesseract-3.03-models. Tesseract 3.03 / 3.04 models
★ 2ocr-schemas. Convert and transform various OCR formats (hOCR, ALTO, PAGE, FineReader)
★ 2models.
★ 1g-ocr. GOCR - State-of-the-Art OCR foundation for German documents. Fast, CPU, no GPU.
★ 9Mongoku. 🔥The Web-scale GUI for MongoDB
★ 1.4kchipotlai-max. The AI coding agent that runs on stolen Chipotle compute 🌯 Fork of OpenCode with Pepper AI as default model. Community project to add providers from Home Depot, Lowes, Target, Starbucks & more.
★ 1.4ktextbite-dataset. TextBite: A Historical Czech Document Dataset for Logical Page Segmentation
★ 6dfine_kraken. D-FINE for document region segmentation
★ 10sbb_images. Image Annotation Tool and Image Search
★ 17ubma-segmentation-ocr-model. This repository contains a segmentation model for historical and modern prints.
★ 2chr2025-works-on-my-machine. Research artefact for the paper ‘“Works on My Machine”: A Case Study of Replicability Challenges in Computational Humanities Research’ at CHR 2025
★ 4UVDoc. Code for the paper "UVDoc: Neural Grid-based Document Unwarping"
★ 223datasette. Tutorial für die Nutzung der von IDM 4 bereitgestellten Datasette-Instanz
★ 4pagexml-mets-viewer. Web app to upload and display multiple PageXML files
★ 2awesome-ocr.
★ 1kRustPython. A Python Interpreter written in Rust
★ 22kocrd_paddleocr. OCR-D integration for PaddleOCR
★ 2tch-rs. Rust bindings for the C++ api of PyTorch.
★ 5.5kDeepSeek-OCR. Contexts Optical Compression
★ 24kocrd_party. OCR-D processor for the party text recognizer
★ 3ocrd_yolo. OCR-D wrapper for yolo based on the ocrd_detectron2 wrapper
★ 2sbb_binarizer_pytorch_converter. Python
★ 5twisted. Event-driven networking engine written in Python.
★ 6kGround-Truth-CHR2024. Coordinates of manually annotated job ads with a link to ANNO Corpus.
★ 3nicegui. Create web-based user interfaces with Python. The nice way.
★ 16kinventory-card-reader. This repository contains code to read, process, and integrate data from inventory cards.
★ 1Ge-ez-HWR. Python
★ 7scribeocr. Web interface for recognizing text, proofreading OCR, and creating fully-digitized documents.
★ 803HTRMoPo. Schemas for repositories of HTR/OCR models
★ 12party. Page-wise text recognition with lower-supervision line data models
★ 54ocapy. OCR Confidence Analysis script written in python
★ 6tooltuesday. Jupyter Notebook
★ 3elementpath. XPath 1.0/2.0/3.0/3.1 parsers and selectors for ElementTree and lxml
★ 91templating-kubernetes. Templating Kubernetes resources with *real* code
★ 112flynt. A tool to automatically convert old string literal formatting to f-strings
★ 732sbb_page_extractor. Python
★ 1HookFormer. Contextual HookFormer for Glacier Calving Front Segmentation (DOI: 10.1109/TGRS.2024.3368215)
★ 3loghi. Shell
★ 17EasyOCR. Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
★ 30kPagePlus. This script processes PAGE XML files, a format widely used in document layout analysis, to perform various operations like validating, repairing, extending, and modifying text regions and lines.
★ 10loghi. Loghi is a comprehensive toolkit designed for Handwritten Text Recognition (HTR) and Optical Character Recognition (OCR), offering an accessible approach to transcribing historical documents and training models for specialized needs.
★ 148laypa. Layout analysis to find layout elements in documents (similar to P2PaLA)
★ 22alto2txt. Convert ALTO XML to plain text + minimal metadata
★ 17sbb_pixelwise_segmentation. Obsolete repo, merged into eynollah
★ 12types-lxml. Complete lxml external type annotation
★ 84dta-lexdb-applications. formatting and integrating the Deutches Textarchiv dictionary into various applications
★ 2ulb-groundtruth-eval-odem-other. OCR Groundtruth ULB VD18 - OCR-D Phase III
★ 4ulb-groundtruth-eval-odem-lat. OCR Groundtruth ULB VD18 Latin - OCR-D Phase III
★ 4ulb-groundtruth-eval-odem-ger. OCR Grountruth ULB VD18 German Fraktur - OCR-D Phase III
★ 5asciimatics. A cross platform package to do curses-like operations, plus higher level APIs and widgets to create text UIs and ASCII art animations
★ 4.3ktextual-web. Run TUIs and terminals in your browser
★ 1.4knotcurses. blingful character graphics/TUI library. definitely not curses.
★ 4.6kflet. Build realtime web, mobile and desktop apps in Python only. No frontend experience required.
★ 16kpytest-docker. Docker-based integration tests
★ 488madoc-platform. A platform for the display, enrichment, and curation of IIIF-based digital objects
★ 57metha. Command line OAI-PMH harvester and client with built-in cache.
★ 132ocrd_network_tests. Shell
★ 1terminal7. A Multi Platform Terminal Emulator & Multiplexer Running over WebRTC with Touch First UI
★ 179operandi_benchmarking. Benchmarking data for OCR-D - Operandi
★ 1gt-fraktur. Shell
★ 7tessdata_contrib. User contributed (non Google) OCR models for Tesseract
★ 33infocom-zcode-terps. Historical source code for Infocom's Z-machine interpreters
★ 366website_next_generation. The new website
★ 143ocropus4extract. Jupyter Notebook
★ 1ocropus4inf. Jupyter Notebook
★ 4ocropus4train. Jupyter Notebook
★ 2page2tei. A repository for illustrating the transformation of a PAGE XML file into XML-TEI format, resulting from experimentations made for the LECTAUREP project.
★ 17audiveris. Latest generation of Audiveris OMR engine
★ 2.7kodem-ocr. Digitalization Workflows with OCR Backend
★ 4pylib. Personal python library
★ 5HTRVX. HTRVX : HTR Validation with XSD
★ 2citlab-article-separation-new. Modules used for separating articles in (historical) newspapers and similar documents. This repository is part of the European Union's Horizon 2020 project NewsEye. For more information about the project see https://www.newseye.eu/.
★ 22NodeFlow. An Editor for creating simple or complex OCR workflows
★ 17delft. a Deep Learning Framework for Text https://delft.readthedocs.io/
★ 416ocrd-network-setup. Deploy ocrd-processing-server in a VM
★ 4GT-commentaries-OCR. Ground truth data for the Optical Character Recognition of Historical Classical Commentaries.
★ 5gt_corpus_benchmark. This repo provides a collection of ground truth data. The collection was compiled under different aspects (complexity of the layouts and use of the fonts). The individual data are also characterized by metadata. The metadata is based on the labeling scheme of OCR-D/PrimaLab.
★ 2wikimedia-ocr. This repository is now at https://gitlab.wikimedia.org/toolforge-repos/ocr
★ 17gesetze. Bundesgesetze und -verordnungen
★ 1.9kocrd_monitor. Web frontend for ocrd_manager
★ 3quiver-benchmarks. Benchmarking OCR-D workflows in Docker
★ 2gt_structure_text. The OCR-D Ground Truth text and structure corpus was created between 2015 -2017. In the years since 2017, this corpus has been further curated and supplemented with metadata where appropriate. The corpus includes page XML files within annotations of the text and structure include.
★ 5ocrd_froc. Python
★ 8stabi-berlin-gt. Ground truth for digitized publications of Staatsbibliothek zu Berlin
★ 5image.textlinedetector. Segment Images in Text Lines and Words
★ 8requests-unixsocket. Use requests to talk HTTP via a UNIX domain socket
★ 221google-ocr-testbed. A repository containing scripts and output for Google OCR of sample files.
★ 1mean_average_precision. Mean Average Precision for Object Detection
★ 203policy. OCR-D Empfehlungen Volltextdigitalisierung
★ 1quiver-data. Providing the Ground Truth workspaces for the OCR-D QUIVER application
★ 2choco-mufin. Tools for normalizing the use of some characters and checking file consistencies
★ 12choco-mufin. Tools for normalizing the use of some characters and checking file consistencies
★ 1pageattrlib. XSLT Library to work with PAGE XML custom attributes
★ 1docstruct. Document structure detection from PAGE-XML to METS-XML
★ 5awk. One true awk
★ 2.2kMienai.ttf. Mienai (見えない) is a "last-resort" font that does the same job as Adobe Blank in a smaller package, at the cost of some compatibility.
★ 1operandi. Project repository for the OCR-D Implementation Project OPERANDI.
★ 6digi-gt. Ground truth for the digitized historic collections of UB Mannheim
★ 2fix-perspective. C++
★ 4OtoN_Converter. Converter from basic OCRD process workflow to Nextflow workflow script
★ 4ocr-util. Evaluate data from mass digitalization workflows
★ 7digi-gt. Ground truth for the digitized historic collections of UB Mannheim
★ 9DIVA-DAF. Repository for the deep-learning framework DIVA-DAF which is build with historical document image analysis in mind.
★ 19curt. Python
★ 15nmalign. forced alignment of lists of string by fuzzy string matching
★ 9ocrd_kitodo. Docker integration of Kitodo.Production and OCR-D
★ 9frat. Fast Rectangle Annotation Tool
★ 9sbb_binarize_flow_from_directory. Python
★ 2ocrd-webapi-implementation. Python
★ 4ddblabs-iiimets. IIIF to METS/MODS conversion script
★ 3ocr-pipeline. OCR Pipeline Module for Digitalization Workflows
★ 5ocrd_manager. frontend for ocrd_controller and adapter towards ocrd_kitodo
★ 10vim-outlaw. The wanted outliner!
★ 46diffgram. The AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human Supervision, Data Workflow, and UI Catalog to get the most value out of your AI Data.
★ 1.9kocrd_controller. Path to network implementation of OCR-D
★ 6ocrd_page2tei. OCR-D wrapper for page2tei
★ 3kitodo-dev-docker. Shell
★ 4DeslantImg. The deslanting algorithm sets text upright in images. Python, C++ and OpenCL implementations provided.
★ 155OPERANDI_TestRepo. This is the OPERANDI project's test repository.
★ 3sbb_column_classifier. Get the number of columns for a document image
★ 3ocrd_detectron2. OCR-D wrapper for detectron2 based segmentation models
★ 16OCR-Metrics-CER-WER. Sample implementation of OCR metrics (CER, WER) calculation with TesseractOCR and fastwer
★ 30OCR17plus. Data for layout analysis and HTR.
★ 4foliautils. Command-line utilities for working with the Format for Linguistic Annotation (FoLiA), powered by libfolia (C++), written by Ko van der Sloot (CLST, Radboud University)
★ 4ticcltools. Tools for TICCL
★ 14PICCL. A set of workflows for corpus building through OCR, post-correction and normalisation
★ 50qt-agi-studio. AGI (Adventure Game Interpreter) is the adventure game engine used by Sierra On-Line(tm) to create some of their early games. QT AGI Studio (formerly known as Linux AGI Studio) is a program which allows you to view, create and edit AGI games. Basically, it is an enhanced port of the Windows AGI Studio developed by Peter Kelly.
★ 6nagi. New Adventure Game Interpreter
★ 42LibreLingo. 🐢 🌎 📚 a community-owned language-learning platform
★ 2.6kqutebrowser. A keyboard-driven, vim-like browser based on Python and Qt.
★ 12kmets-model. Simple METS Api
★ 2click-repl. Subcommand REPL for click apps
★ 240pymux. A terminal multiplexer (like tmux) in Python
★ 1.5kptpython. A better Python REPL
★ 5.4klibtmux. ⚙️ Python API / wrapper for tmux
★ 1.2ktmuxp-config. Configs for tmuxp (https://github.com/tony/tmuxp)
★ 35tmuxp. 🖥️ Session manager for tmux, built on libtmux.
★ 4.5kflusswerk. Easily create AMQP/RabbitMQ based workflows.
★ 8parcel. 📦🚀 Blazing fast, zero configuration web application bundler
★ 1mocri. Small gRPC microservice for OCR based on kraken
★ 5teihub. Automated listing of repos in GitHub with XML files containing teiHeader. Find a project using TEI today!
★ 17unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kmets2iiif. An implementation of the IIIF Presentation API v2 based on XSLT
★ 7table-transformer. Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS evaluation metric.
★ 2.9kulb-zeitungsprojekt-hp1. Training data from "Hauptphase I" of project "Digitalisierung historischer deutscher Zeitungen"
★ 12metaindex. Official Site: www.metaindex.fr // Cataloger application: HTML5 client // J2EE server // integrated interface with Kibana for statistics.
★ 4im2alto. ImageWare MyBib created OCR Output to ALTO XML XSL Transformation
★ 1archive-hocr-tools. Efficient hOCR tooling
★ 57HandwritingRecognitionSystem. Handwriting Recognition System based on a deep Convolutional Recurrent Neural Network architecture
★ 461tesseract-ocr-for-php. A wrapper to work with Tesseract OCR inside PHP.
★ 3kmodern-unix. A collection of modern/faster/saner alternatives to common unix commands.
★ 33kdhh-text-2021. Slides for the workshop Digital Herrnhut in textwissenschaftlichen Kontexten
★ 1ocrd_typegroups_classifier. Font family detection in historical documents.
★ 7fargv. fast_argv
★ 16browse-ocrd-physical-import. A plugin for browse-ocrd to scan book pages with an android phone camera
★ 2hip21_ocrevaluation. A Survey of OCR Evaluation Tools and Metrics (HIP'21)
★ 7nautilusocr. METS/ALTO OCR enhancing tool by the National Library of Luxembourg (BnL)
★ 56DOOM-FX. Doom/FX for Super Nintendo with SuperFX GSU2A
★ 1.2kJunicode-font. A new version of Junicode font
★ 591WA-map. HTML
★ 1htr-united. Ground Truth Resources for the HTR of patrimonial documents
★ 49dahncorpus. Ground Truth dataset for French 20th typewritten OCR produced by the DAHN project
★ 3Issues. For interactive problem-solving, discussion, help, etc
★ 1fantasy. A curated list of available fantasy consoles/computers.
★ 1.6kPAGETools. Small collection of PAGE XML related scripts used at the ZPD Würzburg
★ 12ExpReal. An expressive realizer for interactive narratives
★ 3ocrd_contrib_ubma. Helper scripts for OCR-D
★ 3List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words. List of Dirty, Naughty, Obscene, and Otherwise Bad Words
★ 3.4kocr-eng-bio-testfiles. OCR English (Bio, Natur) ground truth and testfiles
★ 1Mask_RCNN. Mask R-CNN for object detection and instance segmentation on Keras and TensorFlow
★ 3sfst. Stuttgart Finite State Transducer system
★ 25doctr. docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
★ 6.2kacid-banger. The Endless Acid Banger
★ 1.2kcibuildwheel. 🎡 Build Python wheels for all the platforms with minimal configuration.
★ 2.3ksplitflap. DIY split-flap display
★ 4krhasspy. Offline private voice assistant for many human languages
★ 2.8kleptseg. Python
★ 6tesseract_jpn-vert. Python
★ 2page2tsv. PAGE-XML to TSV
★ 4Assemblies-of-putative-SARS-CoV2-spike-encoding-mRNA-sequences-for-vaccines-BNT-162b2-and-mRNA-1273. RNA vaccines have become a key tool in moving forward through the challenges raised both in the current pandemic and in numerous other public health and medical challenges. With the rollout of vaccines for COVID-19, these synthetic mRNAs have become broadly distributed RNA species in numerous human populations. Despite their ubiquity, sequences are not always available for such RNAs. Standard methods facilitate such sequencing. In this note, we provide experimental sequence information for the RNA components of the initial Moderna (https://pubmed.ncbi.nlm.nih.gov/32756549/) and Pfizer/BioNTech (https://pubmed.ncbi.nlm.nih.gov/33301246/) COVID-19 vaccines, allowing a working assembly of the former and a confirmation of previously reported sequence information for the latter RNA. Sharing of sequence information for broadly used therapeutics has the benefit of allowing any researchers or clinicians using sequencing approaches to rapidly identify such sequences as therapeutic-derived rather than host or infectious in origin. For this work, RNAs were obtained as discards from the small portions of vaccine doses that remained in vials after immunization; such portions would have been required to be otherwise discarded and were analyzed under FDA authorization for research use. To obtain the small amounts of RNA needed for characterization, vaccine remnants were phenol-chloroform extracted using TRIzol Reagent (Invitrogen), with intactness assessed by Agilent 2100 Bioanalyzer before and after extraction. Although our analysis mainly focused on RNAs obtained as soon as possible following discard, we also analyzed samples which had been refrigerated (~4 ℃) for up to 42 days with and without the addition of EDTA. Interestingly a substantial fraction of the RNA remained intact in these preparations. We note that the formulation of the vaccines includes numerous key chemical components which are quite possibly unstable under these conditions-- so these data certainly do not suggest that the vaccine as a biological agent is stable. But it is of interest that chemical stability of RNA itself is not sufficient to preclude eventual development of vaccines with a much less involved cold-chain storage and transportation. For further analysis, the initial RNAs were fragmented by heating to 94℃, primed with a random hexamer-tailed adaptor, amplified through a template-switch protocol (Takara SMARTerer Stranded RNA-seq kit), and sequenced using a MiSeq instrument (Illumina) with paired end 78-per end sequencing. As a reference material in specific assays, we included RNA of known concentration and sequence (from bacteriophage MS2). From these data, we obtained partial information on strandedness and a set of segments that could be used for assembly. This was particularly useful for the Moderna vaccine, for which the original vaccine RNA sequence was not available at the time our study was carried out. Contigs encoding full-length spikes were assembled from the Moderna and Pfizer datasets. The Pfizer/BioNTech data [Figure 1] verified the reported sequence for that vaccine (https://berthub.eu/articles/posts/reverse-engineering-source-code-of-the-biontech-pfizer-vaccine/), while the Moderna sequence [Figure 2] could not be checked against a published reference. RNA preparations lacking dsRNA are desirable in generating vaccine formulations as these will minimize an otherwise dramatic biological (and nonspecific) response that vertebrates have to double stranded character in RNA (https://www.nature.com/articles/nrd.2017.243). In the sequence data that we analyzed, we found that the vast majority of reads were from the expected sense strand. In addition, the minority of antisense reads appeared different from sense reads in lacking the characteristic extensions expected from the template switching protocol. Examining only the reads with an evident template switch (as an indicator for strand-of-origin), we observed that both vaccines overwhelmingly yielded sense reads (>99.99%). Independent sequencing assays and other experimental measurements are ongoing and will be needed to determine whether this template-switched sense read fraction in the SmarterSeq protocol indeed represents the actual dsRNA content in the original material. This work provides an initial assessment of two RNAs that are now a part of the human ecosystem and that are likely to appear in numerous other high throughput RNA-seq studies in which a fraction of the individuals may have previously been vaccinated. ProtoAcknowledgements: Thanks to our colleagues for help and suggestions (Nimit Jain, Emily Greenwald, Lamia Wahba, William Wang, Amisha Kumar, Sameer Sundrani, David Lipman, Bijoyita Roy). Figure 1: Spike-encoding contig assembled from BioNTech/Pfizer BNT-162b2 vaccine. Although the full coding region is included, the nature of the methodology used for sequencing and assembly is such that the assembled contig could lack some sequence from the ends of the RNA. Within the assembled sequence, this hypothetical sequence shows a perfect match to the corresponding sequence from documents available online derived from manufacturer communications with the World Health Organization [as reported by https://berthub.eu/articles/posts/reverse-engineering-source-code-of-the-biontech-pfizer-vaccine/]. The 5’ end for the assembly matches the start site noted in these documents, while the read-based assembly lacks an interrupted polyA tail (A30(GCATATGACT)A70) that is expected to be present in the mRNA.
★ 3.4ksbb_page_extractor. Python
★ 4police-settlements. A FiveThirtyEight/The Marshall Project effort to collect comprehensive data on police misconduct settlements from 2010-19.
★ 151memory_profiler. Monitor Memory usage of Python code
★ 4.6kblitz. Blitz++ Multi-Dimensional Array Library for C++
★ 417GTCheck. Check your modified Ground Truth files with visual support!
★ 10kill-the-newsletter. Convert email newsletters into Atom feeds
★ 3.1k8VIM. A Text Editor inside a keyboard, drawing it's inspiration from 8pen and Vim.
★ 581glyphcollector. Elm
★ 64BitmapFonts. My collection of bitmap fonts pulled from various demoscene archives over the years
★ 1.9kOcrd_anybaseocr_block_segmentation. Jupyter Notebook
★ 3page-xml-to-image. Python
★ 4blitzDrt. Tool to correct perspective distortion. Does not correct verticals (column separators, table lines) yet. Uses Blitz++ library.
★ 7pawls. Software that makes labeling PDFs easy.
★ 434Pipeline. pipeline for xslt and python plugins
★ 2docExtractor. (ICFHR 2020 oral) Code for "docExtractor: An off-the-shelf historical document element extraction" paper
★ 89ocr-data. Jupyter Notebook
★ 6pytesstrain. Python tools for Tesseract OCR training
★ 26kraken_generated-data. Python
★ 3OCR_GS_Data. Double-checked Gold Standard Data for Training and Testing OCR Engines
★ 21tesstrainsh-win. Train Tesseract LSTM with tesstrain.sh on Windows
★ 26vim-visual-multi. Multiple cursors plugin for vim/neovim
★ 4.9ksbb_ocr_postcorrection. Two-Step Approach to OCR Post-Correction
★ 14page2tei. Python snippets that might be useful for exporting transcribed pages from PAGE XML to TEI XML
★ 1