This is your work, valued
DE-GAN. Document Image Enhancement with GANs - TPAMI journal
★ 224DocEnTR. DocEnTr: An end-to-end document image enhancement transformer - ICPR 2022
★ 190SSL-OCR. Text-DIAE: A Self-Supervised Degradation Invariant Autoencoders for Text Recognition and Document Enhancement - AAAI 2023
★ 30HTRbyMatching. Hadwritten Text Recognition in Few-shot Scenario
★ 22FaceMatching. Face matching using deep learning (CNN embedding + triplet loss)
★ 11OCR-TR. Optocal Character Recognition (OCR / HTR) using Transformers
★ 11DocVXQA. DocVXQA: Context-Aware Visual Explanations for Document Question Answering
★ 10watermarking-documents. Jupyter Notebook
★ 4GroundedExpAcc. Python
★ 2dali92002.github.io. HTML
★ 1opik. Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
★ 21kdocling. Get your documents ready for gen AI
★ 64keasymcp. TypeScript
★ 32SVGCraft. [WACV 2026 Round 1] Beyond Single Object Text-to-SVG Synthesis with Comprehensive Canvas Layout
★ 24next-level-bert. Python
★ 16awesome-comics-understanding. The official repo of the Comics Survey: "A missing piece in Vision and Language: A Survey on Comics Understanding"
★ 139awesome-diffusion-categorized. collection of diffusion model papers categorized by their subareas
★ 2.2keinops. Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)
★ 9.6kVAR. [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8.7kannotation. annotation system for labelling bounding boxes using openCV
★ 1GraphKD. [ICDAR 2024] (Best Student Paper🏆) Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
★ 16Hyper-Modulation. Official Implementation for "Transferring Unconditional to Conditional GANs with Hyper-Modulation" CVPRW 22 https://arxiv.org/abs/2112.02219
★ 13table-transformer. Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS evaluation metric.
★ 2.9kfiftyone. Refine high-quality datasets and visual AI models
★ 11ktransformer-inertial-poser. Python implementation accompanying the Transformer Inertial Poser paper at SIGGRAPH Asia 2022
★ 93EgoLocate. A real-time system that simultaneously captures human pose, reconstructs the scene in sparse 3D points, and localizes the human in the scene with 6 IMUs and a body-worn phone camera
★ 113LabelAdaptiveMixup-SER. Python
★ 6PFL-DocVQA-Competition. Python
★ 21AvatarPoser. Official Code for ECCV 2022 paper "AvatarPoser: Articulated Full-Body Pose Tracking from Sparse Motion Sensing"
★ 330ControlNet. Let us control diffusion models!
★ 34kdocile. DocILE: Document Information Localization and Extraction Benchmark
★ 150denoising-diffusion-pytorch. Implementation of Denoising Diffusion Probabilistic Model in Pytorch
★ 11kstable-diffusion. A latent text-to-image diffusion model
★ 73kStructured-Diffusion-Guidance. Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
★ 321latent-diffusion. High-Resolution Image Synthesis with Latent Diffusion Models
★ 14khandwriting-synthesis. Handwriting Synthesis with RNNs ✏️
★ 4.8kdoc2graph. Doc2Graph transforms documents into graphs and exploit a GNN to solve several tasks.
★ 139idl_data. OCR Annotations from Amazon Textract for Industry Documents Library
★ 103ivy. Convert Machine Learning Code Between Frameworks
★ 14kGACNN. Generative Adverserial Convolutional Neural Network
★ 9jNMF. Discovering De-similarities of Modular Structure Between Tumor Cells and Normal Cells by Integrating Multiple Data Sources Through Joint Non-Negative Matrix Factorization
★ 9dct-dft-fft-craft. DCT-DFT-FFT Based Method for Text Detection in Underwater Images
★ 3HAGNN. Gene Selection of Microarray Data using Heatmap Analysis and Graph Neural Network
★ 2Machine-Learning. In the summer 2020, I have get a chance to learn machine learning from Andrew Ng, coursework organised by Stanford University. Here, I am going to upload all the assignment done by me during the coursework.
★ 5HandsonML. The Machine Learning Tsunami
★ 5DCN-DQN-TDF-IDF. A Comprehensive Scheme for Tattoo Text Detection
★ 4Assignment_AIforManufacturing_Quality_Assurance_OpGAN. This repository contains the assignments related to Pix2Pix image translation with GAN
★ 3Deep_Learning_with_Python. The deep in deep learning isn’t a reference to any kind of deeper understanding achieved by the approach; rather, it stands for this idea of successive layers of representations......
★ 7SegIris. a procedure of iris segmentation is presented which was designed on the basis of the natural properties of the iris.
★ 5DocSegTr. A Bottom-Up Instance Segmentation Strategy for segmenting document instances using Transformers
★ 2Sim-GAN. Missing Value Estimation of Microarray Data using Sim-GAN
★ 4ContrastiveSupervisedDistillation. This repo contains the code of "Contrastive Supervised Distillation for Continual Representation Learning", Tommaso Barletti, Niccolò Biondi, Federico Pernici, Matteo Bruni, and Alberto Del Bimbo, ICIAP2021
★ 20MAE-pytorch. Unofficial PyTorch implementation of Masked Autoencoders Are Scalable Vision Learners
★ 2.7kncs_metric. Is an image worth five sentences? A new look into semantics of image-text matching - WACV 2022
★ 12GoodNews. Good News Everyone! - CVPR 2019
★ 130SelectiveTextStyleTransfer. ICDAR 2019
★ 25DocSegTr. A Bottom-Up Instance Segmentation Strategy for segmenting document instances using Transformers
★ 59TextRecognitionDataGenerator. A synthetic data generator for text recognition
★ 3.7kvit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25ksemantic_adaptive_margin. WACV 2022 Paper - Is An Image Worth Five Sentences? A New Look into Semantics for Image-Text Matching
★ 16unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kobject-bias. Let there be clock in the beach - WACV 2022
★ 15vissl. VISSL is FAIR's library of extensible, modular and scalable components for SOTA Self-Supervised Learning with images.
★ 3.3kSwin-Transformer. This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
★ 16kthe-gan-zoo. A list of all named GANs!
★ 15kFractals. Matlab project
★ 3DIGITS. Deep Learning GPU Training System
★ 2caffe. Caffe: a fast open framework for deep learning.
★ 1tensorflow. Computation using data flow graphs for scalable machine learning
★ 1DeepRosetta. An universal deep learning models conversor
★ 2LBP-for-word-spotting. Local Binary Pattern for historical handwritten documents
★ 4nmp_qc. Neural Message Passing for Computer Vision
★ 1tiny-dnn. header only, dependency-free deep learning framework in C++11
★ 2DeepSketchHashing. Deep Sketch Hashing source code
★ 3Lenet_nicicon. LeNet applied to NicIcon dataset
★ 5Hallucination.
★ 1sem-pcyc. PyTorch implementation of the paper "Semantically Tied Paired Cycle Consistency for Zero-Shot Sketch-based Image Retrieval", CVPR 2019.
★ 3GoodNews. Good News Everyone! - CVPR 2019
★ 2kornia. Open Source Differentiable Computer Vision Library for PyTorch
★ 1doodle2search. Doodle to Search: Practical Zero Shot Sketch Based Image Retrieval
★ 80online-cv. A minimal Jekyll Theme to host your resume (CV)
★ 1GCN_classification. Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and Retrieval
★ 1swav. PyTorch implementation of SwAV https//arxiv.org/abs/2006.09882
★ 1allanlab. Allan Lab website
★ 1doodle2search.github.io. Doodle to Search: Practical Zero-Shot Sketch-based Image Retrieval
★ 5SigNet. SigNet: Convolutional Siamese Network for Writer Independent Offline Signature Verification
★ 80Multilabel-Fashion-MNIST. Multilabel-Fashion-MNIST
★ 4Naive_text_dataset. A naive approach that generates images and labels of text in simple colored background.
★ 2phoc_clf. Python
★ 3Pytorch-yolo-phoc. Implementation on pytorch of the code from the ECCV 2018 paper - Single Shot Scene Text Retrieval
★ 13Fine_Grained_Clf. Based on the WACV 2020 paper - Fine Grained Classification and Retrieval by Combining Visual and Locally Pooled Textual Features
★ 25GCN_classification. Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and Retrieval
★ 65synth_doc_generation. Official PyTorch Implementation of DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis - ICDAR 2021
★ 93snake. Code for "Deep Snake for Real-Time Instance Segmentation" CVPR 2020 oral
★ 1.2kStacMR. Scene Text Aware Cross Modal Retrieval (StacMR)
★ 24pix2pixHD. Synthesizing and manipulating 2048x1024 images with conditional GANs
★ 6.9k