This is your work, valued
Welcome to my world! Let's code and deploy awesomeness!!
synth_doc_generation. Official PyTorch Implementation of DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis - ICDAR 2021
★ 93DocSegTr. A Bottom-Up Instance Segmentation Strategy for segmenting document instances using Transformers
★ 59STN_FGC. Analyzing Visual Attention Mechanisms for Handwritten Digit Classification
★ 4transformer_classification. Python
★ 2exploring_gans. Exploring GANs
★ 2doc2graph_multi. Multilingual Version for Doc2Graph
★ 2deep-learning-wizard. Open source guides/codes for mastering deep learning to deploying deep learning in production in PyTorch, Python, C++ and more.
★ 1the-incredible-pytorch. The Incredible PyTorch: a curated list of tutorials, papers, projects, communities and more relating to PyTorch.
★ 1ILearnDeepLearning.py. This repository contains small projects related to Neural Networks and Deep Learning in general. Subject are closely linekd with articles I publish on Medium. I encourage you both to read as well as to check how the code works in the action.
★ 1ML-From-Scratch. Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep learning.
★ 1Facial-Expression-Recognization-using-JAFFE. Facial Expression Recognization using JAFFE
★ 1star-vector. StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
★ 4.5khinton-problems. 53 implementations of synthetic learning problems from Geoffrey Hinton's experimental papers (1981-2022). Pure numpy, laptop-runnable, paper-comparison metrics per stub.
★ 32pydantic-deepagents. Open-source, self-hosted Claude Code - a terminal AI assistant and the Python framework behind it. Tool-calling, sandboxed execution, multi-agent teams, skills, checkpoints, unlimited context - on Pydantic AI, any model.
★ 1kOpenHarness. "OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
★ 15kspec-kit. 💫 Toolkit to help you get started with Spec-Driven Development
★ 125kGraphRAG-SDK. Build fast and accurate GenAI apps with GraphRAG SDK at scale 🌟
★ 980Scrapling. 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
★ 72kno-cost-ai. 80+ free AI services for chat, image, video, voice & APIs (may sometimes include access to lead gen ai models for free)
★ 2kcrawl4ai. 🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
★ 76kawesome-design-md. A collection of DESIGN.md files analysis by popular brand design systems. Drop one into your project and let coding agents generate a matching UI.
★ 106kclaude-code-best-practice. from vibe coding to agentic engineering - practice makes claude perfect
★ 64kMADQA. Multimodal Agentic Document QA benchmark (MADQA)
★ 39OpenClaw-RL. OpenClaw-RL: Train any agent simply by talking
★ 5.6kagents. Multi-harness agentic plugin marketplace for Claude Code, Codex CLI, Cursor, OpenCode, GitHub Copilot, and Gemini CLI
★ 38kgoogle_workspace_mcp. Control Gmail, Google Calendar, Docs, Sheets, Slides, Chat, Forms, Tasks, Search & Drive with AI - Comprehensive Google Workspace / G Suite MCP Server & CLI Tool
★ 2.9kcontainerd. An open and reliable container runtime
★ 21kOpenSpec. Spec-driven development (SDD) for AI coding assistants.
★ 63kfull-stack-ai-agent-template. Full-stack AI app generator — FastAPI + Next.js with AI Agents, RAG, streaming, auth, and 20+ integrations out of the box.
★ 1.7kPageIndex. 📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
★ 35kcolpali. The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.
★ 2.7kawesome-n8n-templates. 280+ free n8n automation templates — ready-to-use workflows for Gmail, Telegram, Slack, Discord, WhatsApp, Google Drive, Notion, OpenAI, and more. AI agents, RAG chatbots, email automation, social media, DevOps, and document processing. The largest open-source n8n template collection.
★ 24kself-hosted-ai-starter-kit. The Self-hosted AI Starter Kit is an open-source template that quickly sets up a local AI environment. Curated by n8n, it provides essential tools for creating secure, self-hosted AI workflows.
★ 15kevalchemy. Automatic evals for LLMs
★ 602prompts.chat. f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
★ 167kproduction-agentic-rag-course. Python
★ 8.2kawesome-python. An opinionated list of Python frameworks, libraries, tools, and resources
★ 311koriginal_performance_takehome. Anthropic's original performance take-home, now open for you to try!
★ 4.1kharness-sdk. Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
★ 6.7kgenerative-ai-for-beginners. 21 Lessons, Get Started Building with Generative AI
★ 114kAwesome-Context-Engineering. 🔥 Comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems. hundreds of papers, frameworks, and implementation guides for LLMs and AI agents.
★ 3.3kA-Curated-List-of-ML-System-Design-Case-Studies. This repository contains a curated collection of 300+ case studies from over 80 companies, detailing practical applications and insights into machine learning (ML) system design. The contents are organized to help you easily find relevant case studies based on industry or specific ML use cases.
★ 11kllm-engineer-toolkit. A curated list of 120+ LLM libraries category wise.
★ 11kagentic-doc. Legacy Python library for Agentic Document Extraction (ADE). Use the landingai-ade library for all new projects.
★ 2.4kawesome-context-engineering. A curated list of awesome open-source libraries for context engineering (Long-term memory, MCP: Model Context Protocol, Prompt/RAG Compression, Multi-Agent)
★ 109vision-agent. This tool has been deprecated. Use Agentic Document Extraction instead.
★ 5.3kintellagent. A framework for comprehensive diagnosis and optimization of agents using simulated, realistic synthetic interactions
★ 1.3kLinly-Talker. Digital Avatar Conversational System - Linly-Talker. 😄✨ Linly-Talker is an intelligent AI system that combines large language models (LLMs) with visual models to create a novel human-AI interaction method. 🤝🤖 It integrates various technologies like Whisper, Linly, Microsoft Speech Services, and SadTalker talking head generation system. 🌟🔬
★ 3.4kawesome-vlm-architectures. Famous Vision Language Models and Their Architectures
★ 1.3kgraphiti. Build Real-Time Knowledge Graphs for AI Agents
★ 29kGenAI_Agents. 50+ tutorials and implementations for Generative AI Agent techniques, from basic conversational bots to complex multi-agent systems.
★ 24kVisDoM. Python
★ 45comref-converter. The Common Optical Music Recognition Framework conversion and evaluation toolset
★ 7courses. This repository is a curated collection of links to various courses and resources about Artificial Intelligence (AI)
★ 6.5kmaestro. streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
★ 2.7klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kCADTransformer. [CVPR 2022]"CADTransformer: Panoptic Symbol Spotting Transformer for CAD Drawings", Zhiwen Fan, Tianlong Chen, Peihao Wang, Zhangyang Wang
★ 130DiffusionPen. Official PyTorch Implementation of "DiffusionPen: Towards Controlling the Style of Handwritten Text Generation" - ECCV 2024
★ 104namedcurves. [ECCV'24] [TPAMI'26] NamedCurves: Learned Image Enhancement via Color Naming
★ 36ai-deadlines. :alarm_clock: AI conference deadline countdowns
★ 6klearnable-typewriter. The Learnable Typewriter: A Generative Approach to Text Line Analysis
★ 34awesome-comics-understanding. The official repo of the Comics Survey: "A missing piece in Vision and Language: A Survey on Comics Understanding"
★ 139DistilDoc_ICDAR24. TeX
★ 3LayeredDoc. Official python implementation of LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach
★ 11InstructionModelling. [NeurIPS 2024 Main Track] Code for the paper titled "Instruction Tuning With Loss Over Instructions"
★ 38DA-TextSpotter. Python
★ 1llama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
★ 19kLlamaGen. Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
★ 2kDocScanner. The official repo for “DocScanner: Robust Document Image Rectification with Progressive Learning”, IJCV, 2025.
★ 339DocEdit-Dataset. Release of the DocEdit Dataset associated with the AAAI 2023 paper "DocEdit: Language-guided Document Editing"
★ 11awesome-cvpr-2024. 🤩 An AWESOME Curated List of Papers, Workshops, Datasets, and Challenges from CVPR 2024
★ 146data-augmentation-review. List of useful data augmentation resources. You will find here some not common techniques, libraries, links to GitHub repos, papers, and others.
★ 1.6kRecommendations-Document-Image-Processing. This repository contains a paper collection of the methods for document image processing, including appearance enhancement, deshadowing, dewarping, deblurring, binarization and so on.
★ 396DocRes. [CVPR 2024] DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
★ 629Awesome-Optimal-Transport-in-Deep-Learning. a collection of AWESOME things about Optimal Transport in Deep Learning
★ 354ClipPrompt. A PyTorch implementation of ClipPrompt based on CVPR 2023 paper "CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or Not"
★ 19multimodal-prompt-learning. [CVPR 2023] Official repository of paper titled "MaPLe: Multi-modal Prompt Learning".
★ 817ProText. [AAAI'25, CVPRW 2024] Official repository of paper titled "Learning to Prompt with Text Only Supervision for Vision-Language Models".
★ 128lmppl. Calculate perplexity on a text with pre-trained language models. Support MLM (eg. DeBERTa), recurrent LM (eg. GPT3), and encoder-decoder LM (eg. Flan-T5).
★ 168DEXPERT. A Transformer-based object-centric approach for date estimation of historical photographs
★ 5doc2graph. Doc2Graph transforms documents into graphs and exploit a GNN to solve several tasks.
★ 139Applied-Deep-Learning. Applied Deep Learning Course
★ 3.6kcontextual-bias. Code for the paper "Does Data Repair Lead to Fair Models? Curating Contextually Fair Data To Reduce Model Bias" accepted in WACV2022
★ 6AdvancedLiterateMachinery. A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
★ 1.8kawesome-vision-and-language. A curated list of awesome vision and language resources (still under construction... stay tuned!)
★ 563flash-attention. Fast and memory-efficient exact attention
★ 25kScene-Text-Recognition-Recommendations. Papers, Datasets, Algorithms, SOTA for STR. Long-time Maintaining
★ 353LATIN-Prompt. Python
★ 52lgssl. [CVPR 2023] Learning Visual Representations via Language-Guided Sampling
★ 150Efficient-AI-Backbones. Efficient AI Backbones including GhostNet, TNT and MLP, developed by Huawei Noah's Ark Lab.
★ 4.4kfew-shot-meta-baseline. Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning, in ICCV 2021
★ 655private-gpt. Complete API layer for private AI applications on local models: RAG, skills, tools, MCP, text-to-sql, and more. Works with any OpenAI-compatible inference server.
★ 57kImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1kImageNet-100-Pytorch. (Pytorch) Training ResNets on ImageNet-100 data
★ 67URL. Universal Representation Learning from Multiple Domains for Few-shot Classification - ICCV 2021, Cross-domain Few-shot Learning with Task-specific Adapters - CVPR 2022
★ 147dessurt. Official implementation for Dessurt: Document end-to-end self-supervised understanding and recognition transformer
★ 62DocLayNet. DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis
★ 453flex-dm. [CVPR 2023 highlight] Towards Flexible Multi-modal Document Models
★ 59layout-dm. [CVPR 2023] LayoutDM: Discrete Diffusion Model for Controllable Layout Generation
★ 300llama-dl. High-speed download of LLaMA, Facebook's 65B parameter GPT model
★ 4.1kself-label. Self-labelling via simultaneous clustering and representation learning. (ICLR 2020)
★ 546KnowledgeMiningWithSceneText. Python
★ 38docile. DocILE: Document Information Localization and Extraction Benchmark
★ 150cte-dataset. CTE: Contextualized Table Extraction Dataset
★ 17SIDTD_Dataset. Dataset for ID Document Verification systems
★ 28SlideVQA. SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images (AAAI2023)
★ 106ML-Papers-Explained. Explanation to key concepts in ML
★ 8.6kknowledge-distillation-papers. knowledge distillation papers
★ 765neural_sp. End-to-end ASR/LM implementation with PyTorch
★ 595Deep-Learning. Course: Deep Learning
★ 193gpytorch. A highly efficient implementation of Gaussian Processes in PyTorch
★ 3.9kAI4newbies. This book is about AI, Machine Learning and Deep learning, and was generated using [chatGPT](https://chat.openai.com)
★ 8tech-interview-handbook. Curated coding interview preparation materials for busy software engineers
★ 141kaiaiart. Course content and resources for the AIAIART course.
★ 567few-shot. Repository for few-shot learning machine learning projects
★ 1.3kocr-vqgan. OCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers
★ 85transformers-tutorials. Github repo with tutorials to fine tune transformers for diff NLP tasks
★ 863sketch-primitives. ECCV 2022: Abstracting Sketches through Simple Primitives
★ 27latexify_py. A library to generate LaTeX expression from Python code.
★ 7.6kDocSegTr. A Bottom-Up Instance Segmentation Strategy for segmenting document instances using Transformers
★ 59deepdoctection. A Repo For Document AI
★ 3.2kG-SFDA. code for our ICCV 2021 paper 'Generalized Source-free Domain Adaptation'
★ 110NRC_SFDA. Code for our NeurIPS 2021 paper 'Exploiting the Intrinsic Neighborhood Structure for Source-free Domain Adaptation'
★ 81AaD_SFDA. Code for our NeurIPS 2022 (spotlight) paper 'Attracting and Dispersing: A Simple Approach for Source-free Domain Adaptation'
★ 75OneRing_SF-OPDA. Code for 'OneRing: A Simple Method for Source-free Open-partial Domain Adaptation'
★ 34Awesome-Transformer-Attention. An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
★ 5.1kSSL-OCR. Text-DIAE: A Self-Supervised Degradation Invariant Autoencoders for Text Recognition and Document Enhancement - AAAI 2023
★ 30dali92002.github.io. HTML
★ 1HTRbyMatching. Hadwritten Text Recognition in Few-shot Scenario
★ 22OCR-TR. Optocal Character Recognition (OCR / HTR) using Transformers
★ 11pytorch-loss. label-smooth, amsoftmax, partial-fc, focal-loss, triplet-loss, lovasz-softmax. Maybe useful
★ 2.3kML-Course-Notes. 🎓 Sharing machine learning course / lecture notes.
★ 6.6kvdoc. Python
★ 15STN_FGC. Analyzing Visual Attention Mechanisms for Handwritten Digit Classification
★ 4latr. Implementation of LaTr: Layout-aware transformer for scene-text VQA,a novel multimodal architecture for Scene Text Visual Question Answering (STVQA)
★ 56Hyper-Modulation. Official Implementation for "Transferring Unconditional to Conditional GANs with Hyper-Modulation" CVPRW 22 https://arxiv.org/abs/2112.02219
★ 13ai-table-recognition. Python
★ 39taming-transformers. Taming Transformers for High-Resolution Image Synthesis
★ 6.5kmmocr. OpenMMLab Text Detection, Recognition and Understanding Toolbox
★ 4.8kML-YouTube-Courses. 📺 Discover the latest machine learning / AI courses on YouTube.
★ 17kawesome-ocr. Links to awesome OCR projects
★ 3.1kDALLE2-pytorch. Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch
★ 11ktable-transformer. Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS evaluation metric.
★ 2.9kmetaseq. Repo for external large-scale work
★ 6.6kad-deadlines. Countdown for all* relevant conferences in the domain of autonomous driving
★ 8Awesome-Sketch-Based-Applications. :books: A collection of sketch based application papers.
★ 712layout-model-training. The scripts for training Detectron2-based Layout Models on popular layout analysis datasets
★ 220Transformers-Tutorials. This repository contains demos I made with the Transformers library by HuggingFace.
★ 12kawesome-document-understanding. A curated list of resources for Document Understanding (DU) topic
★ 1.5kdotfiles. Lua
★ 7HCV_IIRC. code for our BMVC 2021 paper "HCV: Hierarchy-Consistency Verification for Incremental Implicitly-Refined Classification"
★ 15Transformer-Explainability. [CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization, a novel method to visualize classifications by Transformer based networks.
★ 2kchessformers. This is a PyTorch implementation of a Transformer Decoder based model that plays chess.
★ 17ncs_metric. Is an image worth five sentences? A new look into semantics of image-text matching - WACV 2022
★ 12modulatedautoencoder. Variable rate with MAE
★ 30SlimCAE. Slimmable Compressive Autoencoders for Practical Neural Image Compression
★ 54idl_data. OCR Annotations from Amazon Textract for Industry Documents Library
★ 103axcell. Tools for extracting tables and results from Machine Learning papers
★ 441OTTER. This code provides a PyTorch implementation for OTTER (Optimal Transport distillation for Efficient zero-shot Recognition), as described in the paper.
★ 71synthesis-in-style. Code for our Paper "Synthesis in Style: Semantic Segmentation of Historical Documents using Synthetic Data"
★ 7cnnimageretrieval-pytorch. CNN Image Retrieval in PyTorch: Training and evaluating CNNs for Image Retrieval in PyTorch
★ 1.5kpose-transfer. :fire: [PR 2023] Multi-scale Attention Guided Pose Transfer (official code).
★ 72nested-transformer. Nested Hierarchical Transformer https://arxiv.org/pdf/2105.12723.pdf
★ 204KD_Lib. A Pytorch Knowledge Distillation library for benchmarking and extending works in the domains of Knowledge Distillation, Pruning, and Quantization.
★ 650SOTR. SOTR: Segmenting Objects with Transformers
★ 192ssl_for_fgvc. Self-Supervised Learning for Fine-Grained Image Categorization
★ 26applied-ml. 📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.
★ 30kGLIP. Grounded Language-Image Pre-training
★ 2.6kMETER. METER: A Multimodal End-to-end TransformER Framework
★ 377top-10-cv-papers-2021. A curated list of the top 10 computer vision papers in 2021 with video demos, articles, code and paper reference.
★ 130GTR. [SIGIR 2021] Retrieving Complex Tables with Multi-Granular Graph Representation Learning.
★ 47Awesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6kVIMER. 视觉预训练基础模型仓库
★ 500SegFormer. Official PyTorch implementation of SegFormer
★ 3.6kTransformer-in-Computer-Vision. A paper list of some recent Transformer-based CV works.
★ 1.5kdocformer. Implementation of DocFormer: End-to-End Transformer for Document Understanding, a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU)
★ 290grid-feats-vqa. Grid features pre-training code for visual question answering
★ 269ClipBERT. [CVPR 2021 Best Student Paper Honorable Mention, Oral] Official PyTorch code for ClipBERT, an efficient framework for end-to-end learning on image-text and video-text tasks.
★ 730ReAgent. A platform for Reasoning systems (Reinforcement Learning, Contextual Bandits, etc.)
★ 3.7kavalanche. Avalanche: an End-to-End Library for Continual Learning based on PyTorch.
★ 2.1kawesome-mlss. 🤖 Machine Learning Summer School Guide
★ 3kgraph-based-deep-learning-literature. links to conference publications in graph-based deep learning
★ 5.1kSelectiveTextStyleTransfer. ICDAR 2019
★ 25object-bias. Let there be clock in the beach - WACV 2022
★ 15awesome-vision-language-pretraining-papers. Recent Advances in Vision and Language PreTrained Models (VL-PTMs)
★ 1.2kPyTorch-GAN. PyTorch implementations of Generative Adversarial Networks.
★ 17kuniversal-computation. Official codebase for Pretrained Transformers as Universal Computation Engines.
★ 245Multilingual-CLIP. OpenAI CLIP text encoders for multiple languages!
★ 832detectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kml-study-plan. The Ultimate FREE Machine Learning Study Plan
★ 3.2kpytorch-tutorial. PyTorch Tutorial for Deep Learning Researchers
★ 32kFSOD-code. Python
★ 246FUDGE. Code for the ICDAR2021 paper "Visual FUDGE: Form Understanding via Dynamic Graph Editing"
★ 33OpenAI-CLIP. Simple implementation of OpenAI CLIP model in PyTorch.
★ 725FSL_CLIP. Few-Shot Learning using CLIP as a visual feature extractor
★ 7FewX. FewX is an open-source toolbox on top of Detectron2 for data-limited instance-level recognition tasks.
★ 366t-SNE-tutorial. A tutorial on the t-SNE learning algorithm
★ 720webpage-template. Simple project webpage template. Originally used in Colorful Image Colorization. ECCV, 2016.
★ 494CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kDALLE-pytorch. Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5.6kTransformer-in-Vision. Recent Transformer-based CV and related works.
★ 1.3knlp-phd-global-equality. A repo for open resources & information for people to succeed in PhD in CS & career in AI / NLP
★ 1.1ktutorial_notebook. Collection of tutorial notebooks
★ 2NLP-progress. Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
★ 23kluckmatters. Understanding Training Dynamics of Deep ReLU Networks
★ 306DeepSDF. Learning Continuous Signed Distance Functions for Shape Representation
★ 1.6kgroup-wise-iccv19. Implementation of "Group-Wise Deep Object Co-Segmentation With Co-Attention Recurrent Neural Network" ICCV 2019
★ 14HolySheet. Word crawler for ancient documents
★ 2aproximated_ged. Bunch of aproximated graph edit distance algorithms.
★ 36graph_db. Synthetic graph database generation.
★ 4graph_metric.pytorch. Graph Metric Learning in PyTorch
★ 10