This is your work, valued
Deep Leaning & Computer Vision Lab, South China University of Technology
DocRes. [CVPR 2024] DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
★ 629Recommendations-Document-Image-Processing. This repository contains a paper collection of the methods for document image processing, including appearance enhancement, deshadowing, dewarping, deblurring, binarization and so on.
★ 396DocAligner. [PR 2025] DocAligner: Automating the Annotation of Photographed Documents Through Real-virtual Alignment
★ 110GCDRNet. [TAI 2023] Appearance Enhancement for Camera-captured Document Images in the Wild
★ 58DocKylin. [AAAI 2025] DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
★ 36Awesome-Image-based-Meter-Recognition-Reading.
★ 32Marior. [ACM MM 2022] Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild
★ 26WMeter-Reader. [TIM 2025] Towards Accurate Readings of Water Meters by Eliminating Transition Error: New Dataset and Effective Solution
★ 19DocAligner-Distortion. [PRL 2025] Enhancing Document Dewarping Evaluation: A New Metric with Improved Accuracy and Efficiency
★ 7DocHighlight. [PRCV 25] Towards Real-World Document Specular Highlight Removal: The DocHighlight Dataset and DocSHRNet Method
★ 6sub2api. Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
★ 35kclaude-relay-service. CRS-自建Claude Code镜像,一站式开源中转服务,让 Claude、OpenAI、Gemini、Droid 订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
★ 12kUniHIR. [ACL 2026 main] The official GitHub page of "Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLM"
★ 18Awesome-Generate-AI-for-Photography. Awesome List for Generative Models and Photography
★ 6ruankao. 软考达人 - 最新最全免费的软考题库。 高级:系统架构设计师、系统分析师、信息系统项目管理师、系统规划与管理师、网络规划设计师。 中级:软件设计师、网络工程师、系统集成项目管理工程师、数据库系统工程师、信息安全工程师、信息系统管理工程师、信息系统监理师、软件评测师、嵌入式系统设计师、电子商务设计师、多媒体应用设计师。 初级:信息系统运行管理员、信息处理技术员、网络管理员、程序员。
★ 1.8kVisuRiddles. VisuRiddles: Fine-grained Perception is a important thing for Multimodal Large Models in Riddles Solving
★ 20minimalRL. Implementations of basic RL algorithms with minimal lines of codes! (pytorch based)
★ 3.2ktop-cvpr-2026-papers. About This repository is a curated collection of the most exciting and influential CVPR 2026 papers. 🔥 [Paper + Code + Demo]
★ 562VoxCPM. VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
★ 35kopendataloader-pdf. PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
★ 28kawesome-3d-datasets. [CVPRW'26] A collection and survey of 3d dataset
★ 34daily_stock_analysis. LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
★ 60ksupertonic. Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
★ 14kAwesome-Streaming-Video-Understanding. 🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖
★ 420hello-agents. 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
★ 70kSpider_XHS. 小红书爬虫数据采集,小红书全域运营解决方案
★ 7.1kxhs. 基于小红书 Web 端进行的请求封装。https://reajason.github.io/xhs/
★ 2.2knanobot. Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
★ 46kMiroFish. A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物
★ 70kedict. 🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
★ 16kMobileAgent. Mobile-Agent: The Powerful GUI Agent Family
★ 9kAgentCPM-GUI. AgentCPM-GUI: An on-device GUI agent for operating Android apps, enhancing reasoning ability with reinforcement fine-tuning for efficient task execution.
★ 1.4klandlord. 斗地主
★ 401ddz_game. 斗地主游戏
★ 630DouyinLiveRecorder. 可循环值守和多人录制的直播录制软件,支持抖音、TikTok、Youtube、快手、虎牙、斗鱼、B站、小红书、pandatv、sooplive、flextv、popkontv、twitcasting、winktv、百度、微博、酷狗、17Live、Twitch、Acfun、CHZZK、shopee等40+平台直播录制
★ 11krlcard. Reinforcement Learning / AI Bots in Card (Poker) Games - Blackjack, Leduc, Texas, DouDizhu, Mahjong, UNO.
★ 3.5kTrendRadar. ⭐AI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.🎯 告别信息过载,你的 AI 舆情监控助手与热点筛选工具!聚合多平台热点 + RSS 订阅,支持关键词精准筛选。AI 智能筛选新闻 + AI 翻译 + AI 分析简报直推手机,也支持接入 MCP 架构,赋能 AI 自然语言对话分析、情感洞察与趋势预测等。支持 Docker ,数据本地/云端自持。集成微信/飞书/钉钉/Telegram/邮件/ntfy/bark/slack 等渠道智能推送。
★ 61kRoMaV2. Python
★ 653DeepSeek-OCR. Contexts Optical Compression
★ 24kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kAwesome-GUI-Agents. A curated collection of resources, tools, and frameworks for developing GUI Agents.
★ 445gpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20kForCenNet.
★ 82QT-TextSR. This repository is the implementation of "QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a Query-aware Transformer", Neurocomputing 2024..
★ 20Dolphin. The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
★ 9kWildDoc. The official repo for “WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?“
★ 75DocAligner-Distortion. [PRL 2025] Enhancing Document Dewarping Evaluation: A New Metric with Improved Accuracy and Efficiency
★ 7OCR-Reasoning. [ICLR 2026] OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
★ 76AutoScaler. [PR 2026] The official GitHub page of "AutoScaler: Self Scale Alignment for Handwritten Mathematical Expression Recognition"
★ 9TokenFD. [ICCV2025] A Token-level Text Image Foundation Model for Document Understanding
★ 135DocSAM. Python
★ 33Awesome-Generative-Models-for-OCR. [arXiv 25] OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
★ 273M6Doc.
★ 166PromptDA. [CVPR 2025] Prompt Depth Anything
★ 1.1kDocLayLLM. [CVPR 2025] DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
★ 30LGGPT. [IJCV 2025] Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
★ 158worth-calculator. Calculating the actual value of your job beyond just salary
★ 3.3kcherry-studio. AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs
★ 49kRedundancyLens. Python
★ 8DocDewarpHV. Python
★ 13RWMD_dataset. The Real-world Mobile Document (RWMD) Database
★ 8FastV. [ECCV 2024 Oral] Code for paper: An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
★ 593AnimatedDrawings. Code to accompany "A Method for Animating Children's Drawings of the Human Figure"
★ 13ksurya. OCR, layout analysis, reading order, table recognition in 90+ languages
★ 21kHDR. [AAAI2025 Oral] Predicting the Original Appearance of Damaged Historical Documents
★ 111PDFMathTranslate. [EMNLP 2025 Demo] PDF scientific paper translation with preserved formats - 基于 AI 完整保留排版的 PDF 文档全文双语翻译,支持 Google/DeepL/Ollama/OpenAI 等服务,提供 CLI/GUI/MCP/Docker/Zotero
★ 36kDocKylin. [AAAI 2025] DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
★ 36docling. Get your documents ready for gen AI
★ 64kCTRNet-plus. The official implement of CTRNet++.
★ 15doc-matcher. Inference, training and evaluation code for our paper "DocMatcher: Document Image Dewarping via Structural and Textual Line Matching" (WACV) 2025.
★ 55TextHarmony. The official code for NeurIPS 2024 paper: Harmonizing Visual Text Comprehension and Generation
★ 127samurai. Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
★ 7.1kTextHawk. Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
★ 68Rectified-Diffusion. [ICLR 2025] Rectified Diffusion: Straightness Is Not Your Need
★ 250inksight. Jupyter Notebook
★ 1kVisualWebBench. Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"
★ 68LLM-Agent-Paper-List. The paper list of the 86-page SCIS cover paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.
★ 8.2kLLM_MultiAgents_Survey_Papers. Large Language Model based Multi-Agents: A Survey of Progress and Challenges (In IJCAI 2024)
★ 1.3kUniMERNet. UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
★ 494DTSM. Code and data for the paper: DTSM: Toward Dense Table Structure Recognition with Text Query Encoder and Adjacent Feature Aggregator
★ 14Show-o. [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
★ 2kHowToCook. Programmer's guide about how to cook at home.
★ 101kScaleDoc. Python
★ 7Campus2025. 2025届互联网校招信息汇总
★ 840DiffMatch. Official implementation of "Diffusion Model for Dense Matching" (ICLR'24 Oral)
★ 189UPOCR. Official implementation of UPOCR: Towards unified pixel-level OCR interface (ICML 2024)
★ 70Awesome-Chinese-LLM. 整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。
★ 23kDocRes. [CVPR 2024] DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
★ 629KVP10k. Repository for the KVP10k dataset
★ 23M2Doc. Jupyter Notebook
★ 43TexTeller. TexTeller can convert image to latex formulas (image2latex, latex OCR) with higher accuracy and exhibits superior generalization ability, enabling it to cover most usage scenarios.
★ 755Bridging-Text-Spotting. (CVPR 2024) Bridging the Gap Between End-to-End and Two-Step Text Spotting.
★ 75TextCoT. [ACM TOMM] Official implementation of "TextCoT: Zoom-In for Enhanced Multimodal Text-Rich Image Understanding"
★ 45AI-on-the-edge-device. Easy to use device for connecting "old" measuring units (water, power, gas, ...) to the digital world
★ 8.6kT-GATE. T-GATE: Temporally Gating Attention to Accelerate Diffusion Model for Free!
★ 418Draw-and-Understand. [ICLR2025] Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
★ 94Recommendations-Document-Image-Processing. This repository contains a paper collection of the methods for document image processing, including appearance enhancement, deshadow, dewarping, deblur, and binarization.
★ 1DocNLC. Official code for DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations (AAAI 2024)
★ 44RFUND. [MM'2024] Official release of RFUND introduced in the MM'2024 paper "PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction"
★ 21Visual-Text-Processing-survey. The official project of paper "Visual Text Processing: A Comprehensive Review and Unified Evaluation""
★ 103Awesome-LLMs-Datasets. Summarize existing representative LLMs text datasets.
★ 1.5kDeepEraser. The official code for “DeepEraser: Deep Iterative Context Mining for Generic Text Eraser”, TMM, 2024.
★ 53Recommendations-Document-Image-Processing. This repository contains a paper collection of the methods for document image processing, including appearance enhancement, deshadowing, dewarping, deblurring, binarization and so on.
★ 396FontDiffuser. [AAAI2024] FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
★ 540T-Rex. [ECCV2024] API code for T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
★ 2.7kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kLLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871Awesome-Deblurring. A curated list of resources for Image and Video Deblurring
★ 2.9kGPT-4V_OCR. Evaluation of the Optical Character Recognition (OCR) capabilities of GPT-4V(ision)
★ 128LA-DocFlatten. Code and Dataset for our paper: Layout-Aware Single-Image Document Flattening
★ 24illtrtemplate-model. Code from our paper "Template-guided Illumination Correction for Document Images with Imperfect Geometric Reconstruction " (ICCVW) 2023.
★ 29GCDRNet. [TAI 2023] Appearance Enhancement for Camera-captured Document Images in the Wild
★ 61Document-Image-Dewarping. Python
★ 69Recommendations-Diffusion-Text-Image. A paper collection of recent diffusion models for text-image generation tasks, e,g., visual text generation, font generation, text removal, text image super resolution, text editing, handwritten generation, scene text recognition and scene text detection.
★ 273WMeter-Reader. [TIM 2025] Towards Accurate Readings of Water Meters by Eliminating Transition Error: New Dataset and Effective Solution
★ 19Mnist-99.7-Accuracy-with-Pytorch. A CNN model builds with Pytorch and reaches 99.7% accuracy
★ 6BGShadowNet. Python
★ 40DSRFGD. Document Shadow Removal with Foreground detection learning from Fully Synthetic Images
★ 3BoxDiff. [ICCV 2023] BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
★ 275Union14M. [ICCV 2023] Code base for Revisiting Scene Text Recognition: A Data Perspective
★ 206SMT. [ICCV2023] This is an official implementation for "Scale-Aware Modulation Meet Transformer".
★ 215DocAligner. [PR 2025] DocAligner: Automating the Annotation of Photographed Documents Through Real-virtual Alignment
★ 110augraphy. Augmentation pipeline for rendering synthetic paper printing, faxing, scanning and copy machine processes
★ 561TabRecSet. A large scale camera-taken table detection and recognition dataset.
★ 151LATIN-Prompt. Python
★ 52VisorGPT. [NeurIPS 2023] Customize spatial layouts for conditional image synthesis models, e.g., ControlNet, using GPT
★ 138UDOP.
★ 250LLMsPracticalGuide. A curated list of practical guide resources of LLMs (LLMs Tree, Examples, Papers)
★ 10kSDT. This repository is the official implementation of Disentangling Writer and Character Styles for Handwriting Generation (CVPR 2023)
★ 1.4kImage2Paragraph. [Image 2 Text Para] Transform Image into Unique Paragraph with ChatGPT, BLIP2, OFA, GRIT, Segment Anything, ControlNet.
★ 822ColossalAI. Making large AI models cheaper, faster and more accessible
★ 41kDocTr-Plus. The official code for “Deep Unrestricted Document Image Rectification”, TMM, 2023.
★ 533Awesome-Document-Image-Rectification. A comprehensive list of awesome document image rectification papers.
★ 559UDoc-GAN. Official PyTorch implementation for ACM MM22 "UDoc-GAN: Unpaired Document Illumination Correction with Background Light Prior"
★ 25layout-parser. A Unified Toolkit for Deep Learning Based Document Image Analysis
★ 5.8kDocumentLayoutAnalysis. Document Layout Analysis resources repos for development with PdfPig.
★ 637TableMASTER-mmocr. 2nd solution of ICDAR 2021 Competition on Scientific Literature Parsing, Task B.
★ 470HorNet. [NeurIPS 2022] HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions
★ 344v2ray-core. A platform for building proxies to bypass network restrictions.
★ 47kExternal-Attention-pytorch. 🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
★ 12kAwesome-Image-based-Meter-Recognition-Reading.
★ 32free. 翻墙、免费翻墙、免费科学上网、免费节点、免费梯子、免费ss/v2ray/trojan节点、蓝灯、谷歌商店、翻墙梯子
★ 41kMarior. [ACM MM 2022] Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild
★ 26yolov7. Implementation of paper - YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
★ 14kLibtorchTutorials. This is a code repository for pytorch c++ (or libtorch) tutorial.
★ 839mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4kawesome-image-rectification. 📝 A curated list of image rectification papers.
★ 179deepfillv2. The PyTorch implementation of ICCV 2019 oral paper: free-form inpainting (deepfillv2), especially Gated Conv
★ 271SauvolaNet-torch. pytorch implement of SauvolaNet
★ 2Document_Binarization_Collection. This repository is a concise collection of well known deep learning based document binarization models.
★ 30vit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kAwesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6klabelbee-client. Out-of-the-box Annotation Toolbox
★ 395FocusOnDepth. A Monocular depth-estimation for in-the-wild AutoFocus application.
★ 158DocSegTr. A Bottom-Up Instance Segmentation Strategy for segmenting document instances using Transformers
★ 59DocEnTR. DocEnTr: An end-to-end document image enhancement transformer - ICPR 2022
★ 190sim2real-docs. Synthesize image datasets of documents in natural scenes with Python+Blender3D
★ 61BEDSR-Net_A_Deep_Shadow_Removal_Network_from_a_Single_Document_Image. Unofficial implementation of ''BEDSR-Net: A Deep Shadow Removal from a Single Document Image'' with PyTorch
★ 66DE-GAN-Implementation-Using-PyTorch. Document Enhancement Generative Adversarial Networks
★ 9DocTr. The official code for “DocTr: Document Image Transformer for Geometric Unwarping and Illumination Correction”, ACM MM, Oral Paper, 2021.
★ 438torch-toolbox. 🛠 Toolbox to extend PyTorch functionalities
★ 419Deformable-Image-Registration-Projects. Deformable Image Registration Projects
★ 265Document-Dewarping-with-Control-Points. Document Dewarping with Control Points
★ 199FCN-for-Semantic-Segmentation. Implemention of FCN-8 and FCN-16 with Keras and uses CRF as post processing
★ 177DeOldify. A Deep Learning based project for colorizing and restoring old images (and video!)
★ 18kBest_AI_paper_2020. A curated list of the latest breakthroughs in AI by release date with a clear video explanation, link to a more in-depth article, and code
★ 2.2kpylsd-nova. Python bindings for Line Segment Detector (LSD)
★ 35document_warp. A script using OpenCV to perform perspective transforms on images of documents to give a head-on view.
★ 8waveCorrection. OCR Document image deformation correction.复现阿里OCR皱巴巴文档图像形变矫正
★ 93pytorch-deeplab-xception. DeepLab v3+ model in PyTorch. Support different backbones.
★ 3ktps_stn_pytorch. PyTorch implementation of Spatial Transformer Network (STN) with Thin Plate Spline (TPS)
★ 957recursive-cnns. Implementation of the paper "Real-time Document Localization in Natural Images by Recursive Application of a CNN."
★ 145pytorch-wgan. Pytorch implementation of DCGAN, WGAN-CP, WGAN-GP
★ 807Document-Scanner. Python
★ 146Dewarping-Document-Image-By-Displacement-Flow-Estimation. Dewarping Document Image By Displacement Flow Estimation with Fully Convolutional Network
★ 195BOOK-CONTENT-SEGMENTATION-AND-DEWARPING. Using FCN to segment the book's content and background, then dewarping the pages,
★ 21page_dewarp. Text page dewarping using a "cubic sheet" model
★ 1.5kdeep-learning-for-document-dewarping. An application of high resolution GANs to dewarp images of perturbed documents
★ 153Document-Image-Dewarping. Document Image Dewarping
★ 428SCUT-HCCDoc_Dataset_Release.
★ 104Object-Detection-Evaluation-Tool. Object Detection Evaluation Tools
★ 61simple-faster-rcnn-pytorch. A simplified implemention of Faster R-CNN that replicate performance from origin paper
★ 4kDocIIW. Repository for Intrinsic Decomposition of Document Images In-the-Wild (BMVC '20)
★ 51RectiNet. A Gated and Bifurcated Stacked U-Net Module for Document Image Dewarping
★ 109EDSR-PyTorch. PyTorch version of the paper 'Enhanced Deep Residual Networks for Single Image Super-Resolution' (CVPRW 2017)
★ 2.6kGuided-Attention-Inference-Network. Contains implementation of Guided Attention Inference Network (GAIN) presented in Tell Me Where to Look(CVPR 2018). This repository aims to apply GAIN on fcn8 architecture used for segmentation.
★ 237pytorch-auto-augment. PyTorch implementation of AutoAugment.
★ 161fast-autoaugment. Official Implementation of 'Fast AutoAugment' in PyTorch.
★ 1.6kUnsupervised-Domain-Specific-Deblurring. Implementation of "Unsupervised Domain-Specific Deblurring via Disentangled Representations"
★ 113voxelmorph. Unsupervised Learning for Image Registration
★ 2.7knetwork-slimming. Network Slimming (Pytorch) (ICCV 2017)
★ 919pytorch-slimming. Learning Efficient Convolutional Networks through Network Slimming, In ICCV 2017.
★ 573pytorch-ssim. pytorch structural similarity (SSIM) loss
★ 1.9kpytorch-msssim. PyTorch differentiable Multi-Scale Structural Similarity (MS-SSIM) loss
★ 464conv_arithmetic. A technical report on convolution arithmetic in the context of deep learning
★ 15kPyNET. Generating RGB photos from RAW image files with PyNET
★ 355crnn_ctc_ocr_tf. Extremely simple implement for CRNN by Tensorflow
★ 524tesseract. Tesseract Open Source OCR Engine (main repository)
★ 76kpytorch-cnn-visualizations. Pytorch implementation of convolutional neural network visualization techniques
★ 8.2kPyramid-Attention-Networks-pytorch. Implementation of Pyramid Attention Networks for Semantic Segmentation.
★ 238CRAFT-pytorch. Official implementation of Character Region Awareness for Text Detection (CRAFT)
★ 3.4kpython-turtle-draw-svg. A python program for draw SVG file using turtle package.
★ 198gans-in-action. Companion repository to GANs in Action: Deep learning with Generative Adversarial Networks
★ 1kocr-open-dataset. list all open dataset about ocr.
★ 99DeepLearning-500-questions. 深度学习500问,以问答形式对常用的概率知识、线性代数、机器学习、深度学习、计算机视觉等热点问题进行阐述,以帮助自己及有需要的读者。 全书分为18个章节,50余万字。由于水平有限,书中不妥之处恳请广大读者批评指正。 未完待续............ 如有意合作,联系scutjy2015@163.com 版权所有,违权必究 Tan 2018.06
★ 58kSRNet. A tensorflow reproducing of paper “Editing Text in the wild”
★ 234pytorch-CycleGAN-and-pix2pix. Image-to-Image Translation in PyTorch
★ 25kImage-Inpainting. A paper summary of image inpainting
★ 840the-gan-zoo. A list of all named GANs!
★ 15k