This is your work, valued
Statistical-Learning-Method_Code. 手写实现李航《统计学习方法》书中全部算法
★ 12kNLP-practice-program. 力求囊括主流NLP模型练手项目,不断更新中
★ 295LeetCode. Python
★ 28VT-SSum.
★ 23kleister-nda.
★ 1copilot-api. GitHub Copilot, OpenAI Codex, and third-party AI provider gateway with OpenAI and Anthropic API compatibility.
★ 944claude-code-router. One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.
★ 36kAIstudioProxyAPI. FastAPI + Playwright + Camoufox 中间层代理服务器,兼容OpenAI API且支持参数转发。项目通过浏览器自动化将API请求转发到 Google AI Studio Chat,并同样按照OpenAI标准格式返回的工具。内置调试WebUI面板。
★ 2.5kAIStudio2API. 将AI Studio反代成OpenAI兼容的API | OpenAI-compatible API proxy for Google AI Studio
★ 133office-fluent-ui-command-identifiers. Office Fluent User Interface Control Identifiers
★ 154HEU_KMS_Activator.
★ 43kBettaFish. 微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。
★ 42kdocx. Easily generate and modify .docx files with JS/TS with a nice declarative API. Works for Node and on the Browser.
★ 5.9kxiaohongshu-mcp. MCP for xiaohongshu.com
★ 15kclaude-code. Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
★ 140kimg2dataset. Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4.4kTraceBoard. 统计最常用的键盘按键
★ 57worktool. 一款安全稳定的Android无障碍服务工具,支持控制企微/微信来运行的无人值守群管理企业微信机器人
★ 2.7kAwesome-Zlibrary.
★ 2.7kAwesome-MCoT. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1koi. CSP-J/S/X, NOIP, NOI, IOI, 信息学奥林匹克竞赛历年真题收录 | QQ交流群529507453
★ 706TACO. Python
★ 240DeepSeek-R1.
★ 92kPanda-70M. [CVPR 2024] Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
★ 702MiniCPM-V. A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26kTkinter-Designer. An easy and fast way to create a Python GUI 🐍
★ 10kMMLU-CF. A Contamination-free Multi-task Language Understanding Benchmark [Official, ACL 2025]
★ 126Stable-Retro. A fork of gym-retro with additional games, emulators and supported platforms
★ 697BizHawk. BizHawk is a multi-system emulator written in C#. BizHawk provides nice features for casual gamers such as full screen, and joypad support in addition to full rerecording and debugging tools for all system cores.
★ 2.7kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12krules. Repository of yara rules
★ 4.9kMiSTer-FPGA-PORTABLE.
★ 28TIP-I2V. [ICCV 2025] TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
★ 41MMDocBench. MMDocBench: data, instructions, source code for evaluating LVLMs in fine-grained visual document understanding and grounding
★ 11warcio. Streaming WARC/ARC library for fast web archive IO
★ 463Awesome-Code-LLM. 👨💻 An awesome and curated list of best code-LLM for research.
★ 1.3kP-tuning. A novel method to tune language models. Codes and datasets for paper ``GPT understands, too''.
★ 939lm-evaluation-harness. A framework for few-shot evaluation of language models.
★ 13kwukong-robot. 🤖 wukong-robot 是一个简单、灵活、优雅的中文语音对话机器人/智能音箱项目,支持ChatGPT多轮对话能力,还可能是首个支持脑机交互的开源智能音箱项目。
★ 7.1kvector-quantize-pytorch. Vector (and Scalar) Quantization, in Pytorch
★ 4kCyrillicHandwritingPOC. Repository for contributions for Data Generation for Post-OCR correction of Cyrillic handwriting paper
★ 23arxiv-compiler. Service to compile LaTeX source packages into PDF, PostScript, and other formats
★ 14dy-auto. 抖音ffmpeg自动生成视频、字幕、自动上传发布
★ 130mobile_monitor_android_simple. 一款轻量级的监听Android手机短信电话通知->推送到微信、企业微信、自定义接口、邮箱、Bark的软件
★ 234python-docx-template. Use a docx as a jinja2 template
★ 2.7kmathpix-markdown-it. Markdown rendering + Latex extras (equations, tables, ...), with conversion features, for the scientific community
★ 677arxiv-tools. Tools to bulk download arxiv data
★ 134SynthText. Code for generating synthetic text images as described in "Synthetic Data for Text Localisation in Natural Images", Ankush Gupta, Andrea Vedaldi, Andrew Zisserman, CVPR 2016.
★ 2.1kdeep-text-recognition-benchmark. Text recognition (optical character recognition) with deep learning methods, ICCV 2019
★ 3.9kchatgpt-gzh. 一个基于公众号的chatpgt项目
★ 91imgaug. Image augmentation for machine learning experiments.
★ 15kList-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words. List of Dirty, Naughty, Obscene, and Otherwise Bad Words
★ 3.4kdeep-text-recognition-benchmark. PyTorch code of my ICDAR 2021 paper Vision Transformer for Fast and Efficient Scene Text Recognition (ViTSTR)
★ 312pypdf. A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
★ 10karxiv-latex-cleaner. arXiv LaTeX Cleaner: Easily clean the LaTeX code of your paper to submit to arXiv
★ 7kllama. Inference code for Llama models
★ 60ktuning_playbook. A playbook for systematically maximizing the performance of deep learning models.
★ 30kpix2struct. Python
★ 686torchscale. Foundation Architecture for (M)LLMs
★ 3.1kbigcode-analysis. Repository for analysis and experiments in the BigCode project.
★ 126text-dedup. All-in-one text de-duplication
★ 765ModelCenter. Efficient, Low-Resource, Distributed transformer implementation based on BMTrain
★ 270BMCook. Model Compression for Big Models
★ 171BMInf. Efficient Inference for Big Models
★ 583BMTrain. Efficient Training (including pre-training and fine-tuning) for Big Models
★ 624BMList. A List of Big Models
★ 343CPM-Live. Live Training for Open-source Big Models
★ 499pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37kSlurm4Azure. Deploy SLURM Workload Manager on Ubuntu to Microsoft Azure
★ 6stopes. A library for preparing data for machine translation research (monolingual preprocessing, bitext mining, etc.) built by the FAIR NLLB team.
★ 309polyglot. Multilingual text (NLP) processing toolkit
★ 2.4kMegatron-LM. Ongoing research training transformer models at scale
★ 17kmetaseq. Repo for external large-scale work
★ 6.6ktaming-transformers. Taming Transformers for High-Resolution Image Synthesis
★ 6.5kpython-docx. Create and modify Word documents with Python
★ 5.7kmmocr. OpenMMLab Text Detection, Recognition and Understanding Toolbox
★ 4.8kColossalAI. Making large AI models cheaper, faster and more accessible
★ 41kmmdetection. OpenMMLab Detection Toolbox and Benchmark
★ 33kDALLE-pytorch. Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5.6kcord. CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
★ 487Deformable-DETR. Deformable DETR: Deformable Transformers for End-to-End Object Detection.
★ 4kdetr. End-to-End Object Detection with Transformers
★ 15kdetectron2. Detectron2 for Document Layout Analysis
★ 189PyMuPDF. PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
★ 10kPubLayNet. Jupyter Notebook
★ 1.1kctdar_measurement_tool. Evaluation Tool for the ICDAR 2019 Competition on Table Detection and Recognition
★ 42detectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kICDAR2019_cTDaR. The ICDAR 2019 cTDaR is to evaluate the performance of methods for table detection (TRACK A) and table recognition (TRACK B). For the first track, document images containing one or several tables are provided. For TRACK B two subtracks exist: the first subtrack (B.1) provides the table region. Thus, only the table structure recognition must be performed. The second subtrack (B.2) provides no a-priori information. This means, the table region and table structure detection has to be done.
★ 180kleister-nda.
★ 61kleister-charity. Shell
★ 40L-ink_Card. Smart NFC & ink-Display Card
★ 7.6kstanza. Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages
★ 7.9konnxruntime. ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
★ 21kMatchSum. Code for ACL 2020 paper: "Extractive Summarization as Text Matching"
★ 520PreSumm. code for EMNLP 2019 paper Text Summarization with Pretrained Encoders
★ 1.3kawesome-public-datasets. A topic-centric list of HQ open datasets.
★ 78kdataset-sts. Semantic Text Similarity Dataset Hub
★ 730control-over-copying. (AAAI'20) The source code for the paper "Controlling the Amount of Verbatim Copying in Abstractive Summarization".
★ 38unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22kXLM. PyTorch original implementation of Cross-lingual Language Model Pretraining.
★ 2.9kgym. A toolkit for developing and comparing reinforcement learning algorithms.
★ 37kannotated-transformer. An annotated implementation of the Transformer paper.
★ 7.4kNLP-progress. Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the most common NLP tasks.
★ 23kallennlp. An open-source NLP research library, built on PyTorch.
★ 12ksentencepiece. Unsupervised text tokenizer for Neural Network-based text generation.
★ 12k