This is your work, valued
CodexGuide. CodexGuide:面向全球初学者、创作者、开发者与团队的 Codex 实践指南
★ 2.9khermes-agent. The agent that grows with you
★ 223klearn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kVoxBox. A large-scale speech corpus introduced in Spark-TTS, built from diverse open-source datasets for training text-to-speech (TTS) systems.
★ 115DIAL. Python
★ 99claw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195kagency-agents. A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.
★ 138kAutoCLI. AutoCLI is a Blazing fast, memory-safe command-line tool — Fetch information from any website with a single command. Covers Twitter/X, Reddit, YouTube, HackerNews, Bilibili, Zhihu, Xiaohongshu, and 55+ sites, with support for controlling Electron desktop apps, integrating local CLI tools (gh, docker, kubectl), now powered by AutoCLI.ai .
★ 2.9kgstack. Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
★ 125kdaily_stock_analysis. LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
★ 60kStar-Office-UI. A pixel office for your OpenClaw: turn invisible work states into a cozy little space with characters, daily notes, and guest agents. Code under MIT; art assets for non-commercial learning only.
★ 7.4kawesome-openclaw-tutorial. 从零开始玩转OpenClaw:最全面的中文教程,涵盖安装、配置、实战案例和避坑指南(github版)
★ 4.5kawesome-openclaw-skills. The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
★ 52kPageIndex. 📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
★ 35kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kOpenBB. Open Data Platform for analysts, quants and AI agents.
★ 71kopencode. The open source coding agent.
★ 191kclaude-code. Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
★ 140keasy-rl. 强化学习中文教程(蘑菇书🍄),在线阅读地址:https://datawhalechina.github.io/easy-rl/
★ 14kcosyvoice-paimon-sft. Fine-tune the Paimon speech using the CosyVoice2 model
★ 10multiwoz. Source code for end-to-end dialogue model from the MultiWOZ paper (Budzianowski et al. 2018, EMNLP)
★ 954MT-LLM. The implementation for "Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions"
★ 51starVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3kARPO. [ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)
★ 1.1kaxlearn. An Extensible Deep Learning Library
★ 2.4kEmbodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kRD-Agent. Research and development (R&D) is crucial for the enhancement of industrial productivity, especially in the AI era, where the core aspects of R&D are mainly focused on data and models. We are committed to automating these high-value generic R&D processes through R&D-Agent, which lets AI drive data-driven AI. 🔗https://aka.ms/RD-Agent-Tech-Report
★ 14kai-hedge-fund. An AI Hedge Fund Team
★ 62kUI-TARS. Pioneering Automated GUI Interaction with Native Agents
★ 11kverl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kA2A. Agent2Agent (A2A) is an open protocol enabling communication and interoperability between opaque agentic applications.
★ 25kchunking_evaluation. This package, developed as part of our research detailed in the Chroma Technical Report, provides tools for text chunking and evaluation. It allows users to compare different chunking methods and includes implementations of several novel chunking strategies.
★ 501Langchain-Chatchat. Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
★ 38kevalscope. A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
★ 3.2kQwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kdeep-research. An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine its research direction overtime and deep dive into a topic.
★ 19kPIKE-RAG. PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation
★ 2.5kczsc. 缠中说禅技术分析工具;缠论;股票;期货;Quant;量化交易
★ 5.6kstock. stock,股票系统。使用python进行开发。
★ 7.9kcamel. 🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org
★ 18kowl. 🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
★ 20kOpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58kAwesome-System2-Reasoning-LLM. Latest Advances on System-2 Reasoning
★ 1.4kai. This repository will have different projects using AutoGen and Tutorials
★ 1.1kdeep-searcher. Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.
★ 8kFlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13kautogen. A programming framework for agentic AI
★ 60kAutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kVLM-R1. Solve Visual Understanding with Reinforced VLMs
★ 6kdataline. Chat with your data - AI data analysis and visualization on CSV, Postgres, MySQL, Snowflake, SQLite...
★ 1.6kpandas-ai. Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.
★ 24ks1. s1: Simple test-time scaling
★ 6.7krllm. Democratizing Reinforcement Learning for LLMs
★ 5.7ksimple-evals. Python
★ 4.6kLogic-RL. Reproduce R1 Zero on Logic Puzzle
★ 2.5kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26kJanus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18knicegui. Create web-based user interfaces with Python. The nice way.
★ 16kstorm. An LLM-powered knowledge curation system that researches a topic and generates a full-length report with citations.
★ 30kgpt-assistant-android. 【新增智能体模式】安卓端全场景GPT助手,可用音量键唤起并进行语音交流,支持联网、拍照、模板、附件解析、智能体模式等 | GPT assistant for Android, activated via volume keys for voice interaction, supporting features such as networking, taking photos, templates, parsing PDF and Office documents, and agent mode.
★ 888HuatuoGPT-II. HuatuoGPT2, One-stage Training for Medical Adaption of LLMs. (An Open Medical GPT)
★ 411KAG. KAG is a logical form-guided reasoning and retrieval framework based on OpenSPG engine and LLMs. It is used to build logical reasoning and factual Q&A solutions for professional domain knowledge bases. It can effectively overcome the shortcomings of the traditional RAG vector similarity calculation model.
★ 8.9kBert-VITS2. vits2 backbone with multilingual-bert
★ 8.8kChatTTS. A generative speech model for daily dialogue.
★ 40kpiper. A fast, local neural text to speech system
★ 11kedge-tts. Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
★ 12kfish-speech. SOTA Open Source TTS
★ 32kCosyVoice. Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
★ 22kgenesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30klarge_concept_model. Large Concept Models: Language modeling in a sentence representation space
★ 2.4kmPLUG-DocOwl. mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2.4kFinMem-LLM-StockTrading. FinMem: A Performance-Enhanced LLM Trading Agent with Layered Memory and Character Design
★ 931MockingBird. 🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 37kShowUI. [CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.
★ 1.9kPapers-in-100-Lines-of-Code. Implementation of papers in 100 lines of code.
★ 2.8kFinRobot. FinRobot: An Open-Source AI Agent Platform for Financial Applications using LLMs 🚀 🚀 🚀
★ 7.7kreplicate-python. Python client for Replicate
★ 911HunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kgenerative_agents. Generative Agents: Interactive Simulacra of Human Behavior
★ 22kdata-juicer. Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27kSWE-bench. SWE-bench: Can Language Models Resolve Real-world Github Issues?
★ 5.5kAwesome-Unified-Multimodal-Models. 📖 This is a repository for organizing papers, codes and other resources related to unified multimodal models.
★ 829open-instruct. AllenAI's post-training codebase
★ 3.8kChinese-LLaMA-Alpaca-3. 中文羊驼大模型三期项目 (Chinese Llama-3 LLMs) developed from Meta Llama 3
★ 2kMedicalGPT. MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。
★ 5.7kProX. [ICML 2025] Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
★ 273BitNet. Official inference framework for 1-bit LLMs
★ 40kEmu3. Next-Token Prediction is All You Need
★ 2.4kswarm. Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
★ 22kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kOpenHands. 🙌 OpenHands: AI-Driven Development
★ 83kChatDev. ChatDev 2.0: Dev All through LLM-powered Multi-Agent Collaboration
★ 34kMetaGPT. 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
★ 70kLightRAG. [EMNLP2025] "LightRAG: Simple and Fast Retrieval-Augmented Generation"
★ 38kopenai-python. The official Python library for the OpenAI API
★ 31kAwesome-Mamba-Papers. Awesome Papers related to Mamba.
★ 1.4kAwesome-Mamba-Collection. A curated collection of papers, tutorials, videos, and other valuable resources related to Mamba.
★ 753Awesome-Code-LLM. [TMLR] A curated list of language modeling researches for code (and other software engineering activities), plus related datasets.
★ 3.4kFirefly. Firefly: 大模型训练工具,支持训练Qwen2.5、Qwen2、Yi1.5、Phi-3、Llama3、Gemma、MiniCPM、Yi、Deepseek、Orion、Xverse、Mixtral-8x7B、Zephyr、Mistral、Baichuan2、Llma2、Llama、Qwen、Baichuan、ChatGLM2、InternLM、Ziya2、Vicuna、Bloom等大模型
★ 6.7kBayLing-Speech. LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
★ 3.1kMemoRAG. Empowering RAG with a memory-based data interface for all-purpose applications!
★ 2.3kQwen-Agent. Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.
★ 17kkotaemon. An open-source RAG-based tool for chatting with your documents.
★ 26kMambaOut. MambaOut: Do We Really Need Mamba for Vision? (CVPR 2025)
★ 2.7kLaw_of_Vision_Representation_in_MLLMs. [COLM'25] Official implementation of the Law of Vision Representation in MLLMs
★ 176llama-models. Utilities intended for use with Llama models.
★ 7.7kmamba. The Fast Cross-Platform Package Manager
★ 8.1kflux. Official inference repo for FLUX.1 models
★ 26kmini-omni. open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
★ 3.6kwriting-in-the-margins. Python
★ 121TAG-Bench. TAG-Bench: A benchmark for table-augmented generation (TAG)
★ 764llm-action. 本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
★ 25kMegatron-LM. Ongoing research training transformer models at scale
★ 17kms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kPai-Megatron-Patch. The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
★ 1.6kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kMegatron-DeepSpeed. Ongoing research training transformer language models at scale, including: BERT & GPT-2
★ 1.4kMegatron-DeepSpeed. Ongoing research training transformer language models at scale, including: BERT & GPT-2
★ 2.3kDecryptPrompt. 总结Prompt&LLM论文,开源数据&模型,AIGC应用
★ 3.4kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kmem0. Universal memory layer for AI Agents
★ 62kllm-paper-daily. Daily updated LLM papers. 每日更新 LLM 相关的论文,欢迎订阅 👏 喜欢的话动动你的小手 🌟 一个
★ 1.3kSEED-Story. SEED-Story: Multimodal Long Story Generation with Large Language Model
★ 884cambrian. Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2kmllm_interview_note. 主要记录大语言大模型(LLMs) 算法(应用)工程师多模态相关知识
★ 288llm_interview_note. 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
★ 15kllama3-from-scratch. llama3 implementation one matrix multiplication at a time
★ 15kMQ-Det. Official PyTorch implementation of "Multi-modal Queried Object Detection in the Wild" (accepted by NeurIPS 2023)
★ 346GroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10ktransformers. 🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
★ 163kHunyuanDiT. Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
★ 4.3kleetcode-editor. Do Leetcode exercises in IDE, support leetcode.com and leetcode-cn.com, to meet the basic needs of doing exercises.Support theoretically: IntelliJ IDEA PhpStorm WebStorm PyCharm RubyMine AppCode CLion GoLand DataGrip Rider MPS Android Studio
★ 4kleetcode. LeetCode Solutions: A Record of My Problem Solving Journey.( leetcode题解,记录自己的leetcode解题之路。)
★ 56kleetcode. LeetCode题解,151道题完整版。
★ 11kmergekit. Tools for merging pretrained large language models.
★ 7.3kJSON2YOLO. Legacy JSON-to-YOLO dataset converter for COCO, LabelMe, Labelbox, VoTT, INFOLKS, and ATH annotations. Superseded by convert_coco() in the Ultralytics package.
★ 1.2kultralytics. Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
★ 60kdocx2tex. Converts Microsoft Word docx to LaTeX
★ 669Open-Sora-Plan. This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
★ 12konnx2tf. A tool for converting ONNX files to LiteRT/TFLite/TensorFlow, PyTorch native code (nn.Module), TorchScript (.pt), state_dict (.pt), Exported Program (.pt2), and Dynamo ONNX. It also supports direct conversion from LiteRT to PyTorch.
★ 985nngen. Jupyter Notebook
★ 522YOLO-World. [CVPR 2024] Real-Time Open-Vocabulary Object Detection
★ 6.5kRPG-DiffusionMaster. [ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
★ 1.8kresume. An elegant \LaTeX\ résumé template. 大陆镜像 https://gods.coding.net/p/resume/git
★ 11kawesome-resume-for-chinese. :page_facing_up: 适合中文的简历模板收集(LaTeX,HTML/JS and so on)由 @hoochanlon 维护
★ 8.1kInstantID. InstantID: Zero-shot Identity-Preserving Generation in Seconds 🔥
★ 12konnx-tensorflow. Tensorflow Backend for ONNX
★ 1.3kPyTorch-ONNX-TFLite. Conversion of PyTorch Models into TFLite
★ 401onnx2tflite. Tool for onnx->keras or onnx->tflite. Hope this tool can help you.
★ 579unstructured. Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
★ 15kAuto-PPT. Auto generate pptx using gpt-3.5, Free to use online / 通过gpt-3.5生成PPT,免费在线使用
★ 755pptxtojson. Office PowerPoint(.pptx) file to JSON | 将 PPTX 文件转为可读的 JSON 数据
★ 447PPTX2HTML. Convert pptx file to HTML by using pure javascript
★ 629python-pptx. Create Open XML PowerPoint documents in Python
★ 3.5kmd2googleslides. Generate Google Slides from markdown
★ 4.7kEmu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kpptx2md. a pptx to markdown converter
★ 1.3kemu2. Simple x86 and DOS emulator for the Linux terminal.
★ 463pdf2pptx. Convert your (Beamer) PDF slides to (Powerpoint) PPTX
★ 462CogVLM. a state-of-the-art-level open visual language model | 多模态预训练模型
★ 6.7kinstruct-pix2pix. Python
★ 6.9kUniControl. Unified Controllable Visual Generation Model
★ 662Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kLLaMA2-Accessory. An Open-source Toolkit for LLM Development
★ 2.8kvllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kOBELICS. Code used for the creation of OBELICS, an open, massive and curated collection of interleaved image-text web documents, containing 141M documents, 115B text tokens and 353M images.
★ 217VisualGLM-6B. Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型
★ 4.2kSALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kViT-Lens. [CVPR 2024] ViT-Lens: Towards Omni-modal Representations
★ 189AudioLDM2. Text-to-Audio/Music Generation
★ 2.6kSEED. Official implementation of SEED-LLaMA (ICLR 2024).
★ 642SEED-Bench. (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
★ 366audiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24ksft_datasets. 开源SFT数据集整理,随时补充
★ 583datasets. 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
★ 22kspeechbrain. A PyTorch-based Speech Toolkit
★ 12kCIF-HieraDist. [INTERSPEECH 2023] Knowledge Transfer from Pre-trained Language Models to Cif-based Recognizers via Hierarchical Distillation
★ 41X-LLM. X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages
★ 318Click4Caption. A visual LLM for image region description or QA.
★ 16QA-CLIP. Chinese CLIP models with SOTA performance.
★ 63Video-LLaMA. [EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
★ 3.1kwhisper. Robust Speech Recognition via Large-Scale Weak Supervision
★ 106kMacaw-LLM. Macaw-LLM: Multi-Modal Language Modeling with Image, Video, Audio, and Text Integration
★ 1.6kmusiclm-pytorch. Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch
★ 3.3kSpeechGPT. SpeechGPT Series: Speech Large Language Models
★ 1.4kmPLUG-Owl. mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
★ 2.5kAwesome-Multimodal-LLM. Research Trends in LLM-guided Multimodal Learning.
★ 355FlagAI. FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.
★ 3.9kONE-PEACE. A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
★ 1.1kImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1ki-Code. Jupyter Notebook
★ 1.7kpeft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21klora-scripts. SD-Trainer. LoRA & Dreambooth training scripts & GUI use kohya-ss's trainer, for diffusion model.
★ 6.1kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kChinese-BERT-wwm. Pre-Training with Whole Word Masking for Chinese BERT(中文BERT-wwm系列模型)
★ 10k