This is your work, valued
Actions speak louder than words
AMRBART. Code for our paper "Graph Pre-training for AMR Parsing and Generation" in ACL2022
★ 106RGAT-ABSA. Code for our paper "Investigating Typed Syntactic Dependencies for Targeted Sentiment Classification Using Graph Attention Neural Network" in TASLP
★ 53Sem-Dialogue. Code for our paper "Semantic Representation for Dialogue Modeling" in ACL2021
★ 46AMR-DomainAdaptation. Code for our paper "Cross-domain Generalization" in EMNLP2022
★ 10Sem-PLM. Python
★ 9Gen-Backparsing. Code for our paper "Online Back-Parsing for AMR-To-Text Generation" in EMNLP 2020
★ 6AMR-Process. Code for preprocessing AMR graphs.
★ 6Med-LLaMA. Python
★ 5BiAAE. Code for our paper "A Bilingual Adversarial Autoencoder for Unsupervised Bilingual Lexicon Induction" in TASLP
★ 5DoubanSpyder. 豆瓣电影及影评的爬虫
★ 5Word-Representations-Reading-List. Papers for mono/cross-lingual word representations
★ 4TinyNLP. 自然语言处理简单工具包
★ 4LLM-ConParsing.
★ 4Dockerfile-pytorch. Dockerfile for building pytorch docker images
★ 4GraduationProject. Cross-lingual word embeddings via KCCA
★ 4BiKCCA. Code for our paper "Improving Vector Space Word Representations Via Kernel Canonical Correlation Analysis" in TALLIP
★ 3goodbai-nlp.github.io. HTML
★ 3Paper-Reimplement. Reimplement of some interesting papers
★ 3MTPE. Python
★ 2S-LSTM-nmt.
★ 2MlInAction. 人工智障&&深度胡扯
★ 2Bilingual-lexicon-survey. 词典抽取任务论文调研
★ 2NLP-assignment2016. 自然语言处理基础--张岳课程作业
★ 2amr-masking. Python
★ 1BlogDoc. muyeby的学习笔记
★ 1MT-Reading-List. A machine translation reading list maintained by Tsinghua Natural Language Processing Group
★ 1DialogRE-baselines. Simple yet strong baselines for dialogue relation extraction task.
★ 1Papers. Paper Reading List
★ 11303101-6. the system of managing paper
★ 1CheatSheet. 各种常用的CheatSheet
★ 1ChineseMedLLaMA. Python
★ 1Embedded-Systems-Programming. 哈工大实时嵌入式系统
★ 1TB_Spider. 百度贴吧爬虫,爬取百度贴吧的帖子,图片
★ 1Biling-Embeddings. Some code for unsupervised bilingual embedding survey
★ 1Gladoscheckin. Glados autocheckin
★ 1PatternClassification. Homework for Pattern Classification (2017 Autumn), Harbin Institute of Technology
★ 1slime. slime is an LLM post-training framework for RL Scaling.
★ 7.7kBetterDisplay. Unlock your displays on your Mac! Flexible HiDPI scaling, XDR/HDR extra brightness, virtual screens, DDC control, extra dimming, PIP/streaming, EDID override and lots more!
★ 33kAwesome-Reasoning. A curated list of papers and resources on reasoning based on LLMs, MLLMs, and graph-augmented LLMs.
★ 6LLM-RL-Visualized. 🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )
★ 4.7kmobaxterm-crack. 破解MobaXterm的高级版,生成密钥,支持几乎所有版本。
★ 549ChatKBQA. [ACL 2024] Official resources of "ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models".
★ 339Nabla-Reasoner. [ICLR'26] "Nabla-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space" by Peihao Wang*, Ruisi Cai*, Zhen Wang, Hongyuan Mei, Qiang Liu, Pan Li, Zhangyang Wang
★ 35Qwen3.6. Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
★ 3.7knsfc. nsfc - 国家自然科学基金项目LaTeX模版(面地青CBA)
★ 1.3kDynasor. [NeurIPS 2025] Simple extension on vLLM to help you speed up reasoning model without training.
★ 232PyAV. Pythonic bindings for FFmpeg's libraries.
★ 3.2kparametric-faithfulness. Jupyter Notebook
★ 23nano-vllm. Nano vLLM
★ 2nano-vllm. Nano vLLM
★ 15kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kVisionLLM. VisionLLM Series
★ 1.2kLMM-Det. Make Large Multimodal Models excel in object detection, ICCV 2025
★ 65LatentCoT-Horizon. 📖 This is a repository for organizing papers, codes, and other resources related to Latent Reasoning.
★ 406continual-learning. PyTorch implementation of various methods for continual learning (XdG, EWC, SI, LwF, FROMP, DGR, BI-R, ER, A-GEM, iCaRL, Generative Classifier) in three different scenarios.
★ 1.9kawesome-llm-implicit-reasoning.
★ 117iina. The modern video player for macOS.
★ 46kGladoscheckin. Glados autocheckin
★ 1Glados-Railgun-checkin. Glados/Space/Railgun 自动签到
★ 202Awesome-LLM4Graph-Papers. [KDD'2024] "LLM4Graph: A Survey of Large Language Models for Graphs"
★ 369glados_automation. [自用]glados自动签到
★ 7TTRL. [NeurIPS 2025] TTRL: Test-Time Reinforcement Learning
★ 1.1kAwesome-RL-for-LRMs. A Survey of Reinforcement Learning for Large Reasoning Models
★ 2.5kAwesome-Efficient-Reasoning-LLMs. [TMLR 2025] Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
★ 786acl-anthology. Data and software for building the ACL Anthology.
★ 768LLM-ConParsing.
★ 4transformers-copy-mechanism. Overwrite huggingface BART and GPT with copy mechanism
★ 21sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kcutlass. CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10kCLAP. Python
★ 7acad-homepage.github.io. AcadHomepage: A Modern and Responsive Academic Personal Homepage
★ 2.9kOverleaf-Bib-Helper. Enhances Overleaf by allowing article searches and BibTeX retrieval from DBLP and Google Scholar | 通过允许从 DBLP 和 Google Scholar 进行文章搜索和获取 BibTeX 来增强 Overleaf。
★ 128awesome-llm-self-reflection. augmented LLM with self reflection
★ 144GNP. Code for AAAI'24 paper "Graph Neural Prompting with Large Language Models".
★ 39AgenticRAG-Survey. Agentic-RAG explores advanced Retrieval-Augmented Generation systems enhanced with AI LLM agents.
★ 1.7kdart. Dataset for NAACL 2021 paper: "DART: Open-Domain Structured Data Record to Text Generation"
★ 158bert_score. BERT score for text generation
★ 1.9kmemorisation-profiles. This is the official implementation for our ACL 2024 paper: "Causal Estimation of Memorisation Profiles".
★ 25AnyGraph. "AnyGraph: Graph Foundation Model in the Wild"
★ 226markitdown. Python tool for converting files and office documents to Markdown.
★ 170kDFireDataset. D-Fire: an image data set for fire and smoke detection.
★ 415leetcode. LeetCode in pure C
★ 3.2kLangBridge. [ACL 2024] LangBridge: Multilingual Reasoning Without Multilingual Supervision
★ 97KG-LLM-Papers. [Paper List] Papers integrating knowledge graphs (KGs) and large language models (LLMs)
★ 2.2kdygiepp. Span-based system for named entity, relation, and event extraction.
★ 590MIC. MMICL, a state-of-the-art VLM with the in context learning ability from ICL, PKU
★ 361Awesome-LLM-KG. Awesome papers about unifying LLMs and KGs
★ 2.6kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kQwen3. Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
★ 27klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kLLaVA-NeXT. Python
★ 4.7kGraph-LLM. Exploring the Potential of Large Language Models (LLMs) in Learning on Graphs
★ 319cuda-samples. Samples for CUDA Developers which demonstrates features in CUDA Toolkit
★ 9.4knccl-tests. NCCL Tests
★ 1.6kSimPO. [NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
★ 956Megatron-LM. Ongoing research training transformer models at scale
★ 17kAwesome-LLM-for-NLP.
★ 105intro-llm-rag. LLM Models and RAG Hands-on guide
★ 307ccf-deadlines. ⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
★ 9.2kWhispering-LLaMA. EMNLP 23 - Integrating Whisper Encoder to LLaMA Decoder for Generative ASR Error Correction
★ 271awesome_Chinese_medical_NLP. 中文医学NLP公开资源整理:术语集/语料库/词向量/预训练模型/知识图谱/命名实体识别/QA/信息抽取/模型/论文/etc
★ 2.6kChinese-medical-dialogue-data. Chinese medical dialogue data 中文医疗对话数据集
★ 1.7kChinese-LLaMA-Alpaca-3. 中文羊驼大模型三期项目 (Chinese Llama-3 LLMs) developed from Meta Llama 3
★ 2kmultilingual-t5. Python
★ 1.3kChatGLM3. ChatGLM3 series: Open Bilingual Chat LLMs | 开源双语对话语言模型
★ 14kgraph_ensemble_learning. Graph Ensemble Learning
★ 40llama3. The official Meta Llama 3 GitHub site
★ 29kOneLLM. [CVPR 2024] OneLLM: One Framework to Align All Modalities with Language
★ 666cobra. [AAAI-25] Cobra: Extending Mamba to Multi-modal Large Language Model for Efficient Inference
★ 295llava-phi. Python
★ 400CausalNLP_Papers. A reading list for papers on causality for natural language processing (NLP)
★ 696causal-text-papers. Curated research at the intersection of causal inference and natural language processing.
★ 819LanguageBind. 【ICLR 2024🔥】 Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
★ 883BakLLaVA. Python
★ 724Huatuo-Llama-Med-Chinese. Repo for BenCao [original name: HuaTuo (华驼)], Instruction-tuning Large Language Models with Chinese Medical Knowledge. 本草(原名:华驼)模型仓库,基于中文医学知识的大语言模型指令微调
★ 5kLLM-Agent-Paper-List. The paper list of the 86-page SCIS cover paper "The Rise and Potential of Large Language Model Based Agents: A Survey" by Zhiheng Xi et al.
★ 8.2kNSFC-application-template-latex. 国家自然科学基金申请书正文(面上项目)LaTeX 模板(非官方)
★ 1.1kChineseResearchLaTeX. 中国科研常用LaTeX模板集
★ 2.6kGraphGPT. [SIGIR'2024] "GraphGPT: Graph Instruction Tuning for Large Language Models"
★ 834LLaVA-Med. Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
★ 2.2kOFA. Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
★ 2.6kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kBLIP. PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
★ 5.7kBiomedGPT. BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
★ 708SALMONN. SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kICD. Code & Data for our Paper "Alleviating Hallucinations of Large Language Models through Induced Hallucinations"
★ 71ContrastiveDecoding. contrastive decoding
★ 206ray. Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
★ 43kAwesome-Language-Model-on-Graphs. A curated list of papers and resources based on "Large Language Models on Graphs: A Comprehensive Survey" (TKDE)
★ 996GRIT. This is an official implementation for "GRIT: Graph Inductive Biases in Transformers without Message Passing".
★ 133modelscope. ModelScope: bring the notion of Model-as-a-Service to life.
★ 9.1kAwesome-LLM. Awesome-LLM: a curated list of Large Language Model
★ 27kAwesome-Graph-LLM. A collection of AWESOME things about Graph-Related LLMs.
★ 2.4ktrlx. A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)
★ 4.8kcode4struct. Official repo for ACL 2023 paper Code4Struct: Code Generation for Few-Shot Structured Prediction from Natural Language.
★ 44xuniren. HTML
★ 601Fay. fay是一个帮助数字人(2.5d、3d、移动、pc、网页)或大语言模型(openai兼容、deepseek)连通业务系统的agent框架。
★ 13kzotcard. ZotCard is a plug-in for Zotero, which is a card note-taking enhancement tool. It provides card templates (such as concept card, character card, golden sentence card, etc., by default, you can customize other card templates), so you can write cards quickly. In addition, it helps you sort cards and standardize card formats.
★ 639zotero-style. Ethereal Style for Zotero
★ 5.2kStanford-CS231n-2021-and-2022. Notes and slides for Stanford CS231n 2021 & 2022 in English. I merged the contents together to get a better version. Assignments are not included. 斯坦福cs231n的课程笔记(英文版本,不含实验代码),将2021与2022两年的课程进行了合并,分享以供交流。
★ 27awesome-digital-human. Digital Human Resource: 2D/3D/4D Human Modeling, Avatar Generation & Animation, Clothed People Digitalization, Virtual Try-On, etc.
★ 2kFlagEmbedding. Retrieval and Retrieval-augmented LLMs
★ 12kReading-List. A paper reading list maintained by ICI-MT, contains all papers and PPTs (if available) shared by our groups.
★ 11ControlPrefixes. Python
★ 90OpenBA. Python
★ 95trl. Train transformer language models with reinforcement learning.
★ 19kdrawio-desktop. Official electron build of draw.io
★ 62kgpt-neox. An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries
★ 7.4kLightLLM. LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
★ 4.2kmedmcqa. A large-scale (194k), Multiple-Choice Question Answering (MCQA) dataset designed to address realworld medical entrance exam questions.
★ 287alpaca_farm. A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
★ 845Chinese-LLaMA-Alpaca. 中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
★ 19kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kpubmedqa. PubMedQA: A Dataset for Biomedical Research Question Answering
★ 434ninja. a small build system with a focus on speed
★ 13kconvsumx. Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved Annotation
★ 19lit-llama. Implementation of the LLaMA language model based on nanoGPT. Supports flash attention, Int8 and GPTQ 4bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed.
★ 6.1ktransformers-bloom-inference. Fast Inference Solutions for BLOOM
★ 566FasterTransformer. Transformer related optimization, including BERT, GPT
★ 6.4ktext-generation-inference. Large Language Model Text Generation Inference
★ 11kllm-security. New ways of breaking app-integrated LLMs
★ 2.1kUIE. Unified Structure Generation for Universal Information Extraction
★ 958vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kalpaca_lora_4bit. Python
★ 534flash-attention. Fast and memory-efficient exact attention
★ 25kopen-instruct. AllenAI's post-training codebase
★ 3.8kBaichuan-7B. A large-scale 7B pretraining language model developed by BaiChuan-Inc.
★ 5.7kmamba. The Fast Cross-Platform Package Manager
★ 8.1kPandaLM. Python
★ 926openconnect. Mirror of the official openconnect repository
★ 462BioMedLM. Python
★ 643composer. Supercharge Your Model Training
★ 5.5kexamples. Fast and flexible reference benchmarks
★ 466Open-Assistant. OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
★ 37kpandallm. Panda项目是于2023年5月启动的开源海外中文大语言模型项目,致力于大模型时代探索整个技术栈,旨在推动中文自然语言处理领域的创新和合作。
★ 1kzotero-updateifs. 从唯问更新影响因子Update IFs from Justscience,其他一系列工具,详见Readme
★ 395zotero-updateifsE. Green Frog https://github.com/redleafnew/zotero-updateifs 的easyScholar数据版。更新影响因子,其他一系列工具,详见Readme
★ 912zotero-better-notes. Everything about note management. All in Zotero.
★ 8kPMC-LLaMA. The official codes for "PMC-LLaMA: Towards Building Open-source Language Models for Medicine"
★ 679LLM-Zoo. LLM Zoo collects information of various open- and close-sourced LLMs
★ 270open_llama. OpenLLaMA, a permissively licensed open source reproduction of Meta AI’s LLaMA 7B trained on the RedPajama dataset
★ 7.5kflan-alpaca-lora. This repository contains the code to train flan t5 with alpaca instructions and low rank adaptation.
★ 50awesome-instruction-dataset. A collection of open-source dataset to train instruction-following LLMs (ChatGPT,LLaMA,Alpaca)
★ 1.2kLMFlow. An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
★ 8.5kDeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kNER-datasets. Datasets to train supervised classifiers for Named-Entity Recognition in different languages (Portuguese, German, Dutch, French, English)
★ 343peft. 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
★ 21kpyllama. LLaMA: Open and Efficient Foundation Language Models
★ 2.8kMegatron-DeepSpeed. Ongoing research training transformer language models at scale, including: BERT & GPT-2
★ 1.4kalpaca-lora. Instruct-tune LLaMA on consumer hardware
★ 19kmedAlpaca. LLM finetuned for medical question answering
★ 563AGIEval. Python
★ 774ColossalAI. Making large AI models cheaper, faster and more accessible
★ 41kllama. Inference code for Llama models
★ 60kstanford_alpaca. Code and documentation to train Stanford's Alpaca models, and generate the data.
★ 30kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kGPT-4-LLM. Instruction Tuning with GPT-4
★ 4.3kgpt_academic. 为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
★ 71kawesome-ChatGPT-resource-zh. 精选 OpenAI 的 [ChatGPT](https://chat.openai.com) 资源清单, 跟随最新资源并添加中文相关Work
★ 676sacremoses. Python port of Moses tokenizer, truecaser and normalizer
★ 497AMRBART. Code for our paper "Graph Pre-training for AMR Parsing and Generation" in ACL2022
★ 106XAMR. Python
★ 11self-attentive-parser. High-accuracy NLP parser with models for 11 languages.
★ 911Awesome-ChatGPT. ChatGPT资料汇总学习,持续更新......
★ 4.2kbest_AI_papers_2022. A curated list of the latest breakthroughs in AI (in 2022) by release date with a clear video explanation, link to a more in-depth article, and code.
★ 3.2kTop-AI-Conferences-Paper-with-Code. MLNLP:本仓库整理人工智能会议(如 ACL、EMNLP、NAACL、COLING、AAAI、IJCAI、ICLR、NeurIPS、ICML 等)中开源代码的论文。
★ 2.7kevaluate. 🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
★ 2.5konedrive. OneDrive Client for Linux
★ 13kbamboo-amr-benchmark. Yacc
★ 2DeepSpeedExamples. Example models using DeepSpeed
★ 6.8kpytorch_distributed. The test of different distributed-training methods on High-Flyer AIHPC
★ 27bert-as-language-model. BERT as language model, fork from https://github.com/google-research/bert
★ 248parsing-as-pretraining. Parsing only with Pretraining Networks
★ 16xl-sum. This repository contains the code, data, and models of the paper titled "XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages" published in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021.
★ 277vscode-languagetool-linter. A from scratch redesign of LanguageTool integration for VS Code.
★ 186vscode-languagetool. LanguageTool grammar checking for Visual Studio Code.
★ 76vscode-ltex. LTeX: Grammar/spell checker :mag::heavy_check_mark: for VS Code using LanguageTool with support for LaTeX :mortar_board:, Markdown :pencil:, and others
★ 906languagetool. Style and Grammar Checker for 25+ Languages
★ 15kpointer-net-for-nested. The official implementation of ACL2022``Bottom-Up Constituency Parsing and Nested Named Entity Recognition with Pointer Networks''
★ 34clash-verge. A Clash GUI based on tauri. Supports Windows, macOS and Linux.
★ 22kbiblatex. biblatex is a sophisticated bibliography system for LaTeX users. It has considerably more features than traditional bibtex and supports UTF-8
★ 590ChatGPT. Reverse engineered ChatGPT API
★ 28kChatGPT_Extension. ChatGPT Extension is a really simple Chrome Extension (manifest v3) that you can access OpenAI's ChatGPT from anywhere on the web.
★ 459gluon-nlp-1. Code repo for "Language Models with Transformers" paper
★ 22amr-guidelines.
★ 278Crossling-AMR-Eval. Python
★ 1EasyNMT. Easy to use, state-of-the-art Neural Machine Translation for 100+ languages
★ 1.3kThesis-latex. my Ph.D. thesis (Zhejiang University)
★ 38Awesome-Dataset-Distillation. A curated list of awesome papers on dataset distillation and related applications.
★ 2keffect-language-amr-structure. Annotations for the LAW 2022 paper "Effect of Source Language on AMR Structure" (Wein et al., 2022).
★ 1paws. This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification.
★ 571zjuthesis. Zhejiang University Graduation Thesis LaTeX Template
★ 3.7kPrefixTuning. Prefix-Tuning: Optimizing Continuous Prompts for Generation
★ 961MSE-AMR. Python
★ 7