北京海淀

WangZeJun

Expert
@zejunwang1

NLP算法工程师,硕士毕业于清华大学信息科学技术学院,关注中文自然语言处理技术与机器学习基础算法

CSTS. 中文自然语言推理与语义相似度数据集

366

LLMTuner. 大语言模型指令调优工具(支持 FlashAttention)

177

bert4vec. 一个基于预训练的句向量生成工具

138

bert_text_classification. 基于 BERT 模型的中文文本分类工具

71

chatglm_tuning. 基于 LoRA 和 P-Tuning v2 的 ChatGLM-6B 高效参数微调

55

bertorch. 基于 pytorch 的 bert 实现和下游任务微调

52

CTCDataset. 中文文本纠错数据集汇总

47

easytokenizer. 高性能文本 Tokenizer 库

31

bloom_tuning. BLOOM 模型的指令微调

24

gpt2classifier. 基于中文 GPT2 预训练模型的文本分类微调

23

fastMatch. Large-scale exact string matching tool

17

darmatch. 一个非常高效的字符串匹配工具,支持正向/反向最大匹配分词和多模式字符串精确匹配

16

gpt2ppl-zh. 基于中文 GPT2 预训练模型的语句困惑度计算

15

lightltp. 基于 onnxruntime 推理引擎的中文 ltp 词法分析

14

GPTDetector. AI生成内容检测分类器

12

simbert_distill. Two-stage SimBERT distillation

10

fastlcs. An effective tool for solving LCS problems

9

ElectraForSpellingCheck. 基于 Electra 预训练模型的中文拼写检查

3

stringutils. A bridge between Unicode encoded strings and std::string.

2

sentence-splitter. Split Chinese text and English text into sentences.

2

learn-unicode. Learn Unicode the easy way

2

confuse. 汉语字词混淆集

1

learn. Some technical learning and summary

1

csc_hard. Chinese spelling correction dataset

1

Firefly. Firefly(流萤): 中文对话式大语言模型

1

TurboTransformers. a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.

1
26
Apply