Beijing

Zilin Zhu

Elite
@zhuzilin

☀️ RL infra @Z.ai, ex WeChat AI

ring-flash-attention. Ring attention implementation with flash attention

1k

NP_ML. A tool library of classical machine learning algorithms with only numpy.

221

whisper-openvino. openvino version of openai/whisper

184

faster-nougat. Implementation of nougat that focuses on processing pdf locally.

85

pdf-with-its-own-md5. A PDF template that contains its own MD5!

45

monkey. A C++ version monkey language interpreter. From Write An Interpreter In Go

39

flash-attention-with-sink. Python

37

ncnn-swift. An example on using ncnn with Swift.

35

es. A JavaScript interpreter from scratch, supporting ES5 syntax.

30

chatgpt-desktop. Desktop version of ChatGPT, support manually set cookie

19

seastar-cn. Seastar教程(文档)中文翻译。A Chinese translation for Seastar tutorial and relevant documents.

19

pytorch-malloc. An external memory allocator example for PyTorch.

16

google-translate-desktop. Google Translate Desktop built with Electron

16

vllm-group. Python

12

NeZha. Organizing ssh servers in one shell.

10

simple-pandas. A much simpler pandas!!!

8

wandb-discord-bot. A discord bot for monitoring wandb project and runs.

8

SwiftPEG. A PEG parser generator written in swift 5.3.

8

aqt-pytorch. Python

7

li. another mini text editor with 500 loc and simple interfaces for short cut.

6

ecdh-psi. A Go Implementation of ECDH-PSI

5

cumem_allocator. Python

5

PIXEL.css. the pixel art code from the fun and famous NES.css

5

pytorch-extension-gcc. How to use gcc instead of setup.py to build an pytorch extension.

4

pytorch-reloadable-pg. PyTorch Reloadable Process Group

3

OpenRLHF. An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & Mixtral)

3

blog. my blog~

3

llama. Inference code for LLaMA models

2

on-device_recommendation_tflite. Swift

2

zhuzilin.

2

dafny-exercises. some formal verification exercises using dafny.

2

torchrec_mapper. C++

2

nvcc_daxpy_example. C++

1

sglang. SGLang is a fast serving framework for large language models and vision language models.

1

scattermoe. Triton-based implementation of Sparse Mixture of Experts.

1

unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities

1

bfc. A really fast non-JIT brainfuck interpreter.

1

triton. Development repository for the Triton language and compiler

1

electron-fc. A electron based famicom(NES) emulator

1

flash-attention. Fast and memory-efficient exact attention

1

garbled_circuit. A python implementation of Yao's GC.

1

torch_memory_saver. Allow torch tensor memory to be released and resumed later

1

autodiff. pytorch like automatic differentiation library in numpy

1
43
Apply