☀️ RL infra @Z.ai, ex WeChat AI
ring-flash-attention. Ring attention implementation with flash attention
1kNP_ML. A tool library of classical machine learning algorithms with only numpy.
221whisper-openvino. openvino version of openai/whisper
184faster-nougat. Implementation of nougat that focuses on processing pdf locally.
85pdf-with-its-own-md5. A PDF template that contains its own MD5!
45monkey. A C++ version monkey language interpreter. From Write An Interpreter In Go
39flash-attention-with-sink. Python
37ncnn-swift. An example on using ncnn with Swift.
35es. A JavaScript interpreter from scratch, supporting ES5 syntax.
30chatgpt-desktop. Desktop version of ChatGPT, support manually set cookie
19seastar-cn. Seastar教程(文档)中文翻译。A Chinese translation for Seastar tutorial and relevant documents.
19pytorch-malloc. An external memory allocator example for PyTorch.
16google-translate-desktop. Google Translate Desktop built with Electron
16vllm-group. Python
12NeZha. Organizing ssh servers in one shell.
10simple-pandas. A much simpler pandas!!!
8wandb-discord-bot. A discord bot for monitoring wandb project and runs.
8SwiftPEG. A PEG parser generator written in swift 5.3.
8aqt-pytorch. Python
7li. another mini text editor with 500 loc and simple interfaces for short cut.
6ecdh-psi. A Go Implementation of ECDH-PSI
5cumem_allocator. Python
5PIXEL.css. the pixel art code from the fun and famous NES.css
5pytorch-extension-gcc. How to use gcc instead of setup.py to build an pytorch extension.
4pytorch-reloadable-pg. PyTorch Reloadable Process Group
3OpenRLHF. An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & Mixtral)
3blog. my blog~
3llama. Inference code for LLaMA models
2on-device_recommendation_tflite. Swift
2zhuzilin.
2dafny-exercises. some formal verification exercises using dafny.
2torchrec_mapper. C++
2nvcc_daxpy_example. C++
1sglang. SGLang is a fast serving framework for large language models and vision language models.
1scattermoe. Triton-based implementation of Sparse Mixture of Experts.
1unilm. Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
1bfc. A really fast non-JIT brainfuck interpreter.
1triton. Development repository for the Triton language and compiler
1electron-fc. A electron based famicom(NES) emulator
1flash-attention. Fast and memory-efficient exact attention
1garbled_circuit. A python implementation of Yao's GC.
1torch_memory_saver. Allow torch tensor memory to be released and resumed later
1autodiff. pytorch like automatic differentiation library in numpy
1