Ph.D. student of Sun Yat-Sen University, prior intern @Tencent, @rednote-hilab and @MoonshotAI. Simulators, GPU, architecture, AI Infra, MLSys
gLLM. An Efficient and Versatile Inference Engine for Distributed LLM Serving
66GEMM_MMA. Optimize GEMM with tensorcore step by step
40PTX-EMU. PTX-EMU is a simple emulator for CUDA program.
40GEMM_WMMA. GEMM by WMMA (tensor core)
15SimpleUseGpgpuSim. GPGPU-SIM 使用篇
14EFIM. Python
3ConvNN. A simple CNN training framework support on CPU and GPU(CUDNN)
3BCI. Jupyter Notebook
1