This is your work, valued
GPU_Programming. Python
98Penny. Hand-Rolled GPU communications library
96FastSoftmax. Step by step implementation of a fast softmax kernel in CUDA
70vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
7CUDA_Matmul. Cuda
7PongAssembly. Assembly
2CPU_Matmul. C++
2Deep-Learning-In-Survival-Analysis. Jupyter Notebook
1SpaceInvaders-FromPlanetZig. Zig
1LLM_Entropy. Python
1Triton-Vs-CUDA-Profiling. Cuda
1Verilog_Playground. Verilog
1