This is your work, valued
ML performance engineer currently focusing on LLM and proteinLM pretraining efficiency and distributed training infra. Open to positions in industry
nanogpt-fp8. Nanochat inspired LLM pretraining using Transformer-Engine with MXFP8 and NVFP4 support. Up to 30% faster than nanochat
6matmul_assembly_x86. Hyper-optimized FP32 GEMM kernels in handwritten AVX2 ASM with a worklog of optimizations implemented
5llm.c. 3x faster LLM training on CPU than Karpathy's original repo
2flash-mHC. Fast implementation of manifold-constrained hyperconnections in Triton.
2cmri-yolosam. Python
1particle-simulator. C
1