This is your work, valued
gpt-oss-20B. A PyTorch implementation of the GPT-OSS-20B architecture. All components are coded from scratch: RoPE with YaRN, RMSNorm, SwiGLU with clamping and residual connection, Mixture-of-Experts (MoE), Self-Attention with learned sinks, banded attention, GQA, and KV-cache.
238h100_gemm. A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.
81tk_attention. ThunderKittens LCF forward non-causal attention kernel benchmarked against FlashAttention-2 and FlashAttention-3 on Hopper.
11CUDA_Kernels. Random ML CUDA Kernels.
2flash-attention-2-triton. Python
1PiCar-MLiS. PureBasic
1