This is your work, valued
doing things at the edge of stability
flash-attention-residuals. Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)
LinearKAN. LinearKAN: A very fast implementation of Kolmogorov-Arnold Networks
kernels. a kernel a day keeps the doctor away