simplegemm. Cuda
134llama2.so. Inference Llama 2 with a model compiled to native code by TorchInductor
14tf32_gemm. Example of binding a TF32 CUTLASS GEMM kernel to PyTorch
12bitserial. Hacking around with ultra-low precision GEMM using TVM
3pytorch. Tensors and Dynamic neural networks in Python with strong GPU acceleration
3membench. C++
2midwit-matmul. A simplistic approach to high-performance GPU matmul
2learn_cuda. Simple programs for learning CUDA
1pyhpc-benchmarks. A suite of benchmarks for CPU and GPU performance of the most popular high-performance libraries for Python :rocket:
1deepstream-video-pipeline. Python
1