cutlass_fpA_intB_gemm. A standalone GEMM kernel for fp16 activation and quantized weight, extracted from FasterTransformer

github.com/junrushao/cutlass_fpA_intB_gemm

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.