Currently a MLSys engineer @ bytedance. Previously researched and worked at CCEM, University of Illinois
INT8-Flash-Attention-FMHA-Quantization. Cuda
165CUDA-INT8-GEMM. CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA API
37eigenMHA. Forward and backward Attention DNN operators implementationed by LibTorch, cuDNN, and Eigen.
31Finite-Element-Domain-Decomposition. Overlapping Schwarz Domain Decomposition Finite Element Algorithm in both Matlab and serial/parallel C++
18Time-Domain-CEM. The only known (by 2022) open-source, easy-to-understand basic algorithm implementations in TD-CEM. (Please star and fork this project if you find it useful!)
15PDE-Net-FDTD. Python
13DNN-2d-FDTD. Python
12RNN-1d-FDTD. Python
9GPU-Tensor-Permute. permute sequence data on GPU with high bandwidth
8dnn-test-framework. DNN unit test framework
7regex-gpu. Block-optimized regular expression matching engine on GPU
6cutlass-kernel-volta-gemm. volta fp16 gemm kernel
4Flash-LightSeq. C++
3eigenDNN.
2adaptive-filtering-algorithms. Adaptive Algorithms
2FETD-MFEM. A simple Finite element time domain example built with MFEM
2DNN-Discrete-Hilbert-Transform. Python
1cutlass-b2bgemm. an extension to the cutlass half-precision b2b gemm example
1Heterogeneous-GPUs. Heterogeneous Nvidia (CUDA) and Intel (OpenCL) GPU Programming
1method-of-moments. method of moments (MoM) for conducting wire problem
1FelixFu520-README. A pupil in the computer world.(Felix Fu)
1FlyDSL. Python
1