NV/AMD GPU Kernel Performance Analysis & Optimization. Now in Bytedance(Data => AML => Seed). Used to be in AIACC Team@Alibaba Cloud.
hp_rms_norm. High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)
30expert_specialization_moe. Expert Specialization MoE Solution based on CUTLASS
27cute_reduce. Reduce kernel based on CUTLASS CuTe and TMA.
12cutlass_cute_experiments. Cuda
10cuda_prefetch_experiment. Prefetch experiments codes.
9CUDASynchronizePrimitives. Benchmark sync primitives
5cutedsl_experiment. Some code snippet for CuTeDSL
5swizzled_layout_gemm. Using a swizzled hierarchical layout for GEMM
5GPU-Cache-Operator. Cuda
3CUDA_Practice. This is a repository with some CUDA examples.
3tma_multicast_demo. Simple Code Snippet for TMA Multicast.
3MPI_practice. A repository with some MPI programming examples.
2OpenMP_practice. This is a repository with some OpenMP examples.
2sglang. SGLang is a fast serving framework for large language models and vision language models.
2NVGPUMicroBenchmark. Instruction-level benchmarks for NVGPUs
2peer_access_demo. Cache operator demo code for peer access
1POSIX_pthread_practice. This is a repository with some POSIX pthread interface.
1