NVIDIA Distinguished Engineer. CUDA, C++, GPU Computing, RAPIDS, Data Analytics, HPC.
hemi. Simple utilities to enable code reuse and portability between CUDA C/C++ and standard C/C++.
349numba_examples. Examples using Numba.
140mini-nbody. A simple gravitational N-body simulation in less than 100 lines of C code, with CUDA optimizations.
109sublimetext-cuda-cpp. CUDA C++ package for Sublime Text 2 & 3
68nsys_easy. Easier, quicker command-line CUDA profiling
63ranger. Generate simple index ranges in C++ and CUDA C++
39cpp11-range. Range-based for loops to iterate over a range of numbers or values
34cuda_event_benchmark. Unit benchmarks of CUDA event APIs.
17culeidoscope. Parallel extension of the Kaleidoscope toy language (from the LLVM project) on the CUDA platform.
11stream-safety-first. Examples from Mark Harris's 2023 GTC presentation on Stream Safety with stream-ordered allocation.
4numba. NumPy aware dynamic Python compiler using LLVM
3raft. Rapids Analytics Framework Toolset to share building blocks between cuGraph and cuML
1flaky-extensions-installer. Batch install vscode extensions with retry for flaky network connections
1cuspatial. CUDA-accelerated GIS and spatiotemporal algorithms
1cuml. cuML - RAPIDS Machine Learning Library
1new-ubuntu-setup. My personal setup for new Ubuntu instances
1rmm. RAPIDS Memory Manager
1forage_maps. Source code for foragemaps.com website
1tHogbomCleanHemi. Portable CUDA / OpenMP implementation of the Hogbom Clean Benchmark
1libgdf. C GPU Dataframe Library
1cudf. Python GPU DataFrame Library
1