This is your work, valued
CUDA_gemm. A simple high performance CUDA GEMM implementation.
★ 437Allocator_MemoryPool. MemoryPool based c++ STL Allocator
★ 12KgeN. A TVM-like CUDA/C code generator.
★ 9CUDA_resources.
★ 2cse-231-final-project. TypeScript
★ 2Pyflow. A toy Pytorch implementation.
★ 2enigne-hcraes. enigne hcraes
★ 1data_mining_homework. data_mining_homework
★ 1SPLCompiler. Pascal subset compiler written in c++
★ 1ray_tracing. ray tracer implemented in C++, inspired by peter shirley's book about ray tracing.
★ 1lure. compiler written in Go
★ 1minisql. implement a simple SQL engine from scratch
★ 1cutile-python. cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
★ 2.1kJetStream. JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
★ 451pprof. pprof is a tool for visualization and analysis of profiling data
★ 9.3kTriton-Puzzles. Puzzles for learning Triton
★ 2.5kJAX-Toolbox. JAX-Toolbox
★ 425aqt. Python
★ 358maxtext. A simple, performant, and scalable Jax LLM!
★ 2.4ktriton. Development repository for the Triton language and compiler
★ 20kthreefry. C++
★ 9paxml. Pax is a Jax-based machine learning framework for training large scale models. Pax allows for advanced and fully configurable experimentation and parallelization, and has demonstrated industry leading model flop utilization rates.
★ 556mlc-llm. Universal LLM Deployment Engine with ML Compilation
★ 23kminivm. A VM That is Dynamic and Fast
★ 1.7kneedle. A machine learning framework project motivated by CMU-10414
★ 1Halide. a language for fast, portable data-parallel computation
★ 1pyxis. Container plugin for Slurm Workload Manager
★ 456enroot. A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.
★ 984Summer2027-Internships. Summer 2026 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.
★ 46kostep-translations. Various translations of OSTEP can be found here. Help the cause and contribute!
★ 3.1kmit-6.824-labs. MIT 6.824 (Distributed Systems) labs in Go
★ 238PySyncObj. A library for replicating your python class between multiple servers, based on raft protocol
★ 750HIPIFY. HIPIFY: Convert CUDA to Portable C++ Code
★ 716gpumembench. A GPU benchmark suite for assessing on-chip GPU memory bandwidth
★ 113ai-edu. AI education materials for Chinese students, teachers and IT professionals.
★ 14kportion. portion, a Python library providing data structure and operations for intervals.
★ 523awesome-machine-learning-in-compilers. Must read research papers and links to tools and datasets that are related to using machine learning for compilers and systems optimisation
★ 1.7klibfuse. The reference implementation of the Linux FUSE (Filesystem in Userspace) interface
★ 6.1kbert. TensorFlow code and pre-trained models for BERT
★ 40kaccel-sim-framework. This is the top-level repository for the Accel-Sim framework.
★ 631tiramisu. A polyhedral compiler for expressing fast and portable data parallel algorithms
★ 960pytorch-cifar100. Practice on cifar100(ResNet, DenseNet, VGG, GoogleNet, InceptionV3, InceptionV4, Inception-ResNetv2, Xception, Resnet In Resnet, ResNext,ShuffleNet, ShuffleNetv2, MobileNet, MobileNetv2, SqueezeNet, NasNet, Residual Attention Network, SENet, WideResNet)
★ 4.8kcontinuation. JavaScript asynchronous Continuation-Passing Style transformation (deprecated).
★ 386pytorch-cifar. 95.47% on CIFAR10 with PyTorch
★ 6.4kdatum. A easy maintain(read/write) language for transform from/to other languages. 下一代企业级编程语言。
★ 139tvm-in-action. TVM stack: exploring the incredible explosion of deep-learning frameworks and how to bring them together
★ 65iree. A retargetable MLIR-based machine learning compiler and runtime toolkit.
★ 3.9kTASO. The Tensor Algebra SuperOptimizer for Deep Learning
★ 743SGEMM-Implementation-and-Optimization. :pencil: Some source code about matrix multiplication implementation on CUDA
★ 34autograd. Efficiently computes derivatives of NumPy code.
★ 7.5kjax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36kcppcoro. A library of C++ coroutine abstractions for the coroutines TS
★ 3.9kfairseq. Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
★ 32klearn-tt. A collection of resources for learning type theory and type theory adjacent fields.
★ 2.5kmcr2. Official Implementation of Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate Reduction (2020)
★ 205xbyak. A JIT assembler for x86/x64 architectures supporting FPU, MMX, SSE (1-4), AVX (1-2, 512), APX, and AVX10.2
★ 2.3kAIChip_Paper_List.
★ 672antares. Antares: an automatic engine for multi-platform kernel generation and optimization. Supporting CPU, CUDA, ROCm, DirectX12, GraphCore, SYCL for CPU/GPU, OpenCL for AMD/NVIDIA, Android CPU/GPU backends.
★ 465cpplinks. A categorized list of C++ resources.
★ 5.3krefl-cpp. Static reflection for C++17 (compile-time enumeration, attributes, proxies, overloads, template functions, metaprogramming).
★ 1.2kAwesome-Model-Quantization. A list of papers, docs, codes about model quantization. This repo is aimed to provide the info for model quantization research, we are continuously improving the project. Welcome to PR the works (papers, repositories) that are missed by the repo.
★ 2.4kgraffitist. Graph Transforms to Quantize and Retrain Deep Neural Nets in TensorFlow
★ 170BPPSA-open. The (open-source part of) code to reproduce "BPPSA: Scaling Back-propagation by Parallel Scan Algorithm".
★ 13nnfusion. A flexible and efficient deep neural network (DNN) compiler that generates high-performance executable from a DNN model description.
★ 1klibuv-tutorial. http://nikhilm.github.io/uvbook/
★ 42Catch2. A modern, C++-native, test framework for unit-tests, TDD and BDD - using C++14, C++17 and later (C++11 support is in v2.x branch, and C++03 on the Catch1.x branch)
★ 21klibnop. libnop: C++ Native Object Protocols
★ 581cpp_exception_handling_abi. A mini ABI capable of handling throw/catch statements for C++ without libstdc++
★ 173ZoneMod. A Competitive L4D2 Configuration.
★ 77learnGitBranching. An interactive git visualization and tutorial. Aspiring students of git can use this app to educate and challenge themselves towards mastery of git!
★ 34kllvm-tutor. A collection of out-of-tree LLVM passes for teaching and learning
★ 3.4klinux. Linux kernel source tree
★ 260gllvm. Whole Program LLVM: wllvm ported to go
★ 342uthash. C macros for hash tables and more
★ 4.7kjittor. Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.
★ 3.2kdevguide. The Python developer's guide
★ 2.1kkenali-kernel. Modified Nexus 9 kernel for Kenali Project
★ 30kernel-analyzer. C++
★ 73llvm-pass-tutorial. A step-by-step tutorial for building an LLVM sample pass
★ 220sort. Repository of sort algorithms in C and CUDA
★ 34pluto. Pluto: An automatic polyhedral parallelizer and locality optimizer
★ 335taichi. Productive, portable, and performant GPU programming in Python.
★ 28kcutlass. CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10kvexcl. VexCL is a C++ vector expression template library for OpenCL/CUDA/OpenMP
★ 721WAVM. WebAssembly Virtual Machine
★ 2.8kmaxas. Assembler for NVIDIA Maxwell architecture
★ 1.1kBinarized-Neural-networks-using-pytorch. Pytorch Implementation using Binary Weighs and activation.Accuracies are comparable .
★ 45wenyan. 文言文編程語言 A programming language for the ancient Chinese.
★ 20kcusplibrary. CUSP : A C++ Templated Sparse Matrix Library
★ 424BinaryNet.pytorch. Binarized Neural Network (BNN) for pytorch
★ 532Lantern. Cuda
★ 172HalideIR. Symbolic Expression and Statement Module for new DSLs
★ 207tangent. Source-to-Source Debuggable Derivatives in Pure Python
★ 2.3ktaco. The Tensor Algebra Compiler (taco) computes sparse tensor expressions on CPUs and GPUs
★ 1.4kHalide. a language for fast, portable data-parallel computation
★ 6.6kpytorch-slimming. Learning Efficient Convolutional Networks through Network Slimming, In ICCV 2017.
★ 573tvm. Open Machine Learning Compiler Framework
★ 14kserenity. The Serenity Operating System 🐞
★ 34kextension-cpp. C++ extensions in PyTorch
★ 1.2kblocksparse. Efficient GPU kernels for block-sparse matrix multiplication and convolution
★ 1.1kgit-recipes. 🥡 Git recipes in Chinese by Zhongyi Tong. 高质量的Git中文教程.
★ 15kulmBLAS. ulmBLAS
★ 111how-to-optimize-gemm. C
★ 2kSparseConvNet. Submanifold sparse convolutional networks
★ 2.1kpybind11. Seamless operability between C++11 and Python
★ 18kspconv. Spatial Sparse Convolution Library
★ 2.3kMIPP. Portable wrapper for SIMD and vector instructions written in C++11. Compatible with NEON, SSE, AVX, AVX-512 and SVE (length specific).
★ 529BinaryNet. Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
★ 1.1kbinarynet. Python
★ 19pytorch_sparse. PyTorch Extension Library of Optimized Autograd Sparse Matrix Operations
★ 1.1ktinyflow. Tutorial code on how to build your own Deep Learning System in 2k Lines
★ 122assignment1-2018. Assignment 1: automatic differentiation
★ 474