PeachPy. x86-64 assembler embedded in Python
2.1kNNPACK. Acceleration package for neural networks on multi-core CPUs
1.7kpthreadpool. Portable (POSIX/Windows/Emscripten) thread pool for C/C++
394FP16. Conversion to/from half-precision floating point formats
384Opcodes. Database of CPU Opcodes
265FXdiv. C99/C++ header-only library for division via fixed-point multiplication by inverse
61psimd. Portable 128-bit SIMD intrinsics
61caffe-nnpack. Caffe with NNPACK integration
59FPplus. Scientific library for high-precision computations and research
49clcc. OpenCL offline compiler
21confu. Ninja-based configuration system
11caffe. Caffe: a fast open framework for deep learning.
10laff-demos. Live demos for Linear Algebra - Foundations to Frontiers course on edX
8CSE6230. High Performance Computing: Tools and Applications course examples
4blis-bench. Benchmark of matrix-matrix multiplication implementations for Web browsers
3vision. Datasets, Transforms and Models specific to Computer Vision
2cpuinfo. CPU INFOrmation library (x86/x86-64/ARM/ARM64, Linux/Windows/Android/macOS/iOS)
2onnx. Open Neural Network Exchange
1stm32_bare_lib. System functions and example code for programming the "Blue Pill" STM32-compatible micro-controller boards.
1XNNPACK. High-efficiency floating-point neural network inference operators for mobile and Web
1caffe2. Caffe2 is a lightweight, modular, and scalable deep learning framework.
1asm.yeppp.info. Run GCC (and other compilers) interactively from your web browser and experiment with its generated code
1simd. Branch of the spec repo scoped to discussion of SIMD in WebAssembly
1onnx-tensorrt. ONNX-TensorRT: TensorRT backend for ONNX
1