I love everything and create value.
FPGA. 帮助大家进行FPGA的入门,分享FPGA相关的优秀文章,优秀项目
5.6kZYNQ-NVDLA. NVDLA (An Opensource DL Accelerator Framework) implementation on FPGA.
390AICS-Course. 《智能计算系统 AI Computing Systems》习题答案、实验答案、课程笔记
216tvm_gpu_gemm. play gemm with tvm
91DigitalAlarmClock. njtech digital design. a fpga digital alarm system with Nexys A7 100T
54AutoGPTQ.tvm. GPTQ inference TVM kernel
41Stream-k.tvm. Python
20VehicleFlowDetection. Implement of vehicle flow statistics based on tensorflow and yolo3 with pyqt5 GUI.
19EthernetVideo. Use FPGA to Transfer Image with Gigabits Ethernet
19Pynq-Accelerator. A easy general acc.
18TVM.CMakeExtend. Tutorials of Extending and importing TVM with CMAKE Include dependency.
16nvdla_loadables. some sample caffemodel, prototxt, test images and pre compiled loadabes .
14nvdla-parser. A NVDLA Loadable Parser.
13iCARDCrack. Tools and Tutorials for IC Card Crack.
10T-MAG. GPU Implementation of T-MAC
7leiblog.wang. My New Blog Powered by HEXO http://leiblog.wang
6mlc-benchmark. Python
6SchoolApi. zhengfang academic affairs system spider with egg.js
5PrincelingModuleHub. VHDL
5Ladder. Python
5HPC-Course. Makefile
5LeiBlog. 用Vuetify.js+Vue.js+Node.js(KOA 自己撸一个博客。http://leiblog.wang
5CPPTorchExecutable. Python
5BitBLAS. Python
5rocblas-benchmark. Makefile
5tvm. Open deep learning compiler stack for cpu, gpu and specialized accelerators
5bitblas-benchmark. Python
5EGO1_XADC. Example project use xadc with E-element ego1 board, read the resistance of the potentiometer.
4vllm-bitblas. A high-throughput and memory-efficient inference and serving engine for LLMs
4Roller. Build and Train AlexNet with PyTorch and Predict with TVM and Pytorch, compare the performance between them
4deepwiki-open. Open Source DeepWiki: AI-Powered Wiki Generator for GitHub/Gitlab Repositories
3compiler-and-arch. A list of tutorials, paper, talks, and open-source projects for emerging compiler and architecture
3tilelang. Domain-specific language designed to streamline the development of high-performance GPU/CPU kernels
3LearnBasys3. some so large project .
3cv. resume.
3memfusion_artifact. Python
3kali. 🐍🐍🐍kali linux审计系统工具使用方法
2MSBitBLAS. BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.
2cutlass. C++
2Tengine. Tengine is a lite, high performance, modular inference engine for embedded device
2NTU-Machine-learning. 台湾大学李宏毅老师机器学习
2AutoGPTQ_nf. Python
2RISC-V-TensorCore. Transactional Verilog design and Verilator Testbench for a RISC-V TensorCore Vector co-processor for reproducible linear algebra
2cutlass_fpA_intB_gemm. A standalone GEMM kernel for fp16 activation and quantized weight, extracted from FasterTransformer
2flash-linear-attention. 🚀 Efficient implementations of state-of-the-art linear attention models in Pytorch and Triton
2How-To-Ask-Questions-The-Smart-Way. How To Ask Questions The Smart Way 《提问的智慧》中文版
2sw. NVDLA SW
2optimize-gemm-on-macbook2019. 在MacBook2019上优化gemm.
2FlashMLA. FlashMLA: Efficient MLA decoding kernels
1stm32f407_hdb3. Njtech EE Project3.1, HDB3 Encoder and Decoder.
1LeiWang1999.
1autoencoder.pytorch. Pytorch implement of Auto-Encoder with MNIST dataset.
1