This is your work, valued

Peking

Lei Wang

Elite
@LeiWang1999

I love everything and create value.

FPGA. 帮助大家进行FPGA的入门,分享FPGA相关的优秀文章,优秀项目

5.6k

ZYNQ-NVDLA. NVDLA (An Opensource DL Accelerator Framework) implementation on FPGA.

390

AICS-Course. 《智能计算系统 AI Computing Systems》习题答案、实验答案、课程笔记

216

tvm_gpu_gemm. play gemm with tvm

91

DigitalAlarmClock. njtech digital design. a fpga digital alarm system with Nexys A7 100T

54

AutoGPTQ.tvm. GPTQ inference TVM kernel

41

Stream-k.tvm. Python

20

VehicleFlowDetection. Implement of vehicle flow statistics based on tensorflow and yolo3 with pyqt5 GUI.

19

EthernetVideo. Use FPGA to Transfer Image with Gigabits Ethernet

19

Pynq-Accelerator. A easy general acc.

18

TVM.CMakeExtend. Tutorials of Extending and importing TVM with CMAKE Include dependency.

16

nvdla_loadables. some sample caffemodel, prototxt, test images and pre compiled loadabes .

14

nvdla-parser. A NVDLA Loadable Parser.

13

iCARDCrack. Tools and Tutorials for IC Card Crack.

10

T-MAG. GPU Implementation of T-MAC

7

leiblog.wang. My New Blog Powered by HEXO http://leiblog.wang

6

mlc-benchmark. Python

6

SchoolApi. zhengfang academic affairs system spider with egg.js

5

PrincelingModuleHub. VHDL

5

Ladder. Python

5

HPC-Course. Makefile

5

LeiBlog. 用Vuetify.js+Vue.js+Node.js(KOA 自己撸一个博客。http://leiblog.wang

5

CPPTorchExecutable. Python

5

BitBLAS. Python

5

rocblas-benchmark. Makefile

5

tvm. Open deep learning compiler stack for cpu, gpu and specialized accelerators

5

bitblas-benchmark. Python

5

EGO1_XADC. Example project use xadc with E-element ego1 board, read the resistance of the potentiometer.

4

vllm-bitblas. A high-throughput and memory-efficient inference and serving engine for LLMs

4

Roller. Build and Train AlexNet with PyTorch and Predict with TVM and Pytorch, compare the performance between them

4

deepwiki-open. Open Source DeepWiki: AI-Powered Wiki Generator for GitHub/Gitlab Repositories

3

compiler-and-arch. A list of tutorials, paper, talks, and open-source projects for emerging compiler and architecture

3

tilelang. Domain-specific language designed to streamline the development of high-performance GPU/CPU kernels

3

LearnBasys3. some so large project .

3

cv. resume.

3

memfusion_artifact. Python

3

kali. 🐍🐍🐍kali linux审计系统工具使用方法

2

MSBitBLAS. BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.

2

cutlass. C++

2

Tengine. Tengine is a lite, high performance, modular inference engine for embedded device

2

NTU-Machine-learning. 台湾大学李宏毅老师机器学习

2

AutoGPTQ_nf. Python

2

RISC-V-TensorCore. Transactional Verilog design and Verilator Testbench for a RISC-V TensorCore Vector co-processor for reproducible linear algebra

2

cutlass_fpA_intB_gemm. A standalone GEMM kernel for fp16 activation and quantized weight, extracted from FasterTransformer

2

flash-linear-attention. 🚀 Efficient implementations of state-of-the-art linear attention models in Pytorch and Triton

2

How-To-Ask-Questions-The-Smart-Way. How To Ask Questions The Smart Way 《提问的智慧》中文版

2

sw. NVDLA SW

2

optimize-gemm-on-macbook2019. 在MacBook2019上优化gemm.

2

FlashMLA. FlashMLA: Efficient MLA decoding kernels

1

stm32f407_hdb3. Njtech EE Project3.1, HDB3 Encoder and Decoder.

1

LeiWang1999.

1

autoencoder.pytorch. Pytorch implement of Auto-Encoder with MNIST dataset.

1