Research and Materials on Hardware implementation of Transformer Model
-
Updated
Feb 28, 2025 - Jupyter Notebook
Research and Materials on Hardware implementation of Transformer Model
AutoSA: Polyhedral-Based Systolic Array Compiler
Tiny ASIC implementation for "The Era of 1-bit LLMs All Large Language Models are in 1.58 Bits" matrix multiplication unit
A systolic array simulator for multi-cycle MACs and varying-byte words, with the paper accepted to HPCA 2022.
IHP 130nm ASIC tapeout of a 2x2 bfloat16 matrix matrix multiplication with DFT infrastructure. Iteration on the previous accelerator taped out on GF180.
A general framework for optimizing DNN dataflow on systolic array
Systolic-array based Deep Learning Accelerator generator
This work implements a dynamic programming algorithm for performing local sequence alignment. Through parallelism, it can run 136X times faster than a software running the same algorithm.
Working 8x8 systolic array hardware implemented in Xilinx Vivado, operated and controlled in software using Xilinx Vitis
Systolic Three Matrix Multiplier for Graph Convolutional Networks using High Level Synthesis
Design a Low-cost-AI-Accelerator based on Google's Tensor Processing Unit Version 1 in TSMC 16nm
Template for project1 TPU
Silicon-proven INT8 systolic NPU (8×8 MAC array) taped out on SkyWater 130nm via LibreLane. Features a custom 32-bit ISA, UART–APB host interface, and fused streaming datapath. Validated on chest X-ray pneumonia detection. Silicon Sprint 2026 — AUC.
This project is focused on the design and verification of digital logic circuits, particularly targeting chip design using Verilog, SystemVerilog, and SVA. The main objectives included designing modules compliant with industry standards such as APB (Advanced Peripheral Bus), memory systems, and systolic matrix multiplication.
A TPU you can watch run - real SystemVerilog systolic array, compiled to WASM, visualized live in your browser.
High-performance systolic-array accelerator for FP32 matrix multiplication in deep learning.
Garuda: CVXIF coprocessor optimizing batch-1 attention microkernels with 7.5-9× lower p99 latency. RISC-V INT8 MAC accelerator for transformer inference.
EE599 Accelerated Computing on FPGA
This is my senior project. Aims to implement the AI accelerator self-test and self-recovery architecture proposed in the paper "STRAIT: Self-Test and Self-Recovery for AI Accelerator". STRAIT is a unified solution that provides self-test, self-diagnosis, and self-recovery functions for systolic array-based AI accelerators.
To associate your repository with the systolic-arrays topic, visit your repo's landing page and select "manage topics."