Records of experiments testing the reproducibility of LLM workloads on datacenter GPUs. Key novelty: Non-associativity is a "fingerprint" of an inference stack and implementation.
-
Updated
Apr 19, 2026 - Jupyter Notebook
Records of experiments testing the reproducibility of LLM workloads on datacenter GPUs. Key novelty: Non-associativity is a "fingerprint" of an inference stack and implementation.
I-GCN graph convolutional network accelerator on Xilinx Zynq UltraScale+ ZCU9EG. Vitis HLS 2025.1, INT8 mixed-precision, 4-PE parallelism, up to 53% pruning and 35% latency reduction over baseline GCN inference.
UVM verification of an 8×8 INT8 weight-stationary systolic-array ML accelerator — closed to 100% functional coverage on an open-source Verilator + UVM toolchain.
To associate your repository with the ml-accelerator topic, visit your repo's landing page and select "manage topics."