Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090
-
Updated
Aug 3, 2026 - C++
Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090
Epsilon is a library with functions for machine learning and statistics written in plain C. It is intended to run on microcontrollers.
Long-term solar activity forecast for solar cycles 25 and 26 with libraries numpy, pandas, scipy, sympy, sklearn. A science project by physics undergrad student.
Un experimento educativo para construir, entrenar y evaluar un Modelo de Lenguaje Pequeño/Grande experimental (~100M parámetros) completamente desde cero usando PyTorch.
A neural architecture framework exploring low-rank multiplicative gating, spectral orthogonal bases (DCT/Walsh), complex-valued phase mixers, and conformal geometry over frozen substrates. Learning to equalize, not to sculpt.
delphi unit to compute the fast Walsh-Hadamard transform
Add a description, image, and links to the fwht topic page so that developers can more easily learn about it.
To associate your repository with the fwht topic, visit your repo's landing page and select "manage topics."