Turn a Hugging Face model into a working demo, through a conversation. Built with aparté.
-
Updated
Aug 30, 2026 - TypeScript
Turn a Hugging Face model into a working demo, through a conversation. Built with aparté.
In-Browser WebGPU, ONNX Runtime Web & CoreML Inference Latency Matrix
A 19.5M parameter next word prediction transformer trained from scratch RMSNorm, RoPE, SwiGLU, grouped query attention quantised to int8 and running entirely in your browser via WebAssembly. Perplexity 28.3, top 5 accuracy 60.7%, 16ms per prediction, zero inference servers. Ships with the writing app, user dashboard, and admin console.
To associate your repository with the in-browser-inference topic, visit your repo's landing page and select "manage topics."