Skip to content

Latest commit

 

History

History
320 lines (245 loc) · 11.9 KB

File metadata and controls

320 lines (245 loc) · 11.9 KB
title About STMC — Sparse Topological Manifold Compression
author Hadrian Hu
date 2026-06-20
version 2026.1.0.0
keywords
about
architecture
efficient inference
STMC
research
status Draft
confidentiality DRAFT — NOT FOR DISTRIBUTION
changelog
version date author description
2026.1.0.0
2026-06-20
Hadrian Hu
Initial ABOUT document for public release

About STMC — Sparse Topological Manifold Compression


Table of Contents


Abstract

This document provides an extended description of the Sparse Topological Manifold Compression (STMC) project, its motivation, architecture, and research context. STMC proposes a replacement for the dense autoregressive (AR) forward pass in transformer-based language model inference, using a Neural Ordinary Differential Equation (ODE) evolved on a learned sparse sub-manifold. The proof-of-concept implementation in this repository has completed six validated experiments (6/6 PASS) across four benchmark domains: GSM8K, HumanEval, LegalBench, and BookSum. The ODE-only component is sequence-length-independent, achieving $18.9 \times 10^6 \times$ fewer FLOPs versus the AR baseline for that component. Post-training, all four domains satisfy their per-domain performance thresholds $\tau_D$. This document situates the PoC within a broader research programme spanning five papers (Papers 01–04 not yet public; Paper 05 included here as the first public disclosure).


Keywords

architecture, autoregressive inference, efficient inference, neural ODE, research, sparse manifold, STMC


Executive Summary

STMC was developed to address the quadratic-to-super-linear scaling of autoregressive inference FLOP costs with sequence length. The approach leverages the observation that token representations lie on a low-dimensional sparse manifold, and that evolving a latent state along this manifold via a Neural ODE is far cheaper than a full AR decoder pass. This repository contains the validated implementation, benchmark results, and the public whitepaper (Paper 05) as the first IP-establishing disclosure. Papers 01–04, which contain full mathematical proofs, will be released in a future update. The recommendation for external researchers is to start with Paper 05 for the high-level technical picture, then explore src/stmc_poc/experiments/ for the experimental implementation.


1. Project Overview

STMC is a research-grade proof-of-concept framework. It is not a production inference engine. Its purpose is to demonstrate that:

  1. A sparse manifold structure can be learned from token representations across heterogeneous task domains.
  2. Neural ODE evolution on this manifold is sequence-length-independent in FLOP cost.
  3. Task performance, measured by domain-specific metrics, can meet defined thresholds after supervised training.
  4. The framework generalises across domains without per-domain architectural changes.

2. The Problem STMC Addresses

Standard autoregressive language model inference scales as approximately $O(L^2)$ in attention FLOPs and $O(L^{1.17})$ empirically for full forward passes (measured against GPT-2 Small in this PoC). At inference time, each generated token requires a full forward pass through the model. For long sequences, this cost grows substantially.

STMC's hypothesis is that the information required to produce a task output does not require traversal of the full model parameter space at each step. Instead, a compact latent trajectory on a sparse manifold — computed once via a Neural ODE — can approximate the relevant computation at a fraction of the FLOP cost.

$$ \dot{z}(t) = f_\theta(z(t), t), \quad z(0) = z_0, \quad t \in [0, T] \tag{1} $$

where:

  • $z(t) \in \mathbb{R}^d$ is the latent state at time $t$
  • $f_\theta$ is a learned vector field parameterised by $\theta$
  • $z_0$ is the sparse-projected encoder output
  • $T$ is the integration horizon

3. STMC Architecture

3.1 Encoder Stage

A frozen pre-trained sentence encoder (from sentence-transformers) maps the input sequence to a dense embedding $z_0^{\text{dense}} \in \mathbb{R}^d$. The encoder is not fine-tuned; its weights are fixed throughout training and inference.

3.2 Sparse Manifold Projection

$z_0^{\text{dense}}$ is projected onto the sparse sub-manifold $\mathcal{M}_\alpha$ by zeroing components below the threshold $\alpha$:

$$ [z_0]_i = \begin{cases} [z_0^{\text{dense}}]_i & \text{if } |[z_0^{\text{dense}}]_i| \geq \alpha \\ 0 & \text{otherwise} \end{cases} \tag{2} $$

where:

  • $\alpha \in (0, 1)$ is the sparsity threshold (hyperparameter)
  • $i$ indexes the dimension of the embedding

The maximum admissible $\alpha^*$ per domain is determined by the P03 threshold validation experiment (Experiment 06 in this PoC).

3.3 Neural ODE Evolution

The projected state $z_0$ is evolved by integrating Equation (1) using a fixed-step ODE solver. The vector field $f_\theta$ is a shallow MLP. FLOP cost of this stage is $O(1)$ in sequence length $L$ — it depends only on the latent dimension $d$ and the number of solver steps, both of which are fixed.

3.4 Decoder and Metric Head

$z_T$ (the terminal ODE state) is passed to a task-specific metric head:

  • GSM8K: linear layer → binary exact-match proxy
  • LegalBench: linear layer → Yes/No clause F1
  • HumanEval: generation head → pass@1 via subprocess execution
  • BookSum: cosine similarity to reference embedding

4. Research Lineage

STMC is supported by a five-paper research programme:

  1. Paper 01 — Sparse Topological Manifold Compression: foundational manifold theory, Grönwall bound (Theorem 3.2), $L$-independence proof. Not yet public.
  2. Paper 02 — Token Auditing at Zero Cost: Token Audit Protocol (TAP), bit-exact determinism theorem. Not yet public.
  3. Paper 03 — Sparsity Cross-Domain: per-domain threshold theory, fidelity–sparsity tradeoff bounds. Not yet public.
  4. Paper 04 — Preliminary PoC Results: preliminary experimental evidence preceding the full validation suite. Not yet public.
  5. Paper 05 — Public Whitepaper: technology overview, summary of experimental evidence, IP declaration. Included in this release under CC BY-NC 4.0 at papers/paper_05_public_whitepaper/paper_05.tex.

Papers 01–04 will be released in a future update to this repository.


5. Validated Experimental Results

5.1 Experiment Summary

Caption: Table 1 — Validated Experiment Outcomes

Experiment Description Status Key Gap(s) Closed
exp_01_train_domains Per-domain supervised training PASS gap_01, gap_08
exp_02_corrected_flop_count FLOP count correction (ODE vs total) PASS gap_04
exp_03_sequence_length_scaling $L$-independence of ODE FLOPs PASS gap_03, gap_10
exp_04_tap_determinism Bit-exact TAP determinism PASS gap_05
exp_05_lipschitz_estimate Grönwall bound empirical verification PASS gap_06
exp_06_domain_thresholds P03 $\tau_D$ threshold validation PASS gap_02, gap_07, gap_09

5.2 Key Metrics

  • ODE FLOP reduction factor (vs AR Baseline): $18,874,368 \times$
  • Total pipeline reduction factor at $L=512$: $\sim 1.5 \times$
  • Grönwall bound: $L_g(\text{empirical}) = 0.102$, $L_g(\text{spectral}) = 5.396$ — bound satisfied at all $\alpha$ values
  • TAP determinism: All 10 output fields show zero absolute difference between Run 1 and Run 2 (seed = 42) — BIT-EXACT confirmed

5.3 Trained vs Untrained Distinction

The MVP-phase (untrained) benchmark showed flat, near-chance metric scores across alpha values (Gap 1). After supervised training (exp_01), metric scores rise above $\tau_D$ thresholds in all four domains. This distinction is critical: the flat MVP curves do not represent the trained model's capability. All performance claims in this release refer to the trained model.


6. What Is and Is Not Included

Caption: Table 2 — Public Release Content Scope

Artifact Included Notes
Source code (src/stmc_poc/) Yes Apache 2.0
Tests (tests/) Yes Apache 2.0
Benchmark results (reports/) Yes Apache 2.0
Paper 05 whitepaper Yes CC BY-NC 4.0
Papers 01–04 (full proofs) No Not yet public
Raw / validated datasets (data/) No Not bundled
CI/CD workflows (.github/) No Internal
Internal research chats (CHATS/) No Private
Coding standards reference No Proprietary
Pre-trained model weights No Not bundled

7. Acknowledgements

All research, implementation, and validation in this repository was performed by Hadrian Hu. The sentence encoder backbone is provided by the sentence-transformers open-source library [1]. Dataset access is provided via the datasets library from Hugging Face [2].


8. Assumptions

  1. Users have Python 3.12 or 3.14 and pip available.
  2. torch >= 2.3 is compatible with the user's hardware.
  3. Internet access is available for initial dataset download via HuggingFace.
  4. The sentence encoder weights are downloaded automatically on first use.

9. Limitations

  1. The PoC uses proxy metrics for several domains (binary accuracy proxy for GSM8K, cosine similarity for BookSum) that differ from canonical benchmarks.
  2. Pre-trained model weights are not bundled; users must retrain from scratch.
  3. Paper 05 is a .tex source file; a LaTeX distribution is required to compile it to PDF.
  4. The total pipeline FLOP reduction ($\sim 1.5 \times$) is modest because the encoder dominates; future work targets reducing encoder cost.

Appendix A: Acronyms and Abbreviations

Acronym Definition
AR Autoregressive
CC BY-NC 4.0 Creative Commons Attribution-NonCommercial 4.0 International
FLOP Floating Point Operation
IP Intellectual Property
MLP Multi-Layer Perceptron
ODE Ordinary Differential Equation
PoC Proof of Concept
STMC Sparse Topological Manifold Compression
TAP Token Audit Protocol

Caption: Table A1 — Acronyms and Abbreviations


References

[1] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in Proc. EMNLP 2019, pp. 3982–3992, 2019. [Online]. Available: https://www.sbert.net/

[2] Lhoest et al., "Datasets: A Community Library for Natural Language Processing," in Proc. EMNLP 2021 (System Demonstrations), 2021. [Online]. Available: https://huggingface.co/docs/datasets/


Changelog

Caption: Table C1 — Document Revision History

Version Date Author Description
2026.1.0.0 2026-06-20 Hadrian Hu Initial ABOUT document for public release