This page documents the attention registry values tr_mha and tr_mha_v2.
Naming note: the main README uses TR-MHA for the complete
attention_type="mha"+mlp_type="tr_hash_engine"architecture. The components below are a separate experimental branch that routes low-rank residual adapters inside MHA.
The matched result for the complete MHA + TR-MoE architecture is reported in
RESULTS_100M_MPS.md. It must not be attributed to the
attention-adapter prototypes documented on this page.
Both adapter variants preserve standard full MHA:
shared Q/K/V projections ─► causal attention ─► output projection
│
└─ token-routed low-rank residual applied to Q and/or V
Every query head keeps its own K/V head, so
num_key_value_heads == num_attention_heads is required.
The first prototype:
- builds two layer-specific candidates from token ID;
- computes a contextual score over all route experts;
- combines fixed prior and contextual verification logits;
- selects top-k low-rank Q and V adapters;
- adds their weighted deltas to the shared MHA projections.
The routed branch is not guaranteed to be neutral at initialization.
The second prototype is baseline-preserving:
- the full MHA path is unchanged at initialization;
- the routed up-projection is initialized to zero;
- token ID fixes the two candidate experts;
- context only reweights those candidates;
- only selected expert/token pairs are evaluated;
tr_mha_targetschoosesq,v, orqv.
from complexity import ComplexityModel, ModelConfig
config = ModelConfig(
hidden_size=384,
num_hidden_layers=10,
num_attention_heads=8,
num_key_value_heads=8,
attention_type="tr_mha_v2",
mlp_type="tr_hash_engine",
intermediate_size=256,
shared_expert=True,
shared_intermediate_size=1536,
routing_strategy="token_id_balanced_hash",
vocab_size=32_000,
tr_mha_num_experts=4,
tr_mha_adapter_rank=4,
tr_mha_top_k=2,
tr_mha_targets="qv",
)
model = ComplexityModel(config)The executable construction and neutrality checks are maintained in
tests/test_tr_mha.py. Historical experiment YAMLs were removed with the old
pretraining driver and must not be presented as runnable configurations.
The training runner can report:
- adapter output gate;
- contextual verifier strength and use;
- route entropy;
- per-expert route shares.
These are bounded attention experiments. They are not the default pretraining architecture and should not inherit claims from MHA + TR-MoE pilots.