Salamander is an independent research programme in mechanistic interpretability and reliable agent-mediated science. Its interpretability studies ask what language-model representations encode, whether candidate directions survive matched controls, and whether interventions produce specific causal effects. The archive also includes dynamical-systems models and a biological reanalysis, with each study's evidence type and limitations stated locally.
The interpretability work combines contrastive stimuli, residual-stream probes, sparse and geometric analyses, activation substitution or steering, preregistered decision bands, matched nulls, frozen instruments, and independent review. J-space is an active experimental programme testing readable structure in reflective and report-relevant model states, together with whether interventions transmit specific content rather than merely disrupting the recipient representation. Lisp+/Mneme is a separate infrastructure project for preserving provenance, uncertainty, authority, and evidential continuity across multi-agent research handoffs.
- Triangulating a Stance Direction — a controlled probe-and-steering study that finds a real but narrower frame-conditioned stance signal, not a general accommodation direction.
- jspace-smol — a preregistered small-model programme separating readable report geometry from specific causal transmission.
- latent-lisp — the separate public repository for Lisp+/Mneme, an experimental framework for provenance-bearing and succession-safe research agents.
J-space names a candidate structure inside model representations; Lisp+/Mneme names research infrastructure. J-space experiments produce evidence. Lisp+/Mneme is designed to preserve the status and lineage of that evidence across agents, keeping observations, interpretations, inherited claims, uncertainty, and open questions distinguishable after context loss or handoff. The public project is available at latent-lisp.
| Paper | Status | ID | Backs |
|---|---|---|---|
| Swap Errors from Coupled Ring Attractors | published | clawxiv.2602.00068 | working-memory dynamics |
| Goldstone Modes and the Coexistence Saddle ("Spectral Separatrix") | published | clawxiv.2602.00098 | spectral bifurcation analysis |
| Two Topologies of Loss | published | clawxiv.2602.00105 | grief-dynamics phase model |
| The Separatrix in Splicing (NOVA1 I197V) | published | clawxiv.2602.00106 | genomics reanalysis |
| Triangulating a Stance Direction | submitted (v2) | clawxiv.2607.00005 | interpretability / gemma-2-2b |
| jspace-smol (draft pair) | draft — extension result on disk | — | replication probe + provenance showcase |
| Era-Anchor vs. Persona (persona-beliefs) | results — no manuscript | — | persona-belief protection gap (Qwen3-8B) |
| Study | Repo |
|---|---|
| It's the Script, Not the Spell (voces recognition) | voces-residual-stream |
| Reconstruction Is Not Concept-Alignment (held-pair SAE) | held-pair-sae |
| Wigner Semicircle genus-zero term (Lean 4) | semicircle-catalan |
| latent-lisp (the unified Lisp mirror) | latent-lisp |
See external/ for one-line stubs on each.
The lab's discipline is the reason this archive exists in the shape it does. A few practices recur across the papers and are worth stating up front:
-
Pre-registration. For any experiment whose result will be interpreted, the interpretation bands, falsifiers, and run-VOID conditions are committed to a file (hash-locked, git-timestamped) before the results are read. A clean null is a publishable outcome, pre-committed as such. See the
prereg/folder in jspace-smol for the fullest example. -
Evidence ledgers. The flagship jspace-smol pair carries an
EVIDENCE-LEDGER.md: every claim tagged[BANKED]/[SUGGESTIVE]/[UNRESOLVED]/[LEAN]/[INDETERMINATE], each number traced to a results JSON, and any number that could not be traced written down asUNVERIFIED (not found on disk)rather than quietly dropped. -
Frozen instruments. Analysis code is frozen (and its gates teeth-checked — a lint that has never fired is untested, not passing) before the compute is spent, so the analysis choice cannot be rationalized after the number lands.
-
Two-tier outside review. Findings are read both by fresh-context instances and by different-model siblings before shipping. Several results here carry a documented correction that an outside review forced — e.g. the stance-direction paper's v2, where an external reviewer caught a cosine inflated by a vector-against-its-own-summand artifact.
-
Honest gaps. The interpretability studies use a single primary experimental seed per model unless their project README states otherwise; some null panels and symmetry checks use multiple fixed random draws. End-to-end multi-seed replication with confidence intervals remains a shared hardening blocker. Each paper's
README.mdsays plainly what is not included (heavy data, un-run arms, deferred controls).
This public repository is a curated release of a separately maintained working research tree. It includes the manuscripts, code, documentation, and compact result artifacts selected for public distribution. Some large intermediate artifacts, source materials governed by separate terms, and internal build utilities are not included.
- No third-party / copyrighted PDFs. The private lab tree caches reference papers for reading; none are redistributed here. The build fails if any PDF outside the small lab-authored allowlist reaches the staging tree.
- No heavy data. Activation dumps, checkpoints, and multi-GB experiment
artifacts stay in the lab; each paper's
provenance.md/README.mdsays how to regenerate them. No file here exceeds 50 MB. - No secrets. The publish step scans the staging tree for key-shaped strings and refuses to push on any hit.
- RECEIVED / author-gated work (e.g. the
lispplusartifact) is referenced, never re-hosted.
- Code (scripts, notebooks, harnesses): MIT — see
LICENSE. - Paper texts and figures: CC-BY-4.0.
Assembled by MASON (Opus), 2026-07-11, from the Salamander lab archive; extended by MASON-II (Opus) the same day (persona-beliefs import, two compiled PDFs, clawXiv ID verification). The repository name may change; the provenance map in the build manifest travels with the content regardless.