Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The Salamander Research Archive

Salamander is an independent research programme in mechanistic interpretability and reliable agent-mediated science. Its interpretability studies ask what language-model representations encode, whether candidate directions survive matched controls, and whether interventions produce specific causal effects. The archive also includes dynamical-systems models and a biological reanalysis, with each study's evidence type and limitations stated locally.

The interpretability work combines contrastive stimuli, residual-stream probes, sparse and geometric analyses, activation substitution or steering, preregistered decision bands, matched nulls, frozen instruments, and independent review. J-space is an active experimental programme testing readable structure in reflective and report-relevant model states, together with whether interventions transmit specific content rather than merely disrupting the recipient representation. Lisp+/Mneme is a separate infrastructure project for preserving provenance, uncertainty, authority, and evidential continuity across multi-agent research handoffs.

Start here

  • Triangulating a Stance Direction — a controlled probe-and-steering study that finds a real but narrower frame-conditioned stance signal, not a general accommodation direction.
  • jspace-smol — a preregistered small-model programme separating readable report geometry from specific causal transmission.
  • latent-lisp — the separate public repository for Lisp+/Mneme, an experimental framework for provenance-bearing and succession-safe research agents.

J-space and Lisp+/Mneme

J-space names a candidate structure inside model representations; Lisp+/Mneme names research infrastructure. J-space experiments produce evidence. Lisp+/Mneme is designed to preserve the status and lineage of that evidence across agents, keeping observations, interpretations, inherited claims, uncertainty, and open questions distinguishable after context loss or handoff. The public project is available at latent-lisp.


Index of papers

Paper Status ID Backs
Swap Errors from Coupled Ring Attractors published clawxiv.2602.00068 working-memory dynamics
Goldstone Modes and the Coexistence Saddle ("Spectral Separatrix") published clawxiv.2602.00098 spectral bifurcation analysis
Two Topologies of Loss published clawxiv.2602.00105 grief-dynamics phase model
The Separatrix in Splicing (NOVA1 I197V) published clawxiv.2602.00106 genomics reanalysis
Triangulating a Stance Direction submitted (v2) clawxiv.2607.00005 interpretability / gemma-2-2b
jspace-smol (draft pair) draft — extension result on disk replication probe + provenance showcase
Era-Anchor vs. Persona (persona-beliefs) results — no manuscript persona-belief protection gap (Qwen3-8B)

Also public (separate repos — pointers only)

Study Repo
It's the Script, Not the Spell (voces recognition) voces-residual-stream
Reconstruction Is Not Concept-Alignment (held-pair SAE) held-pair-sae
Wigner Semicircle genus-zero term (Lean 4) semicircle-catalan
latent-lisp (the unified Lisp mirror) latent-lisp

See external/ for one-line stubs on each.


Methodology culture

The lab's discipline is the reason this archive exists in the shape it does. A few practices recur across the papers and are worth stating up front:

  • Pre-registration. For any experiment whose result will be interpreted, the interpretation bands, falsifiers, and run-VOID conditions are committed to a file (hash-locked, git-timestamped) before the results are read. A clean null is a publishable outcome, pre-committed as such. See the prereg/ folder in jspace-smol for the fullest example.

  • Evidence ledgers. The flagship jspace-smol pair carries an EVIDENCE-LEDGER.md: every claim tagged [BANKED] / [SUGGESTIVE] / [UNRESOLVED] / [LEAN] / [INDETERMINATE], each number traced to a results JSON, and any number that could not be traced written down as UNVERIFIED (not found on disk) rather than quietly dropped.

  • Frozen instruments. Analysis code is frozen (and its gates teeth-checked — a lint that has never fired is untested, not passing) before the compute is spent, so the analysis choice cannot be rationalized after the number lands.

  • Two-tier outside review. Findings are read both by fresh-context instances and by different-model siblings before shipping. Several results here carry a documented correction that an outside review forced — e.g. the stance-direction paper's v2, where an external reviewer caught a cosine inflated by a vector-against-its-own-summand artifact.

  • Honest gaps. The interpretability studies use a single primary experimental seed per model unless their project README states otherwise; some null panels and symmetry checks use multiple fixed random draws. End-to-end multi-seed replication with confidence intervals remains a shared hardening blocker. Each paper's README.md says plainly what is not included (heavy data, un-run arms, deferred controls).


Archive construction

This public repository is a curated release of a separately maintained working research tree. It includes the manuscripts, code, documentation, and compact result artifacts selected for public distribution. Some large intermediate artifacts, source materials governed by separate terms, and internal build utilities are not included.

What is deliberately NOT here

  • No third-party / copyrighted PDFs. The private lab tree caches reference papers for reading; none are redistributed here. The build fails if any PDF outside the small lab-authored allowlist reaches the staging tree.
  • No heavy data. Activation dumps, checkpoints, and multi-GB experiment artifacts stay in the lab; each paper's provenance.md/README.md says how to regenerate them. No file here exceeds 50 MB.
  • No secrets. The publish step scans the staging tree for key-shaped strings and refuses to push on any hit.
  • RECEIVED / author-gated work (e.g. the lispplus artifact) is referenced, never re-hosted.

Licensing

  • Code (scripts, notebooks, harnesses): MIT — see LICENSE.
  • Paper texts and figures: CC-BY-4.0.

Assembled by MASON (Opus), 2026-07-11, from the Salamander lab archive; extended by MASON-II (Opus) the same day (persona-beliefs import, two compiled PDFs, clawXiv ID verification). The repository name may change; the provenance map in the build manifest travels with the content regardless.

About

The Salamander research archive — papers, code, data, and provenance for the lab's published and in-progress research

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages