Skip to content

Commit 204c020

Browse files
committed
docs: CHANGELOG 2026-08-13 - classifier prompt fix corrects 08-12 verdict, muse-glimmer promoted
Archived the oldest active entry (2026-07-22 Tenant threading, M2 Wave B1/B2) to CHANGELOG-2026.md to hold the 10-entry cap.
1 parent 0e004a4 commit 204c020

2 files changed

Lines changed: 21 additions & 9 deletions

File tree

‎logs/CHANGELOG/CHANGELOG-2026.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,13 @@
1+
## 2026-07-22 - Tenant threading landed for SessionStore + FactStore (M2 Wave B1/B2)
2+
3+
**Changes:** Threaded non-optional `tenantId: string` as the first parameter through every `SessionStore` port method (both `sqlite-session-store.ts`/`pg-session-store.ts`, plus the concrete-only `recentWrites`/`recentMarkers`/`getSessionScopeById`/`insertSession`/`listBackfillCandidates`/`insertFactsForSession`) and every `FactStore` port method (both backends, incl. `ingestSessionFactsInTxn`). Every SELECT/UPDATE/DELETE on the `sessions`/`facts`/`entities`/`entity_variants`/`session_entities` STAMP tables now routes its WHERE fragment through `tenantClause`/`tenantClausePg` — no inline `tenant_id = ?`/`$N` left in any store file outside `tenant-clause.ts` (verified: the leak-contract's case-11 pre-threading-floor scan is green). Vector-path rule applied to both `semanticSearch` implementations (session chunk KNN and fact KNN both re-resolve candidate ids against the tenant-filtered base table before returning — a same-epsilon cross-tenant embedding no longer surfaces). By-id rule applied: `getById`/`getByIds` for a wrong-tenant id return the same not-found shape as a missing row. Three defense-in-depth `X.tenant_id = s.tenant_id` join conditions were removed as provably redundant (session_entities/facts rows are always stamped with their owning session's tenant at write time) rather than routed through the helper, to keep the guard's literal-scan invariant honest without duplicating protection SQL already provides. Compiler-driven caller threading reached ~25 call sites across `RecallService`/`FactRecallService`/ingest/reprocess/reclassify-oversized/backfill-facts/scheduler/workstream bind+rollup/work-digest/`stampSignalScope`/HTTP routes/MCP handlers/CLI commands — `DEFAULT_TEAM_ID` supplied at each composition root (route handler entry, MCP tool entry, CLI command action, scheduler construction), never hardcoded mid-chain. Leak-contract cases 1, 2, and 4 flipped from `it.todo` to real assertions against the fixture's real (now-threaded) stores; the fixture itself rewritten to seed sessions/facts through `SqliteSessionStore.insertSession`/`SqliteFactStore.insert` instead of raw SQL (code_exemplars/signals/workstreams stay raw — unthreaded, Wave B3-B6). Suite 2177 pass / 0 fail (sqlite lane); pg lane source-compiles but unexecuted (`NLM_PG_TEST_URL` not set in this environment — task #408 still open); `tsc --noEmit` and `build:server` both clean.
4+
5+
**Decisions:** Store test-only helpers (`insertSessionForTest`, `insertEdgeForTest`) keep a `tenantId = DEFAULT_TEAM_ID` default parameter — the one deliberate exception to "no defaults in store signatures," scoped to test seams outside the port contract. `run-eval.ts`'s `Searcher` interface stayed single-arg; callers wrap `RecallService`/`FactRecallService` in a one-line adapter at each of the 3 call sites rather than widening the eval harness's own type.
6+
7+
**State:** main, uncommitted at CHANGELOG-write time (commits land right after, one per Wave-B task per the plan). Zero behavior change in local mode — `DEFAULT_TEAM_ID` is the only value passed anywhere until M3 replaces it with the token-resolved tenant.
8+
9+
**Next:** Wave B3-B6 (CodeExemplarStore, SignalStore/WorkstreamStore/EntityStore/OutcomeStore, Source/Provider registries, `actions-log.ts`/`build-dataset.ts`/data-stats raw-SQL modules), then Wave C (MCP/HTTP surface threading + hosted-mode gates + the full guard test with an allowlist + contract completion). #408 (pg test instance) blocks trusting the pg lane on this wave's changes.
10+
111
## 2026-07-22 - Audit remediation, #352 telemetry wave, v0.21.0 release, causal replay eval PASS, Team-NLM direction set
212

313
**Changes:** Full read-evidence audit (code review, repo/npm/runtime health, telemetry) found and fixed the root cause of a silent 13-day production downgrade (Jul 8 NxtOS install), a dead semantic-embedding lane, an npm CI break (npm@latest skipping install scripts), and a pg serial-test flake. Shipped #352 Phase 2 + Tier-B: derivable session columns (persona/parent/model/tokens/skill), a pure-core outcome rollup with a citation-degraded rule, the `report_outcome` MCP contract tool, and digest/get_session surfacing - backfilled live across 7,405 sessions. Elegance revision on the drift guard: folded into the existing daily digest instead of a second cron; daemon crash-restart moved to launchd KeepAlive. Released and deployed v0.21.0. Ran a pre-registered causal replay eval (100 real prompts, blind cross-family judge, bar locked before the run) - PASS on every gate (decisive 0.840, win 0.881, harm 0.119); published on the README as the project's first causal impact claim, with the harness shipped so anyone can run it on their own corpus. Closed a backlog cluster: re-derivation-pair persistence (#405), `NLM_ALERT_WEBHOOK` self-reporting (#404), durable gold-set builder (#403 tooling), better-sqlite3 13.0.1 (#402, drops the deprecated prebuild-install chain), an NxtOS installer version-downgrade guard, and action-layer polish + nvm warning (#294/#151). `@huggingface/transformers` bump (#400) found genuinely blocked upstream (no version of onnxruntime-node clears the adm-zip CVE yet).

‎logs/CHANGELOG/CHANGELOG.md‎

Lines changed: 11 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,16 @@
22

33
Session-level log per session protocol. Cap: 10 entries — archive older to `CHANGELOG-YYYY.md` when exceeded.
44

5+
## 2026-08-13 - Classifier prompt fix corrects the 08-12 muse-glimmer verdict; model promoted; runtime swap unblocked in the UI
6+
7+
**Changes:** (1) Root-caused the 08-12 eval's "Gemma stays" call: 4 of the 20 gold fixtures are trivial single-turn lookups (version check, `package.json` dump) where the reference expects `entities: []`, but both classifiers were extracting incidental proper nouns from command output regardless — muse-glimmer paid for it worse (its granular extraction style also surfaced file paths and hostnames the reference never listed), which is what produced 08-12's entity-F1 69.3% vs 42.6% gap. Fixed `CLASSIFIER_SYSTEM_PROMPT` (`prompt.ts`): hoisted a single "Session triage: WORK or LOOKUP" instruction above the field list (previously the trivial-session rule was only half-stated, separately, inside `confidence`'s description, with no shared definition across `entities`/`decisions`/`facts`), and added an entity-stability rule (file paths, hostnames, URLs, SHAs are never entities — redirect into `facts[]` under the matching predicate instead of dropping them). (2) Re-ran both models on the fixed prompt, n=20 each: Gemma entity-F1 69.4%->80.7%, decision-F1 46.5%->62.7%, calibration 80%->100%, p50 latency 29.8s->19.1s. muse-glimmer entity-F1 44.4%(a first run poisoned by a model-swap race right after the Gemma baseline, LM Studio hadn't settled — 14/20 fixtures came back empty at a 25ms p50, not a real timeout)/83.5%(clean re-run)->82.9%, label-accuracy 95% throughout, calibration ->100%. Entity-F1 and decision-F1 are now a wash between models (within noise at n=20); the deciding signal is label-accuracy, 70% (Gemma) vs 95% (muse-glimmer) — a ~5-fixture gap, not noise. (3) Promoted `NLM_CLASSIFIER_MODEL` to `meta/muse-glimmer` in `~/.nlm/.env` (backup `.env.bak-20260813-muse-glimmer-promotion`), daemon restarted. (4) `POST /api/classifier` hardcoded provider to `deepseek`/`ollama` only and explicitly rejected `openai` ("set via NLM_CLASSIFIER env + restart") — dead code path for the Settings > Classifier page's existing switch-model UI (provider/model dropdowns, required connection test, save), since the daemon has run classifier provider `openai` (LM Studio) since 2026-06-24. Added `ClassifierBox.openAiSwapAvailable` (true once a baseUrl was configured at daemon startup — the swap is model selection only, never endpoint selection) and let the route accept both `openai` and the Providers-registry's `openai-compatible` kind (both ride the same `DeepSeekClient` construction path). Fixed the Classifier page's active-provider matching (pre-select + dirty check) to treat both kinds as equivalent to a live `openai` classifier.
8+
9+
**Decisions:** Model choice was made on label-accuracy, not entity-F1/decision-F1 — those two were within eval noise at n=20 post-fix, and label-accuracy had the only gap large enough to trust at this sample size (design review via a second-opinion pass: hoist triviality to session level rather than duplicate it per-field, one contrastive example pair per distinct entity failure mode rather than matching `decisions`' 4+3 density — more examples at n=20 fixtures risks overfitting the prompt to this specific gold set, not generalizing the rule). Did not chase Gemma's pre-fix 80% calibration number — 4/20 fixtures is noise, and it corrected to 100% as a side effect of the triage hoist anyway.
10+
11+
**State:** `main` locally fast-forwarded through two commits (`fix(classifier): hoist session-triage rule...`, `feat(classifier): allow runtime model swap...`), full suite green modulo 4 pre-existing failures unrelated to either change (3 missing-`dist/ui`-assets, 1 `plugin.json` version drift already present on `main` before this session). NOT pushed: `main` was already 6 commits ahead of `origin/main` "by choice, public repo, not scrubbed" per the 08-12 entry below — now 48 commits ahead across that unpushed span. One own-commit scrub caught and fixed before landing: a test fixture hardcoded the real Mac Studio LAN IP, swapped for a placeholder hostname.
12+
13+
**Next:** decision-F1 sitting at ~50-60% for both models (flagged in second-opinion review) is the larger unaddressed product weakness this session's prompt fix didn't touch. Grow the classifier gold fixture set past 20 before tuning the prompt further — both this session's swings and 08-12's wrong verdict came from noise at the current sample size. The 08-12 entry below still applies for its other findings (I5a/#434, #435 fact-embed rot, unpushed backlog) — this entry only supersedes its classifier-eval verdict.
14+
515
## 2026-08-12 - classifier gate for build-artifact facts, integrity self-heal shipped, embed rot found
616

717
**Changes:** (1) Model eval: swapped gemma-4-26b for meta/muse-glimmer and qwen3.6-35b-a3b on the same 20 fixtures via `nlm eval --classifier`. Gemma stays. Muse wins label accuracy (90% vs 75%) and calibration (100% vs 80%) but loses badly on extraction - entity F1 69.3% vs 42.6%, decision F1 61.8% vs 48.9% - at 2x the latency (p50 47.5s vs 23.7s). qwen3.6-35b is strictly dominated by gemma except a 10pt calibration edge. Extraction is the half `recall_facts` is built on, so the trade points the wrong way. Context length is a non-factor either way: transcripts are hard-truncated to 15K chars at `prompt.ts:133`, so 262k vs 131k is ~30x more headroom than gets used. (2) 0.21.1: `isEphemeralSubject` added to `coerceFacts`, sibling of `isNonAnswerValue` (#325), same call site. Blocks per-run build artifacts (`tsc|status = clean`, `task-4-review|status = Approved`) from becoming durable facts. The store carried 56 I5a violations, ~16 groups of exactly this shape. (3) 0.21.2: merged `feat/ingest-embed-failure-visibility` - fact and chunk embed failures were swallowed by bare `catch{}` in the SQLite store, a silent regression against Postgres which already logged them. (4) Store cleanup: I5a 56 -> 29, I7 3 -> 0 and holding. 284 build-artifact facts retired, 229 orphaned embeddings reaped, 12 verified-stale facts superseded, ~944MB of superseded backups reclaimed. (5) Shipped a nightly integrity self-heal in the Whtnxt Agent repo (`scripts/nlm-integrity-sweep.sh`, called from the 07:00 digest).
@@ -94,12 +104,4 @@ Session-level log per session protocol. Cap: 10 entries — archive older to `CH
94104

95105
**Next:** M3 per-team token auth (swap DEFAULT_TEAM_ID at composition roots for token-resolved teams), M4 ingest attribution + name re-key, M7 pg test isolation. Program docs private in `.superpowers/sdd/`.
96106

97-
## 2026-07-22 - Tenant threading landed for SessionStore + FactStore (M2 Wave B1/B2)
98-
99-
**Changes:** Threaded non-optional `tenantId: string` as the first parameter through every `SessionStore` port method (both `sqlite-session-store.ts`/`pg-session-store.ts`, plus the concrete-only `recentWrites`/`recentMarkers`/`getSessionScopeById`/`insertSession`/`listBackfillCandidates`/`insertFactsForSession`) and every `FactStore` port method (both backends, incl. `ingestSessionFactsInTxn`). Every SELECT/UPDATE/DELETE on the `sessions`/`facts`/`entities`/`entity_variants`/`session_entities` STAMP tables now routes its WHERE fragment through `tenantClause`/`tenantClausePg` — no inline `tenant_id = ?`/`$N` left in any store file outside `tenant-clause.ts` (verified: the leak-contract's case-11 pre-threading-floor scan is green). Vector-path rule applied to both `semanticSearch` implementations (session chunk KNN and fact KNN both re-resolve candidate ids against the tenant-filtered base table before returning — a same-epsilon cross-tenant embedding no longer surfaces). By-id rule applied: `getById`/`getByIds` for a wrong-tenant id return the same not-found shape as a missing row. Three defense-in-depth `X.tenant_id = s.tenant_id` join conditions were removed as provably redundant (session_entities/facts rows are always stamped with their owning session's tenant at write time) rather than routed through the helper, to keep the guard's literal-scan invariant honest without duplicating protection SQL already provides. Compiler-driven caller threading reached ~25 call sites across `RecallService`/`FactRecallService`/ingest/reprocess/reclassify-oversized/backfill-facts/scheduler/workstream bind+rollup/work-digest/`stampSignalScope`/HTTP routes/MCP handlers/CLI commands — `DEFAULT_TEAM_ID` supplied at each composition root (route handler entry, MCP tool entry, CLI command action, scheduler construction), never hardcoded mid-chain. Leak-contract cases 1, 2, and 4 flipped from `it.todo` to real assertions against the fixture's real (now-threaded) stores; the fixture itself rewritten to seed sessions/facts through `SqliteSessionStore.insertSession`/`SqliteFactStore.insert` instead of raw SQL (code_exemplars/signals/workstreams stay raw — unthreaded, Wave B3-B6). Suite 2177 pass / 0 fail (sqlite lane); pg lane source-compiles but unexecuted (`NLM_PG_TEST_URL` not set in this environment — task #408 still open); `tsc --noEmit` and `build:server` both clean.
100-
101-
**Decisions:** Store test-only helpers (`insertSessionForTest`, `insertEdgeForTest`) keep a `tenantId = DEFAULT_TEAM_ID` default parameter — the one deliberate exception to "no defaults in store signatures," scoped to test seams outside the port contract. `run-eval.ts`'s `Searcher` interface stayed single-arg; callers wrap `RecallService`/`FactRecallService` in a one-line adapter at each of the 3 call sites rather than widening the eval harness's own type.
102-
103-
**State:** main, uncommitted at CHANGELOG-write time (commits land right after, one per Wave-B task per the plan). Zero behavior change in local mode — `DEFAULT_TEAM_ID` is the only value passed anywhere until M3 replaces it with the token-resolved tenant.
104-
105-
**Next:** Wave B3-B6 (CodeExemplarStore, SignalStore/WorkstreamStore/EntityStore/OutcomeStore, Source/Provider registries, `actions-log.ts`/`build-dataset.ts`/data-stats raw-SQL modules), then Wave C (MCP/HTTP surface threading + hosted-mode gates + the full guard test with an allowlist + contract completion). #408 (pg test instance) blocks trusting the pg lane on this wave's changes.
107+
_Older entries archived in CHANGELOG-2026.md_

0 commit comments

Comments
 (0)