|
2 | 2 |
|
3 | 3 | Session-level log per session protocol. Cap: 10 entries — archive older to `CHANGELOG-YYYY.md` when exceeded. |
4 | 4 |
|
| 5 | +## 2026-08-13 - Classifier prompt fix corrects the 08-12 muse-glimmer verdict; model promoted; runtime swap unblocked in the UI |
| 6 | + |
| 7 | +**Changes:** (1) Root-caused the 08-12 eval's "Gemma stays" call: 4 of the 20 gold fixtures are trivial single-turn lookups (version check, `package.json` dump) where the reference expects `entities: []`, but both classifiers were extracting incidental proper nouns from command output regardless — muse-glimmer paid for it worse (its granular extraction style also surfaced file paths and hostnames the reference never listed), which is what produced 08-12's entity-F1 69.3% vs 42.6% gap. Fixed `CLASSIFIER_SYSTEM_PROMPT` (`prompt.ts`): hoisted a single "Session triage: WORK or LOOKUP" instruction above the field list (previously the trivial-session rule was only half-stated, separately, inside `confidence`'s description, with no shared definition across `entities`/`decisions`/`facts`), and added an entity-stability rule (file paths, hostnames, URLs, SHAs are never entities — redirect into `facts[]` under the matching predicate instead of dropping them). (2) Re-ran both models on the fixed prompt, n=20 each: Gemma entity-F1 69.4%->80.7%, decision-F1 46.5%->62.7%, calibration 80%->100%, p50 latency 29.8s->19.1s. muse-glimmer entity-F1 44.4%(a first run poisoned by a model-swap race right after the Gemma baseline, LM Studio hadn't settled — 14/20 fixtures came back empty at a 25ms p50, not a real timeout)/83.5%(clean re-run)->82.9%, label-accuracy 95% throughout, calibration ->100%. Entity-F1 and decision-F1 are now a wash between models (within noise at n=20); the deciding signal is label-accuracy, 70% (Gemma) vs 95% (muse-glimmer) — a ~5-fixture gap, not noise. (3) Promoted `NLM_CLASSIFIER_MODEL` to `meta/muse-glimmer` in `~/.nlm/.env` (backup `.env.bak-20260813-muse-glimmer-promotion`), daemon restarted. (4) `POST /api/classifier` hardcoded provider to `deepseek`/`ollama` only and explicitly rejected `openai` ("set via NLM_CLASSIFIER env + restart") — dead code path for the Settings > Classifier page's existing switch-model UI (provider/model dropdowns, required connection test, save), since the daemon has run classifier provider `openai` (LM Studio) since 2026-06-24. Added `ClassifierBox.openAiSwapAvailable` (true once a baseUrl was configured at daemon startup — the swap is model selection only, never endpoint selection) and let the route accept both `openai` and the Providers-registry's `openai-compatible` kind (both ride the same `DeepSeekClient` construction path). Fixed the Classifier page's active-provider matching (pre-select + dirty check) to treat both kinds as equivalent to a live `openai` classifier. |
| 8 | + |
| 9 | +**Decisions:** Model choice was made on label-accuracy, not entity-F1/decision-F1 — those two were within eval noise at n=20 post-fix, and label-accuracy had the only gap large enough to trust at this sample size (design review via a second-opinion pass: hoist triviality to session level rather than duplicate it per-field, one contrastive example pair per distinct entity failure mode rather than matching `decisions`' 4+3 density — more examples at n=20 fixtures risks overfitting the prompt to this specific gold set, not generalizing the rule). Did not chase Gemma's pre-fix 80% calibration number — 4/20 fixtures is noise, and it corrected to 100% as a side effect of the triage hoist anyway. |
| 10 | + |
| 11 | +**State:** `main` locally fast-forwarded through two commits (`fix(classifier): hoist session-triage rule...`, `feat(classifier): allow runtime model swap...`), full suite green modulo 4 pre-existing failures unrelated to either change (3 missing-`dist/ui`-assets, 1 `plugin.json` version drift already present on `main` before this session). NOT pushed: `main` was already 6 commits ahead of `origin/main` "by choice, public repo, not scrubbed" per the 08-12 entry below — now 48 commits ahead across that unpushed span. One own-commit scrub caught and fixed before landing: a test fixture hardcoded the real Mac Studio LAN IP, swapped for a placeholder hostname. |
| 12 | + |
| 13 | +**Next:** decision-F1 sitting at ~50-60% for both models (flagged in second-opinion review) is the larger unaddressed product weakness this session's prompt fix didn't touch. Grow the classifier gold fixture set past 20 before tuning the prompt further — both this session's swings and 08-12's wrong verdict came from noise at the current sample size. The 08-12 entry below still applies for its other findings (I5a/#434, #435 fact-embed rot, unpushed backlog) — this entry only supersedes its classifier-eval verdict. |
| 14 | + |
5 | 15 | ## 2026-08-12 - classifier gate for build-artifact facts, integrity self-heal shipped, embed rot found |
6 | 16 |
|
7 | 17 | **Changes:** (1) Model eval: swapped gemma-4-26b for meta/muse-glimmer and qwen3.6-35b-a3b on the same 20 fixtures via `nlm eval --classifier`. Gemma stays. Muse wins label accuracy (90% vs 75%) and calibration (100% vs 80%) but loses badly on extraction - entity F1 69.3% vs 42.6%, decision F1 61.8% vs 48.9% - at 2x the latency (p50 47.5s vs 23.7s). qwen3.6-35b is strictly dominated by gemma except a 10pt calibration edge. Extraction is the half `recall_facts` is built on, so the trade points the wrong way. Context length is a non-factor either way: transcripts are hard-truncated to 15K chars at `prompt.ts:133`, so 262k vs 131k is ~30x more headroom than gets used. (2) 0.21.1: `isEphemeralSubject` added to `coerceFacts`, sibling of `isNonAnswerValue` (#325), same call site. Blocks per-run build artifacts (`tsc|status = clean`, `task-4-review|status = Approved`) from becoming durable facts. The store carried 56 I5a violations, ~16 groups of exactly this shape. (3) 0.21.2: merged `feat/ingest-embed-failure-visibility` - fact and chunk embed failures were swallowed by bare `catch{}` in the SQLite store, a silent regression against Postgres which already logged them. (4) Store cleanup: I5a 56 -> 29, I7 3 -> 0 and holding. 284 build-artifact facts retired, 229 orphaned embeddings reaped, 12 verified-stale facts superseded, ~944MB of superseded backups reclaimed. (5) Shipped a nightly integrity self-heal in the Whtnxt Agent repo (`scripts/nlm-integrity-sweep.sh`, called from the 07:00 digest). |
@@ -94,12 +104,4 @@ Session-level log per session protocol. Cap: 10 entries — archive older to `CH |
94 | 104 |
|
95 | 105 | **Next:** M3 per-team token auth (swap DEFAULT_TEAM_ID at composition roots for token-resolved teams), M4 ingest attribution + name re-key, M7 pg test isolation. Program docs private in `.superpowers/sdd/`. |
96 | 106 |
|
97 | | -## 2026-07-22 - Tenant threading landed for SessionStore + FactStore (M2 Wave B1/B2) |
98 | | - |
99 | | -**Changes:** Threaded non-optional `tenantId: string` as the first parameter through every `SessionStore` port method (both `sqlite-session-store.ts`/`pg-session-store.ts`, plus the concrete-only `recentWrites`/`recentMarkers`/`getSessionScopeById`/`insertSession`/`listBackfillCandidates`/`insertFactsForSession`) and every `FactStore` port method (both backends, incl. `ingestSessionFactsInTxn`). Every SELECT/UPDATE/DELETE on the `sessions`/`facts`/`entities`/`entity_variants`/`session_entities` STAMP tables now routes its WHERE fragment through `tenantClause`/`tenantClausePg` — no inline `tenant_id = ?`/`$N` left in any store file outside `tenant-clause.ts` (verified: the leak-contract's case-11 pre-threading-floor scan is green). Vector-path rule applied to both `semanticSearch` implementations (session chunk KNN and fact KNN both re-resolve candidate ids against the tenant-filtered base table before returning — a same-epsilon cross-tenant embedding no longer surfaces). By-id rule applied: `getById`/`getByIds` for a wrong-tenant id return the same not-found shape as a missing row. Three defense-in-depth `X.tenant_id = s.tenant_id` join conditions were removed as provably redundant (session_entities/facts rows are always stamped with their owning session's tenant at write time) rather than routed through the helper, to keep the guard's literal-scan invariant honest without duplicating protection SQL already provides. Compiler-driven caller threading reached ~25 call sites across `RecallService`/`FactRecallService`/ingest/reprocess/reclassify-oversized/backfill-facts/scheduler/workstream bind+rollup/work-digest/`stampSignalScope`/HTTP routes/MCP handlers/CLI commands — `DEFAULT_TEAM_ID` supplied at each composition root (route handler entry, MCP tool entry, CLI command action, scheduler construction), never hardcoded mid-chain. Leak-contract cases 1, 2, and 4 flipped from `it.todo` to real assertions against the fixture's real (now-threaded) stores; the fixture itself rewritten to seed sessions/facts through `SqliteSessionStore.insertSession`/`SqliteFactStore.insert` instead of raw SQL (code_exemplars/signals/workstreams stay raw — unthreaded, Wave B3-B6). Suite 2177 pass / 0 fail (sqlite lane); pg lane source-compiles but unexecuted (`NLM_PG_TEST_URL` not set in this environment — task #408 still open); `tsc --noEmit` and `build:server` both clean. |
100 | | - |
101 | | -**Decisions:** Store test-only helpers (`insertSessionForTest`, `insertEdgeForTest`) keep a `tenantId = DEFAULT_TEAM_ID` default parameter — the one deliberate exception to "no defaults in store signatures," scoped to test seams outside the port contract. `run-eval.ts`'s `Searcher` interface stayed single-arg; callers wrap `RecallService`/`FactRecallService` in a one-line adapter at each of the 3 call sites rather than widening the eval harness's own type. |
102 | | - |
103 | | -**State:** main, uncommitted at CHANGELOG-write time (commits land right after, one per Wave-B task per the plan). Zero behavior change in local mode — `DEFAULT_TEAM_ID` is the only value passed anywhere until M3 replaces it with the token-resolved tenant. |
104 | | - |
105 | | -**Next:** Wave B3-B6 (CodeExemplarStore, SignalStore/WorkstreamStore/EntityStore/OutcomeStore, Source/Provider registries, `actions-log.ts`/`build-dataset.ts`/data-stats raw-SQL modules), then Wave C (MCP/HTTP surface threading + hosted-mode gates + the full guard test with an allowlist + contract completion). #408 (pg test instance) blocks trusting the pg lane on this wave's changes. |
| 107 | +_Older entries archived in CHANGELOG-2026.md_ |
0 commit comments