Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 84 additions & 0 deletions artifacts/rewrite-efficacy-study4/bridge-en-gemini-3.7-flash.jsonl

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
{
"computed_at": "2026-09-03T23:28:49.538Z",
"stage": "en",
"s1_arm": "A1",
"rule": {
"min_auc": 0.85,
"spearman_slack": 0.1,
"reference": "grok-vs-gpt on the same passages"
},
"passages_scored": 84,
"originals": 42,
"n_ai": 21,
"n_human": 21,
"auc_gemini": 0.9943310657596371,
"auc_gpt_reference": 1,
"auc_grok_reference": 0.9977324263038548,
"n_paired": 84,
"spearman_gemini_gpt": 0.9109976002674424,
"spearman_grok_gpt": 0.8943306674236398,
"spearman_gemini_grok": 0.8805498403188557,
"candidate": "judge-gemini-3.7-flash",
"admitted": true,
"judges": [
"judge-gpt",
"judge-gemini-3.7-flash"
],
"label": "panel-v2-deviation-gemini"
}
42 changes: 42 additions & 0 deletions artifacts/rewrite-efficacy-study4/s4-rows-en.jsonl

Large diffs are not rendered by default.

16 changes: 16 additions & 0 deletions docs/research/2026-rewrite-efficacy-prereg.md
Original file line number Diff line number Diff line change
Expand Up @@ -684,6 +684,22 @@ Bridge results on the 108 archived S1-D passages (reference on the same passages

Selection rule (highest AUC among admitted) picks **`judge-gemini-3.7-flash`**. From this note on, stage 1 is a two-judge panel (`judge-gpt` + `judge-gemini-3.7-flash`) plus det: the admitted judge scores every stored P/S body from rows finished so far, and the runner is restarted with both judges for the remaining rows. The results doc drops the *single-perceptual-judge* label and reports every candidate above. The deepseek bridge may be resumed for the record only; it cannot change the selection.

### Study 4 — dated note 2026-09-04 (before any stage 2 row): stage 2 (en) started on the owner's instruction

Stage 1 (ko) closed on 2026-09-03 with H-4b not supported. The owner asked to resume; stage 2 runs under the identical registered rules. Before the first row: (1) the 42 Arm-A1 documents inherit Study 1's `spok` exclusion (pilot Deviation 4); (2) the second-judge seat is re-bridged on the 84 archived Arm-A1 passages (42 originals + 42 S1 rewrites) with the English judge prompt under the same admission rule — `judge-gemini-3.7-flash` was admitted on Korean only. The bridge verdict is recorded in `artifacts/rewrite-efficacy-study4/bridge-en-verdict-gemini-3.7-flash.json`; if it fails, `judge-gemini-3.1-pro` is bridged next, and if neither passes stage 2 runs single-perceptual-judge as registered. Plumbing is smoke-tested on a synthetic English paragraph only.

### Study 4 — dated note 2026-09-04 (before any stage 2 row): stage 2 judge seat

EN bridge on the 84 archived Arm-A1 passages: `judge-gemini-3.7-flash` (Gemini API, the local research key — not the product keys, which the owner rotated to product-only use the same day) scored 84/84, **AUC 0.994**, Spearman vs gpt **0.911** (reference: gpt AUC 1.000, grok AUC 0.998, Spearman(grok, gpt) 0.894) — admitted. A subscription-only transport of the same model (`judge-gemini-3.7-flash-cli`, gemini CLI) was bridged in parallel to avoid API-key use; it produced 8/84 scores and then timed out three times in a row on consecutive passages (180 s each), so it is **not admitted** and its partial rows are discarded. Stage 2 therefore runs with the same two-judge panel as stage 1: `judge-gpt` + `judge-gemini-3.7-flash` (API) plus the det chief.

### Study 4 — correction 2026-09-04: the "CLI" gemini transport was also API-key authenticated

On this machine the gemini CLI's effective auth is `gemini-api-key` (`~/.gemini/settings.json` → `security.auth.selectedType`) and `GEMINI_API_KEY` is exported in the login shell, so every gemini CLI call in this study — the 2026-09-02 gemini-2.5-pro bridge (108 passages) and the 2026-09-04 `judge-gemini-3.7-flash-cli` bridge (8 passages before timeouts) — was billed to the local research key, not to a subscription. The earlier notes' "subscription-only" wording for that transport is withdrawn; the transports differ only in sampling defaults. Nothing about the admitted judge, the panel, or any row changes.

### Study 4 — dated note 2026-09-04 (after stage 2): stage 2 closed; detail-token definition flaw on English

Stage 2 (en, 42/42) closed **not supported** (paired d +3.7 [−3.5, +10.8]; det +12.1 [+4.5, +20.6]; floor met 42/42; meaning gate 38/42 violated). Post-hoc observation, not a criterion change: the registered detail-token regex `[A-Za-z][A-Za-z0-9+._-]+` matches every English word, so H-4b-c retention and guard rail 3 are vocabulary-overlap measures on English and are reported as such in the results doc. Both stages are complete; nothing ships; the next candidate is H-4a.

## Sources
- Self-Preference Bias in LLM-as-a-Judge — arXiv:2410.21819
- TH-Bench (humanizing attacks vs detectors) — arXiv:2503.08708
Expand Down
67 changes: 58 additions & 9 deletions docs/research/2026-rewrite-efficacy-study4.md
Original file line number Diff line number Diff line change
Expand Up @@ -161,12 +161,61 @@ the production prompt. The next candidate on the registered list is **H-4a —
deterministic merge/split with seam-only LLM infill** — the first mechanism
that acts on the architecture the judges keep naming.

## Stage 2 (en)

Registered to run after stage 1 with identical rules. Not started as of
2026-09-03: the runner spends the owner's Claude subscription window (the
same quota interactive sessions use), and the Korean result is decisive. If
the owner wants the English data point, the bridge must first be re-run on the
84 archived Study 1 Arm-A1 passages (the gemini-3.7-flash seat was admitted on
Korean only), then `S4_STAGE=en` runs the same harness. A deferral is
recorded here as a dated note per the registration; it is not a withdrawal.
## Stage 2 (en) — completed 2026-09-04: **NOT supported either**

Run 2026-09-04 07:02 → 10:15 UTC on the owner's instruction, 2 supervisor
passes, no session-limit hits. The 42 Study 1 Arm-A1 documents (21 AI + 21
human; `spok` excluded per pilot Deviation 4), same rewriter, English
production prompt (template `72428aea…`), English constraint block
(`81e2e53c…`) and retry suffix (`b320025b…`). Panel: `judge-gpt` +
`judge-gemini-3.7-flash` (re-admitted on the 84 archived English passages:
AUC 0.994, Spearman vs gpt 0.911 against grok's 0.894; the gemini CLI
transport of the same model timed out and was not admitted). Det chief
fresh-corpus accuracy on the English originals was **0.762 < 0.85**, so per
panel v2 the binary det column is disabled and only the continuous det score
is reported.

| | P (plain) | S (specificity) | paired d = S − P (95% CI) |
|---|---:|---:|---|
| AI docs, panel AI-likeness (n=21) | 81.8 | **85.5** | **+3.7 [−3.5, +10.8]** |
| AI docs, det chief (continuous) | 11.3 | **23.4** | **+12.1 [+4.5, +20.6]** |
| human docs, panel (n=21) | 10.5 | 9.5 | −0.9 [−3.3, +1.9] (guard rail 2 held) |

- **H-4b-a: NOT SUPPORTED.** The point estimate is positive again and the
interval includes zero. AI-call rate 95.2% → 95.2% (discordant 2 vs 2).
- **The floor executed this time.** Unlike Korean, every English S row met
the 98% floor (42/42; mean 1.67 attempts on AI docs). The plain rewrite
compresses English hard (P length ratio **0.806** on AI docs, 0.908 on
human); S restored it to 1.03 / 1.00. So stage 2 is the clean test of the
hypothesis that stage 1 could not deliver: with the length actually held,
the constrained rewrite still reads *more* AI-like to both the panel and
the deterministic scorer.
- **Copy check:** S median trigram Jaccard 0.630 vs P 0.540 on AI docs; on
human docs 0.885 vs 0.746 (just under the 0.90 "rewrote less" line).
- **Guard rail 1: 38/42 = 90.5% — VIOLATED** (no corpus-artifact documents
in this arm; four real dropped numbers in S; P 37/42). Rails 2 and 3 held.
- **Cue mix on AI docs still called "ai":** structural share fell from 35%
(P) to 18% (S) while the specificity-absence class rose from 11 to 17 of 40
cues. English judges name different residual tells than Korean judges
(Korean: 84–92% structural), and keeping the length did not remove them.
- **Descriptive gpt-only:** S1 rw1 74.0 | P 86.0 | S 89.5 on AI docs — the
same upward drift of today's plain rewrite versus July's that stage 1
showed (prompt changed twice since; judge drift possible).

**Metric caveat found after the fact (not a criterion change):** the
registered detail-token definition counts every Latin token of two or more
letters, which in English is every word. The English "retention" (P 45% →
S 64%) and "added tokens" (~130 per document in both arms) therefore measure
vocabulary overlap, not concrete details, and guard rail 3 is vacuous for
English. The numbers-and-quotes part of the definition is unaffected; a
future English stage needs a detail definition that excludes ordinary
words. Recorded here and in the prereg as a dated note.

### Combined verdict

Both stages fail the registered primary criterion in the same direction and
both violate the meaning gate. In Korean the model would not hold the floor;
in English it held the floor and the result got worse. The hypothesis that
the plain rewrite's lost specificity is what judges react to is not
supported in either language. Nothing ships; the next candidate remains
H-4a.
28 changes: 19 additions & 9 deletions scripts/research/rewrite-efficacy-study4-bridge.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,21 @@ import {
} from './study4-common.mjs';

// BRIDGE_JUDGE selects the candidate (default: the original gemini CLI judge).
// BRIDGE_STAGE=en bridges on the archived Study 1 Arm-A1 passages (spok
// excluded per pilot Deviation 4) with the English judge prompt; files carry
// an `en-` prefix. Default stage ko = Arm D.
const JUDGE_ID = process.env.BRIDGE_JUDGE || 'judge-gemini';
const JUDGE = JUDGE_DEFS[JUDGE_ID];
if (!JUDGE) throw new Error(`unknown bridge judge ${JUDGE_ID}`);
const STAGE = process.env.BRIDGE_STAGE === 'en' ? 'en' : 'ko';
const S1_ARM = STAGE === 'en' ? 'A1' : 'D';
const EXCLUDED_REGISTERS = new Set(['spok']);
const SUFFIX = JUDGE_ID === 'judge-gemini' ? '' : `-${JUDGE_ID.replace(/^judge-/u, '')}`;
const ROWS = join(OUT_DIR, `bridge-gemini${SUFFIX}.jsonl`.replace('bridge-gemini-', 'bridge-'));
const VERDICT = join(OUT_DIR, `bridge-verdict${SUFFIX}.json`);
const LOG = join(OUT_DIR, `bridge-run${SUFFIX}.log`);
const STAGE_PREFIX = STAGE === 'en' ? 'en-' : '';
const ROWS = join(OUT_DIR, `bridge-${STAGE_PREFIX}gemini${SUFFIX}.jsonl`.replace(`bridge-${STAGE_PREFIX}gemini-`, `bridge-${STAGE_PREFIX}`));
const VERDICT = join(OUT_DIR, `bridge-${STAGE_PREFIX}verdict${SUFFIX}.json`);
const LOG = join(OUT_DIR, `bridge-${STAGE_PREFIX}run${SUFFIX}.log`);
const EXPECTED = STAGE === 'en' ? 84 : 108;
const RULE = { min_auc: 0.85, spearman_slack: 0.10, reference: 'grok-vs-gpt on the same passages' };

function analyze(log) {
Expand All @@ -36,10 +44,12 @@ function analyze(log) {
const aucGpt = auc(originals.filter((r) => r.source_class === 'ai').map((r) => r.gpt?.ai_likeness).filter(Number.isFinite), originals.filter((r) => r.source_class === 'human').map((r) => r.gpt?.ai_likeness).filter(Number.isFinite));
const aucGrok = auc(originals.filter((r) => r.source_class === 'ai').map((r) => r.grok?.ai_likeness).filter(Number.isFinite), originals.filter((r) => r.source_class === 'human').map((r) => r.grok?.ai_likeness).filter(Number.isFinite));
const complete = rows.length;
const admitted = complete >= 100 && aucGemini !== null && rhoGeminiGpt !== null && rhoGrokGpt !== null
const admitted = complete >= Math.floor(EXPECTED * 0.93) && aucGemini !== null && rhoGeminiGpt !== null && rhoGrokGpt !== null
&& aucGemini >= RULE.min_auc && rhoGeminiGpt >= rhoGrokGpt - RULE.spearman_slack;
const verdict = {
computed_at: new Date().toISOString(),
stage: STAGE,
s1_arm: S1_ARM,
rule: RULE,
passages_scored: complete,
originals: originals.length,
Expand All @@ -58,7 +68,7 @@ function analyze(log) {
label: admitted ? 'panel-v2-deviation-gemini' : 'single-perceptual-judge',
};
writeFileSync(VERDICT, JSON.stringify(verdict, null, 2) + '\n');
log(`bridge verdict [${JUDGE_ID}]: scored ${complete}/108; AUC gemini ${fmt(aucGemini)} (gpt ${fmt(aucGpt)}, grok ${fmt(aucGrok)}); rho gemini-gpt ${fmt(rhoGeminiGpt)} vs grok-gpt ${fmt(rhoGrokGpt)} (gemini-grok ${fmt(rhoGeminiGrok)}); admitted=${admitted}`);
log(`bridge verdict [${JUDGE_ID}] stage ${STAGE}: scored ${complete}/${EXPECTED}; AUC gemini ${fmt(aucGemini)} (gpt ${fmt(aucGpt)}, grok ${fmt(aucGrok)}); rho gemini-gpt ${fmt(rhoGeminiGpt)} vs grok-gpt ${fmt(rhoGrokGpt)} (gemini-grok ${fmt(rhoGeminiGrok)}); admitted=${admitted}`);
return verdict;
}

Expand All @@ -70,8 +80,8 @@ async function main() {
const log = makeLogger(LOG);
if (process.argv.includes('--analyze')) { analyze(log); return; }

const s1Rows = readJsonl(join(S1_DIR, 's1-rows-D.jsonl'));
const texts = loadS1Texts('D');
const s1Rows = readJsonl(join(S1_DIR, `s1-rows-${S1_ARM}.jsonl`)).filter((r) => !EXCLUDED_REGISTERS.has(r.register));
const texts = loadS1Texts(S1_ARM);
const passages = [];
for (const r of s1Rows) {
const t = texts.get(r.original_sha);
Expand All @@ -80,11 +90,11 @@ async function main() {
if (r.rewrite_sha && t.rewritten) passages.push({ key: `${r.original_sha}:rewrite`, original_sha: r.original_sha, cond: 'rewrite', source_class: r.source_class, text: t.rewritten, archived: r.judges?.rewrite ?? {} });
}
const done = new Set(readJsonl(ROWS).filter((r) => r.gemini && Number.isFinite(r.gemini.ai_likeness)).map((r) => r.key));
log(`bridge start [${JUDGE_ID}] — ${passages.length} passages, ${done.size} already scored`);
log(`bridge start [${JUDGE_ID}] stage ${STAGE} (S1 arm ${S1_ARM}) — ${passages.length} passages, ${done.size} already scored`);
let failures = 0;
for (const p of passages) {
if (done.has(p.key)) continue;
const gemini = await judgeOnce(JUDGE, p.text, 'ko');
const gemini = await judgeOnce(JUDGE, p.text, STAGE);
if (gemini.error) {
failures += 1;
log(`${p.key}: gemini FAILED — ${gemini.error}`);
Expand Down
11 changes: 8 additions & 3 deletions scripts/research/rewrite-efficacy-study4.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,12 @@ const PREFIX = SMOKE_FILE ? 's4-smoke' : 's4';
const OUT_JSONL = join(OUT_DIR, `${PREFIX}-rows-${STAGE}.jsonl`);
const TEXTS_JSONL = join(OUT_DIR, `${PREFIX}-texts-${STAGE}.private.jsonl`);
const LOG = join(OUT_DIR, `${PREFIX}-run-${STAGE}.log`);
const VERDICT = join(OUT_DIR, 'bridge-verdict.json');
// Stage-specific bridge verdict for the admitted second judge (ko: the
// gemini-3.7-flash amendment verdict; en: the en bridge). S4_JUDGES overrides.
const VERDICT = join(OUT_DIR, STAGE === 'en' ? 'bridge-en-verdict-gemini-3.7-flash.json' : 'bridge-verdict-gemini-3.7-flash.json');
// Pilot Deviation 4: the HAP-E `spok` register is not written prose; Study 1
// excluded it and stage 2 inherits the exclusion (42 = 21 + 21 documents).
const EXCLUDED_REGISTERS = new Set(['spok']);

const REWRITE_TIMEOUT_MS = 900_000; // raised 600s -> 900s at doc 12 (S attempt 2 timeout), Study 3 precedent; execution note, no criterion change
const FLOOR = 0.98;
Expand Down Expand Up @@ -197,7 +202,7 @@ async function judgeAll(judges, text, lang) {
function resolveJudges(log) {
if (process.env.S4_JUDGES) return process.env.S4_JUDGES.split(',').map((s) => s.trim()).filter(Boolean);
if (process.env.S4_USE_GROK === '1') return ['judge-gpt', 'judge-grok'];
if (!existsSync(VERDICT)) throw new Error('bridge-verdict.json missing — run rewrite-efficacy-study4-bridge.mjs first (registered admission rule)');
if (!existsSync(VERDICT)) throw new Error(`${VERDICT} missing — run rewrite-efficacy-study4-bridge.mjs for this stage first (registered admission rule)`);
const verdict = JSON.parse(readFileSync(VERDICT, 'utf8'));
log(`bridge verdict ${verdict.label}: judges ${verdict.judges.join(', ')} (AUC gemini ${verdict.auc_gemini?.toFixed?.(3)}, rho gemini-gpt ${verdict.spearman_gemini_gpt?.toFixed?.(3)} vs grok-gpt ${verdict.spearman_grok_gpt?.toFixed?.(3)})`);
return verdict.judges;
Expand All @@ -217,7 +222,7 @@ async function main() {
const judges = resolveJudges(log);
for (const id of judges) if (!JUDGE_DEFS[id]) throw new Error(`unknown judge ${id}`);

let s1Rows = readJsonl(join(S1_DIR, `s1-rows-${S1_ARM}.jsonl`)).filter((r) => r.original_sha);
let s1Rows = readJsonl(join(S1_DIR, `s1-rows-${S1_ARM}.jsonl`)).filter((r) => r.original_sha && !EXCLUDED_REGISTERS.has(r.register));
let texts = loadS1Texts(S1_ARM);
if (SMOKE_FILE) {
const smokeText = readFileSync(SMOKE_FILE, 'utf8').trim();
Expand Down
11 changes: 10 additions & 1 deletion scripts/research/study4-common.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -18,21 +18,30 @@ export const JUDGE_TIMEOUT_MS = 180_000;
export const JUDGE_ATTEMPTS = 3;

const HTTP_CLI = [join('scripts', 'research', 'judge-http-cli.mjs')];
// Detached runners can start without nvm on PATH; resolve the gemini CLI once
// (S3 did the same for codex). Falls back to a bare `gemini` lookup.
const GEMINI_BIN = process.env.GEMINI_BIN || ['/home/devswha/.nvm/versions/node/v24.18.0/bin/gemini', join(process.env.HOME || '', '.nvm', 'versions', 'node', 'v24.18.0', 'bin', 'gemini')].find((p) => existsSync(p)) || 'gemini';
const http = (id, family, baseURL, keyEnv, model) => ({
id, family, cmd: 'node', args: HTTP_CLI,
env: { JUDGE_BASE_URL: baseURL, JUDGE_API_KEY_ENV: keyEnv, JUDGE_MODEL: model },
});

export const JUDGE_DEFS = Object.freeze({
'judge-gpt': { id: 'judge-gpt', family: 'gpt', cmd: 'codex', args: ['exec', '--skip-git-repo-check', '--sandbox', 'read-only'] },
'judge-gemini': { id: 'judge-gemini', family: 'gemini', cmd: 'gemini', args: ['-p', '', '--output-format', 'text', '--skip-trust', '--allowed-mcp-server-names', NO_MCP_SERVERS, '-m', 'gemini-2.5-pro'] },
'judge-gemini': { id: 'judge-gemini', family: 'gemini', cmd: GEMINI_BIN, args: ['-p', '', '--output-format', 'text', '--skip-trust', '--allowed-mcp-server-names', NO_MCP_SERVERS, '-m', 'gemini-2.5-pro'] },
'judge-grok': { id: 'judge-grok', family: 'xai', cmd: 'node', args: [join('scripts', 'research', 'xai-cli.mjs')] },
// Bridge candidates without xAI credit (2026-09-02 amendment): API judges on
// the OpenAI-compatible endpoints patina's providers already use.
'judge-gemini-3.7-flash': http('judge-gemini-3.7-flash', 'gemini', 'https://generativelanguage.googleapis.com/v1beta/openai', 'GEMINI_API_KEY', 'gemini-3.7-flash'),
'judge-gemini-3.1-pro': http('judge-gemini-3.1-pro', 'gemini', 'https://generativelanguage.googleapis.com/v1beta/openai', 'GEMINI_API_KEY', 'gemini-3.1-pro-preview'),
'judge-deepseek-v4-pro': http('judge-deepseek-v4-pro', 'deepseek', 'https://api.deepseek.com', 'DEEPSEEK_API_KEY', 'deepseek-v4-pro'),
'judge-kimi-k3': http('judge-kimi-k3', 'moonshot', 'https://api.moonshot.ai/v1', 'KIMI_API_KEY', 'kimi-k3'),
// Same model as judge-gemini-3.7-flash through the gemini CLI. NOTE: on a
// machine where the CLI's auth type is gemini-api-key (or GEMINI_API_KEY is
// exported) this is API-key billed too — it is NOT a subscription path.
// Transport differs (CLI default sampling vs API temperature 0), so it is
// bridged separately before use. Not admitted for Study 4 (timeouts).
'judge-gemini-3.7-flash-cli': { id: 'judge-gemini-3.7-flash-cli', family: 'gemini', cmd: GEMINI_BIN, args: ['-p', '', '--output-format', 'text', '--skip-trust', '--allowed-mcp-server-names', NO_MCP_SERVERS, '-m', 'gemini-3.7-flash'] },
});

export const sha = (s) => createHash('sha256').update(s, 'utf8').digest('hex').slice(0, 16);
Expand Down