Skip to content

research: rewrite-efficacy Study 4 — stage 2 (en) - #720

Merged
devswha merged 5 commits into
devfrom
research/rewrite-efficacy-study4-en
Sep 4, 2026
Merged

research: rewrite-efficacy Study 4 — stage 2 (en)#720
devswha merged 5 commits into
devfrom
research/rewrite-efficacy-study4-en

Conversation

@devswha

@devswha devswha commented Sep 3, 2026

Copy link
Copy Markdown
Owner

Summary

Study 4 stage 2 (English) complete: H-4b NOT supported, same direction as Korean. Both stages closed; nothing ships; next candidate H-4a. Results section added to docs/research/2026-rewrite-efficacy-study4.md; dated notes in the prereg cover the resume, the judge seat, the API-key correction, and a post-hoc metric caveat.

  • Corpus: the 42 Study 1 Arm-A1 documents (21 AI + 21 human; spok excluded). Same rewriter claude-sonnet-4-6; English production prompt + English constraint block, all fixed before the first row.
  • Judges: judge-gpt + judge-gemini-3.7-flash (Gemini API; re-admitted on the 84 archived English passages: AUC 0.994, Spearman vs gpt 0.911 vs grok's 0.894) + det chief with the binary column disabled (fresh-corpus accuracy 0.762). The gemini CLI transport of the same model timed out and was not admitted; a correction note records that on this machine the CLI was API-key authenticated anyway.

Results (42/42)

P S paired d = S − P
AI docs, panel (n=21) 81.8 85.5 +3.7 [−3.5, +10.8] (needed ≤ −3, CI upper < 0)
AI docs, det chief (continuous) 11.3 23.4 +12.1 [+4.5, +20.6]
human docs, panel 10.5 9.5 −0.9 [−3.3, +1.9] (rail 2 held)
  • The floor executed on 42/42 (plain rewrite compresses English to 0.81; S restored 1.03), so this is the clean test stage 1 could not deliver — and S still reads more AI-like to the panel and the deterministic scorer.
  • Guard rail 1 (meaning gate on S) 38/42 = 90.5% → violated; four real dropped numbers; P 37/42.
  • Post-hoc caveat (dated note, no criterion change): the registered detail-token regex matches every English word, so retention and rail 3 are vocabulary-overlap measures on English.
  • Cue mix on AI docs still called "ai": structural 35% (P) → 18% (S); specificity-absence 11 → 17 of 40.

Test plan

  • node --check + eslint on the changed scripts
  • EN bridge verdicts (API: admitted; CLI: not admitted) committed
  • EN smoke run on a synthetic paragraph only
  • stage 2 complete: 42/42 rows clean, 0 quorum loss; analysis per the registered rules
  • prose gate on the changed docs; committed rows are hash-and-score only (verified: no text fields)

Merge with a merge commit: rows reference commits on this branch.

🤖 Generated with Claude Code

https://claude.ai/code/session_018ePnwo4nBWCWvFvpy1YgCN

…k exclusion

Owner asked to resume after stage 1 closed; the note is recorded before any
stage 2 row. The bridge now takes BRIDGE_STAGE=en (Arm A1, English judge
prompt, 84 passages) and the runner inherits Study 1's spok exclusion and
reads the stage-specific admission verdict.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ePnwo4nBWCWvFvpy1YgCN
@vercel

vercel Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
patina Ready Ready Preview Sep 4, 2026 10:19am UTC

Request Review

…iption-path gemini judge

- judge-gemini-3.7-flash (Gemini API) on the 84 archived Arm-A1 passages:
  AUC 0.994, Spearman vs gpt 0.911 (grok reference 0.894) — admitted.
- Adds judge-gemini-3.7-flash-cli: the same model through the gemini CLI's
  subscription login (no API key), bridged separately; the gemini binary is
  resolved once so detached runners work without nvm on PATH.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ePnwo4nBWCWvFvpy1YgCN
…on EN; CLI transport not admitted (timeouts)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ePnwo4nBWCWvFvpy1YgCN
…uthenticated on this machine

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ePnwo4nBWCWvFvpy1YgCN
…ges closed

- 42/42 documents, gpt + gemini-3.7-flash + det (binary disabled: fresh-corpus
  accuracy 0.762).
- Primary paired d(S − P) on AI docs +3.7 [−3.5, +10.8]; det chief +12.1
  [+4.5, +20.6]; AI-call rate unchanged 95.2%.
- The 98% floor was met on 42/42 (plain rewrite compresses English to 0.81),
  so this is the clean test stage 1 could not deliver — and S reads more AI.
- Guard rail 1 violated (38/42, four real dropped numbers); rails 2–3 held,
  rail 3 vacuous on English (detail-token regex matches every word; recorded
  as a dated note, not a criterion change).
- Combined verdict: nothing ships; next candidate H-4a.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ePnwo4nBWCWvFvpy1YgCN
@devswha
devswha merged commit 95e52ea into dev Sep 4, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant