Skip to content

docs(readme): strong-run benefit table replaces stale metric-v1 numbers (EN/JA)#162

Merged
shiro-0x merged 1 commit into
mainfrom
claude/japanese-request-t2fvf0
Jul 12, 2026
Merged

docs(readme): strong-run benefit table replaces stale metric-v1 numbers (EN/JA)#162
shiro-0x merged 1 commit into
mainfrom
claude/japanese-request-t2fvf0

Conversation

@shiro-0x

Copy link
Copy Markdown
Owner

変更内容

  • 新規属性追加
  • 既存属性更新
  • schema 拡張
  • スクリプト修正
  • ドキュメント修正 (README.md / README.ja.md / CHANGELOG.md)
  • テスト追加・修正
  • その他

概要

README の「Measured, not vibes」節に metric v1 の旧数値(mean 9.8 / 8.6 / 8.0 / 2.4 — metric v2 で置き換え済みの計測)が残っていたため、v1.9.0 の strong 実測の効果表に差し替え(EN/JA 同期):

条件 維持率 平均 ロック耐性
blend + persona_lock 92% 86.1 100%
手書き 41tok ベースライン 8% 55.4 0%

正直な注記も同梱: 1 モデル・1 シナリオ対 / moderate では手書きが互角 / 同条件再実行で ±20-40pt 振れる(単発順位は信用しない)。

検証

  • scripts/check_readme_counts.py: OK (346 属性、EN/JA 一致)
  • scripts/build_site.py --check: OK
  • docs のみの変更(コード・テスト・属性 YAML 変更なし。直前の release_check で全ゲート緑確認済み)

関連 Issue

refs: docs/BENCHMARKS.md「Follow-up runs」節

🤖 Generated with Claude Code

https://claude.ai/code/session_014NfK99PNs4u4kzyyxjurza


Generated by Claude Code

…efit table (EN/JA)

The "Measured, not vibes" section still cited the metric-v1 first-run
summary (mean 9.8 / 8.6 / 8.0 / 2.4), which metric v2 superseded. Both
READMEs now show the strong-weight attack-scenario table (a_lock 92%
maintenance / 86.1 mean / 100% lock resistance vs hand-written baseline
8% / 0%) with the honest caveats spelled out: one model/scenario pair,
hand-written prompts stay competitive at moderate, and same-config
repeats swing +/-20-40pts so single runs are not rankings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014NfK99PNs4u4kzyyxjurza
@shiro-0x
shiro-0x merged commit 87adad1 into main Jul 12, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants