docs(readme): strong-run benefit table replaces stale metric-v1 numbers (EN/JA)#162
Merged
Merged
Conversation
…efit table (EN/JA) The "Measured, not vibes" section still cited the metric-v1 first-run summary (mean 9.8 / 8.6 / 8.0 / 2.4), which metric v2 superseded. Both READMEs now show the strong-weight attack-scenario table (a_lock 92% maintenance / 86.1 mean / 100% lock resistance vs hand-written baseline 8% / 0%) with the honest caveats spelled out: one model/scenario pair, hand-written prompts stay competitive at moderate, and same-config repeats swing +/-20-40pts so single runs are not rankings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014NfK99PNs4u4kzyyxjurza
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
変更内容
README.md/README.ja.md/CHANGELOG.md)概要
README の「Measured, not vibes」節に metric v1 の旧数値(mean 9.8 / 8.6 / 8.0 / 2.4 — metric v2 で置き換え済みの計測)が残っていたため、v1.9.0 の strong 実測の効果表に差し替え(EN/JA 同期):
正直な注記も同梱: 1 モデル・1 シナリオ対 / moderate では手書きが互角 / 同条件再実行で ±20-40pt 振れる(単発順位は信用しない)。
検証
scripts/check_readme_counts.py: OK (346 属性、EN/JA 一致)scripts/build_site.py --check: OK関連 Issue
refs: docs/BENCHMARKS.md「Follow-up runs」節
🤖 Generated with Claude Code
https://claude.ai/code/session_014NfK99PNs4u4kzyyxjurza
Generated by Claude Code