Situation
The current 34 benchmark scenarios were co-created with AI, resulting in questions that are too well-formed: no terminology drift, single-question-per-answer format. Real users ask more complex and unstructured questions.
An expert reviewer noted that questions accumulated in an internal Q&A service for Nablarch are increasingly of the type that require knowledge of design rationale to answer — making them a strong candidate for realistic benchmark material (see benchmark-review.md).
Pain
AI-generated scenarios alone cannot measure answer quality against the complex, unstructured questions real users actually ask. The benchmark risks being optimized for idealized inputs rather than real usage patterns.
Benefit
- Developers can objectively evaluate nabledge-6 quality against realistic user questions
- The benchmark functions as a meaningful indicator of practical quality
Success Criteria
🤖 Generated with Claude Code
Situation
The current 34 benchmark scenarios were co-created with AI, resulting in questions that are too well-formed: no terminology drift, single-question-per-answer format. Real users ask more complex and unstructured questions.
An expert reviewer noted that questions accumulated in an internal Q&A service for Nablarch are increasingly of the type that require knowledge of design rationale to answer — making them a strong candidate for realistic benchmark material (see benchmark-review.md).
Pain
AI-generated scenarios alone cannot measure answer quality against the complex, unstructured questions real users actually ask. The benchmark risks being optimized for idealized inputs rather than real usage patterns.
Benefit
Success Criteria
🤖 Generated with Claude Code