Skip to content

As a developer, I want to incorporate real user questions into benchmark scenarios so that evaluation reflects actual usage #406

Description

@kiyotis

Situation

The current 34 benchmark scenarios were co-created with AI, resulting in questions that are too well-formed: no terminology drift, single-question-per-answer format. Real users ask more complex and unstructured questions.

An expert reviewer noted that questions accumulated in an internal Q&A service for Nablarch are increasingly of the type that require knowledge of design rationale to answer — making them a strong candidate for realistic benchmark material (see benchmark-review.md).

Pain

AI-generated scenarios alone cannot measure answer quality against the complex, unstructured questions real users actually ask. The benchmark risks being optimized for idealized inputs rather than real usage patterns.

Benefit

  • Developers can objectively evaluate nabledge-6 quality against realistic user questions
  • The benchmark functions as a meaningful indicator of practical quality

Success Criteria

  • The benchmark includes scenarios derived from real user questions, and nabledge-6 can be evaluated against them
  • Any quality gaps between AI-generated and real-user scenarios are visible in the benchmark results

🤖 Generated with Claude Code

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions