Repository navigation
EvalPort import/export for nemoguardrails.eval datasets and results? #2326
adhabnr-ux
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi maintainers — I'm the author of EvalPort, an open standard (Apache 2.0) for portable LLM evaluation datasets: a JSON spec for test suites, test cases, and result sets, plus Python/TypeScript SDKs (
evalport-sdk) that validate against real JSON Schemas. Posting here rather than as an issue since CONTRIBUTING.md asks that issues be opened by a human, and this is a "how might this fit" idea rather than a bug or a committed feature request — Discussions/Ideas seemed like the right spot.I went through
nemoguardrails/eval/models.pyandnemoguardrails/eval/cli.pybefore writing this, so this is grounded in your actual data model, not a guess:EvalConfig(policies: List[Policy]+interactions: List[InteractionSet]) is structurally very close to an EvalPort test suite: eachInteractionSet(id,inputs,expected_output: List[ExpectedOutput], tags) maps almost 1:1 onto an EvalPort test case, with yourRefusalOutput/SimilarMessageOutput/GenericOutputvariants mapping onto EvalPort's assertion types.EvalOutput(results: List[InteractionOutput], each withinput,output,compliance,compliance_checks) maps onto an EvalPort result set —InteractionOutput.complianceis essentially your per-policy pass/fail, which is exactly the shape EvalPort result records carry.Concretely, what I'm proposing is small and additive — not a refactor of the eval subsystem:
The payoff: a policy-compliance dataset built for
nemoguardrails eval runcould be exported once and reused (or diffed) against other eval tooling that already speaks EvalPort, and — going the other direction — someone could bring in an existing EvalPort test suite (e.g. a jailbreak/refusal benchmark) and run it straight throughcheck_compliancewithout hand-writing the YAML/JSON.Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
No pressure at all if this isn't a priority right now or doesn't fit where the eval module is headed — happy to just leave this here as a pointer, or to sketch a fuller PR if a maintainer thinks it's worth pursuing. Thanks for the very readable eval code either way, it made this easy to check for a real fit before posting.
All reactions