Add ntt-complex.format_ast benchmark scaffold#1491
Draft
vmendelev wants to merge 1 commit into
Draft
Conversation
(cherry picked from commit 72d4fa0d36d0ae0bf57cda2fc4301400de6bc864) Signed-off-by: Valentin Mendelev <vmendelev@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add ntt-complex.format_ast benchmark scaffold
Summary
This PR adds the bootstrap
ntt-complexbenchmark group and its first subtest,ntt-complex.format_ast.format_astevaluates a complex2 task: audio speech translation plus strict output formatting. It renders existing FLEURS and CoVoST2 AST/ST examples into three required output formats:json_objectsrt_single_cuemarkdown_tableThe evaluator validates the requested structure, extracts the translated text, and scores the extracted translation with the standard audio translation BLEU path.
Implementation
Added:
nemo_skills/dataset/ntt-complex/prepare.pynemo_skills/dataset/ntt-complex/ntt_complex_eval.pynemo_skills/dataset/ntt-complex/ntt_complex_metrics.pynemo_skills/dataset/ntt-complex/README.mdtests/test_ntt_complex_format_ast.pyThe prepare script accepts source manifests from both naming layouts used across the current code and canary-dev references:
fleurs/st/test.jsonlorfleurs/ast/test.jsonlcovost2/st/test.jsonlorcovost2/ast/test.jsonlValidation
Local tests on branch
codex/ntt-complex-format-ast:Result:
3 passedDraco/IAD data preflight:
ntt-complex-format-ast-prepare-preflight-v4-2026061810226192/home/vmendelev/.cache/saferun/runner-packages/runner-cluster-configs/cluster_configs/v0.7.1/lustre/fsw/portfolios/llmservice/users/vmendelev/workspace/code/nemo-run/ntt-complex-eval/eval-format-ast/skills_data/ntt-complex/format_ast/test.jsonl1200fleurs=600,covost2=600json_object=400,markdown_table=400,srt_single_cue=400Runtime Evaluation Status
Target checkpoint:
The latest successful startup eval was:
ntt-complex-format-ast-piotr-step7200-eval-v25-open-asr-startup-2026061810228227/lustre/fs12/portfolios/llmservice/users/pzelasko/containers/open-asr-nemotron-omni.sqsh++max_samples=8exit_code=0format_valid=0/8;format_ast_is_correct=0/8The full eval completed successfully:
ntt-complex-format-ast-piotr-step7200-eval-v26-full-2026061810228387/lustre/fsw/portfolios/llmservice/users/vmendelev/workspace/code/nemo-run/ntt-complex-eval/eval-format-ast/evals/piotr_step7200_format_ast_v26_full/eval-results/ntt-complex.format_ast/output.jsonlformat_valid=0/1200;format_ast_is_correct=0/1200; average BLEU0.2322json_object0/400 valid (missing_json_object),markdown_table0/400 valid (too_few_markdown_rows),srt_single_cue0/400 valid (too_few_srt_lines)covost20/600 valid, average BLEU0.2450;fleurs0/600 valid, average BLEU0.2194latest/iad.yamlmust not be used for this work because it is an unfilled template. The active workflows use the renderedv0.7.1config override above.Hosted Nemotron Omni comparison:
ntt-complex-format-ast-hosted-nemotron-omni-api-probe-v2-2026061810228637nvidia/nvidia/nemotron-3-nano-omni-30b-a3b-reasoningerrors=0;format_valid=6/6;format_ast_is_correct=2/6; average BLEU0.1997ntt-complex-format-ast-hosted-nemotron-omni-api-full-v1-2026061810228702/lustre/fsw/portfolios/llmservice/users/vmendelev/workspace/code/nemo-run/ntt-complex-eval/eval-format-ast/evals/hosted_nemotron_omni_format_ast_full_v1/eval-results/ntt-complex.format_ast/output.jsonlerrors=0;format_valid=1124/1200;format_ast_is_correct=120/1200; average BLEU0.1118json_object392/400 valid and 44/400 correct;markdown_table332/400 valid and 33/400 correct;srt_single_cue400/400 valid and 43/400 correct.PR Handoff
Suggested base:
Suggested head:
Branch commit:
Browser PR URL:
Qwen Omni comparison status:
ntt-complex-format-ast-hosted-qwen-omni-model-probe-v1-20260618, Slurm app10228800Qwen/Qwen3-Omni-30B-A3B-Instruct,qwen/qwen3-omni-30b-a3b-instruct,Qwen/Qwen3-Omni-30B-A3B-Thinking,qwen/qwen3-omni-30b-a3b-thinkingkey_model_access_deniedwith the available runner key.Qwen/Qwen3-Omni-30B-A3B-Instructandvllm_omni.OmniLLM, but bounded Draco/IAD cache checks did not find the documented Granary chain script, a Qwen3-Omni HF snapshot, or a usablevllm_omniruntime/package.