Skip to content

Add Blackwell NVFP4/FP8 gsm8k accuracy coverage (1xB200)#39

Open
stecasta wants to merge 1 commit into
vllm-project:mainfrom
stecasta:stecasta/blackwell-nvfp4-small-b200
Open

Add Blackwell NVFP4/FP8 gsm8k accuracy coverage (1xB200)#39
stecasta wants to merge 1 commit into
vllm-project:mainfrom
stecasta:stecasta/blackwell-nvfp4-small-b200

Conversation

@stecasta

Copy link
Copy Markdown

Adds 1xB200 gsm8k accuracy recipes to the nightly, part of the effort to catch Blackwell quantized-model correctness regressions in nightly instead of downstream.

Reference gsm8k (vLLM CI gsm8k harness, 5-shot / 1319q):

  • Gemma-4-31B-IT-NVFP4: 0.702
  • Llama-3.3-70B-Instruct-NVFP4: 0.911
  • Devstral-Small-2-24B (FP8): 0.834
  • Mistral-Small-4-119B-NVFP4: 0.905

Accuracy-only (lm_eval gsm8k, completions path). Validated locally (parse_workload + generate_pipeline); B200 nightly run to follow.

AI-assisted (Claude Code).

Adds 1xB200 accuracy-only recipes (lm_eval gsm8k) for Gemma-4-31B-IT-NVFP4,
Llama-3.3-70B-Instruct-NVFP4, Devstral-Small-2-24B (FP8), and
Mistral-Small-4-119B-NVFP4, to catch Blackwell quantized-model correctness
regressions in nightly instead of downstream.

Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant