Skip to content

Add Kimi-K2-Thinking NVFP4/INT4 gsm8k accuracy coverage (4xB200)#40

Open
stecasta wants to merge 1 commit into
vllm-project:mainfrom
stecasta:stecasta/blackwell-kimi-thinking-b200
Open

Add Kimi-K2-Thinking NVFP4/INT4 gsm8k accuracy coverage (4xB200)#40
stecasta wants to merge 1 commit into
vllm-project:mainfrom
stecasta:stecasta/blackwell-kimi-thinking-b200

Conversation

@stecasta

Copy link
Copy Markdown

Adds Kimi-K2-Thinking accuracy recipes on 4xB200 (TP=4 + EP) to the nightly, extending Blackwell large-MoE coverage so quantized correctness regressions surface in nightly instead of downstream.

Reference gsm8k (vLLM CI gsm8k harness, 5-shot / 1319q):

  • Kimi-K2-Thinking NVFP4: 0.921
  • Kimi-K2-Thinking INT4: 0.914

Accuracy-only (lm_eval gsm8k, completions path). Validated locally (parse_workload + generate_pipeline); B200 nightly run to follow.

AI-assisted (Claude Code).

Adds 4xB200 (TP=4 + EP) accuracy-only recipes (lm_eval gsm8k) for
Kimi-K2-Thinking NVFP4 and INT4, extending Blackwell large-MoE coverage
so quantized correctness regressions surface in nightly instead of downstream.

Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant