Skip to content

feat: add random serving benchmark workload - #5036

Merged
qinxuye merged 5 commits into
xorbitsai:mainfrom
qinxuye:feat/benchmark-random-workload
Jun 18, 2026
Merged

qinxuye merged 5 commits into
xorbitsai:mainfrom
qinxuye:feat/benchmark-random-workload

Conversation

@qinxuye

@qinxuye qinxuye commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add --dataset-name random to benchmark_serving.py for vLLM-style synthetic request generation without ShareGPT.
  • Add random prompt generation helpers with configurable input/output lengths, range ratio, shared prefix length, and seed.
  • Thread ignore_eos through the benchmark runner so synthetic output lengths are more stable.
  • Update benchmark README examples and fix the --num-prompts flag spelling.

Validation

  • python -m black --check benchmark/utils.py benchmark/benchmark_runner.py benchmark/benchmark_serving.py
  • python -m isort --check-only benchmark/utils.py benchmark/benchmark_runner.py benchmark/benchmark_serving.py
  • python -m flake8 benchmark/utils.py benchmark/benchmark_runner.py benchmark/benchmark_serving.py
  • python -m py_compile benchmark/utils.py benchmark/benchmark_runner.py benchmark/benchmark_serving.py
  • python benchmark/benchmark_serving.py --help

@XprobeBot XprobeBot added this to the v2.x milestone Jun 15, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a synthetic prompt generation feature (similar to vLLM's random benchmark mode) to the serving benchmark tool, allowing users to generate random prompts with reproducible token lengths. It adds new command-line arguments, updates the benchmark runner to support ignoring EOS tokens, and implements random request sampling in the utility module. The review feedback highlights opportunities to prevent special tokens from being introduced during random prompt generation, add validation for input and output lengths, and safely access the API response's usage statistics to avoid potential KeyErrors.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread benchmark/utils.py
Comment thread benchmark/utils.py
Comment thread benchmark/utils.py
Comment thread benchmark/utils.py
Comment thread benchmark/benchmark_runner.py Outdated
@qinxuye
qinxuye force-pushed the feat/benchmark-random-workload branch from cb138d1 to 163b31f Compare June 16, 2026 17:35
@qinxuye

qinxuye commented Jun 17, 2026

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for generating synthetic, random benchmark requests (similar to vLLM's random benchmark mode) without requiring a dataset. It updates the documentation, adds new command-line arguments to benchmark_serving.py, and implements robust synthetic request generation in utils.py. Additionally, it introduces an --ignore-eos option and improves token usage tracking in benchmark_runner.py. The review feedback suggests enhancing error handling when parsing usage statistics from API responses and adding input validation to ensure args.request_rate is positive to prevent runtime errors.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread benchmark/benchmark_runner.py
Comment thread benchmark/benchmark_runner.py Outdated
Comment thread benchmark/benchmark_serving.py
@qinxuye

qinxuye commented Jun 17, 2026

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Overall the feature is well-implemented — clean separation between CLI parsing (benchmark_serving.py), benchmark orchestration (benchmark_runner.py), and prompt generation (utils.py). The ignore_eos threading and completion_tokens defensive handling are good improvements. The --num-prompts spelling fix in README is also appreciated.

Issues

  1. Major — _num_special_tokens_to_add returns 0 for non-callable attributes (see inline)
  2. Minor — request_rate validation happens after tokenizer loading (see inline)

Positive Observations

  • ignore_eos is cleanly threaded through all three constructor chains (BenchmarkRunner → ConcurrentBenchmarkRunner → ServingBenchmarkRunner)
  • completion_tokens fallback logic in send_request handles missing usage data gracefully in both streaming and non-streaming paths
  • parse_random_range_ratio provides good error messages for CLI misconfiguration
  • _gen_prompt_decode_to_target_len retry loop correctly handles tokenizer decode-re-encode mismatches
  • Reproducibility is preserved via np.random.default_rng(seed) local RNG for prompt generation alongside global np.random.seed for request timing

Validation

  • py_compile passes on all three changed Python files

Comment thread benchmark/utils.py
Comment thread benchmark/benchmark_serving.py

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review

All 9 prior findings (Gemini ×7, rogercloud ×2) are fully addressed. Three new minor issues found via inline comments below.

Comment thread benchmark/benchmark_runner.py Outdated
Comment thread benchmark/benchmark_runner.py Outdated
Comment thread benchmark/benchmark_serving.py Outdated
Comment thread benchmark/utils.py Outdated

@rogercloud rogercloud left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All three follow-up findings are addressed:

  1. benchmark_runner.py — logger.warning added for non-streaming missing-usage fallback. ✓
  2. benchmark_serving.py — _validate_random_range_ratio helper validates [0, 1) for both the scalar and JSON dict paths (also catches invalid input/output values individually). ✓
  3. utils.py — np.setdiff1d replaces the Python set subtraction. ✓

LGTM.

@qinxuye
qinxuye merged commit 7eef297 into xorbitsai:main Jun 18, 2026
8 of 13 checks passed
@qinxuye
qinxuye deleted the feat/benchmark-random-workload branch June 18, 2026 16:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants