Skip to content

Commit 6d597df

Browse files
momowayclaude
andauthored
[2.0] Add new Frontier-CS 2.0 problem vllm_llm_serving_optimization (#145)
* feat: Add new Frontier-CS 2.0 problem vllm_llm_serving_optimization * feat: Add real SWE-bench resolved fraction * feat(2.0/vllm): BFCL-memory workload + single-H100 contention + scoring/gate fixes Updates the vLLM serving-latency task with this iteration's work: - BFCL switched from AST function-calling to the multi-turn `memory` category (vendored kv backend + 5 pre-baked per-scenario snapshots); real, non-zero, non-ceilinged accuracy used as a guardrail. - Continuum job-FCFS reference.patch added (wins ~1.1-1.6x on SWE, single H100). - Single H100 (was H100:2): creates the KV contention scheduling needs; H100:2 was measured to over-provision (no queueing -> codex ~1.0x). - BFCL at jps=1.0 + down-weighted to 0.2 (SWE 0.8): measured to have no reproducible latency signal at any arrival rate (batch-numerics non-determinism swamps it), so it mainly serves as correctness gate + accuracy guardrail. - BFCL correctness gate uses a 5% abs-OR-rel tolerance to absorb that non-determinism instead of falsely failing good patches. - modal_app serves Qwen3-Coder-30B-A3B; dynamic serving_harness label; misc fixes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
1 parent ed19788 commit 6d597df

52 files changed

Lines changed: 6660 additions & 0 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎2.0/README.md‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -47,6 +47,21 @@ applies the submitted patch to a clean skeleton, runs a hidden arena against
4747
multiple baseline bot families, and scores by mean baseline win rate with a
4848
small faster-win tiebreak. The online generals.io service is not used.
4949

50+
## vLLM LLM-Serving Optimization
51+
52+
This systems problem asks agents to patch a clean upstream vLLM checkout to
53+
reduce the end-to-end latency of an LLM serving system on a multi-turn agentic
54+
workload, while keeping accuracy near a baseline. Its problem ID is
55+
`vllm_llm_serving_optimization`. The served model is
56+
`meta-llama/Llama-3.1-8B-Instruct` on a single Modal L40S, and the workload is a
57+
mini-swe-agent SWE-bench run. The agent submits a Python-only patch and can run
58+
an async public test (a subset of the final eval set) that returns real latency
59+
and accuracy feedback. Scoring is the geometric-mean latency speedup versus a
60+
vanilla-vLLM baseline, gated by an accuracy guardrail: accuracy within 5% of the
61+
baseline does not affect the score, and beyond that the score decays
62+
inverse-proportionally with the accuracy drop. Like duckdb-e2e, the agent and
63+
judge run in separate Docker environments.
64+
5065
## BBOPlace ISPD2005
5166

5267
This VLSI placement problem asks agents to generate macro placement candidates
Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,8 @@
1+
**/__pycache__
2+
**/*.pyc
3+
harbor/app/.public_test
4+
docs
5+
docker/README.md
6+
*.md
7+
reference.patch
8+
harbor/app/solution.patch
Lines changed: 124 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,124 @@
1+
# vllm_llm_serving_optimization — Design & Changes
2+
3+
Agent task: patch clean upstream **vLLM v0.11.0** (Python-only, allowlisted
4+
files) to cut end-to-end latency of an H100 Modal serve of
5+
`meta-llama/Llama-3.1-8B-Instruct`, preserving generation quality.
6+
7+
## What this revision adds
8+
9+
### 1. BFCL as a second judged workload (50/50 with SWE-bench)
10+
11+
`serving_eval/bfcl.py` runs the **Berkeley Function Calling Leaderboard**
12+
`simple` (Python) category as a serving workload: one chat completion per
13+
instance (prompt mode, no native tool-calling), client-side per-instance
14+
latency, and a deterministic **AST-equality** correctness check against a
15+
ground-truth call. The data slice is vendored under `serving_eval/bfcl_data/`
16+
(`BFCL_v4_simple_python.json` + `possible_answer/…`, from `bfcl-eval==2026.3.23`,
17+
Apache-2.0) so the judge runs **real, offline** evaluation with no network and no
18+
heavy `bfcl_eval` dependency.
19+
20+
Why BFCL: an 8B model resolves ~0% of SWE-bench Verified, so the old accuracy
21+
guardrail was mathematically dead (`baseline_accuracy = 0 ⇒ multiplier ≡ 1`).
22+
Llama-3.1-8B gets a **meaningfully non-zero** BFCL `simple` accuracy, so its
23+
guardrail is *live*.
24+
25+
The self-contained decoder + checker were **cross-validated** against
26+
`bfcl_eval`'s `ast_checker` on all 400 records: 395/400 agree, and on the 5
27+
nested dict/list cases ours is only ever *more* lenient (never stricter) — so it
28+
is symmetric and fair for baseline-vs-patched (and identical for both sides).
29+
Wrong-function and prose outputs are correctly scored incorrect.
30+
31+
Scoring (`scoring.py`, `evaluator.full_evaluation`):
32+
```
33+
final_score = swebench_weight * swe_score + bfcl_weight * bfcl_score # 0.5 / 0.5
34+
workload_score = clip(100*log2(geomean(per_instance_speedup)), 0, 100) * accuracy_multiplier
35+
```
36+
37+
### 2. Reference solution (`reference.patch`) — continuum job-level FCFS + long-prefill cap
38+
39+
Two-file, self-contained, correctness-preserving change (both files are in the
40+
strongly-allowed `vllm/v1/core/sched/**`):
41+
42+
- **`request_queue.py` — `JobFCFSRequestQueue` (the continuum soul).** The WAITING
43+
queue is ordered by each conversation's **job_id first-arrival time** instead of
44+
per-request arrival time, so a later turn of an in-flight conversation is
45+
admitted ahead of a brand-new job's first prefill — keeping ongoing multi-turn
46+
work moving and reusing its (already cache-hot) prefix. It is activated by
47+
default via `create_request_queue` (FCFS policy → `JobFCFSRequestQueue`); no
48+
launch-flag change is needed. Requests without a job_id fall back to plain
49+
per-request FCFS, so it is a safe drop-in.
50+
- **`scheduler.py` — long-prefill admission cap.** Caps fresh *uncached* long
51+
prefills per step (deferred via the existing skip-and-requeue when decode work
52+
is in flight), so a burst of new long prompts cannot head-of-line-block decode.
53+
54+
**job_id plumbing (no protocol/request.py changes).** The workload runners
55+
(`agent_runner.py`, `bfcl.py`) send a stable per-conversation id via
56+
`extra_body={"vllm_xargs": {"job_id": <instance_id>}}`. v0.11.0 already forwards
57+
`vllm_xargs` into `sampling_params.extra_args`, which vanilla vLLM ignores and the
58+
reference reads as `request.sampling_params.extra_args["job_id"]`. So the same
59+
requests serve identically on the baseline; only the patched scheduler uses the
60+
signal. Ordering uses `request.arrival_time` only (never wall-clock), so it is
61+
deterministic and changes only admission *order* — never tokens — and the greedy
62+
+ BFCL correctness gates pass.
63+
64+
This is materially more faithful to continuum than a client-signal-free version:
65+
continuum's headline is exactly job-level FCFS keyed on job_id (KV-pinning and the
66+
tool-call-length estimator are the parts it stubs/omits).
67+
68+
**Evidence it beats baseline.** This is the same mechanism a real codex trial
69+
agent used to measure **1.79× geomean speedup** on the 30-instance SWE-bench
70+
slice against this exact baseline (the reference diffs from blob `2b2cd63`, which
71+
matches the trial patch's base). The win comes from smoothing prefill bursts so
72+
each scheduler iteration keeps the running decode batch flowing (lower p50/p95
73+
inter-token latency) and from letting hot-prefix conversations resume without
74+
queueing behind a cold long prefill. On the blended metric the SWE-bench half
75+
improves strongly; the BFCL half (short single-turn prompts, no long prefills to
76+
defer) is roughly neutral, so the reference still scores clearly above the
77+
0-point baseline.
78+
79+
To re-validate live (needs Modal + HF creds and a free H100):
80+
```
81+
MODAL_TOKEN_ID=… MODAL_TOKEN_SECRET=… FRONTIER_SUBMISSION_ROLE=final \
82+
python3 evaluator.py reference.patch # judge path (baseline vs patched)
83+
# or, agent-side: bash harbor/app/public_test.sh run
84+
```
85+
86+
### 3. Audit fixes folded in
87+
88+
- **Live accuracy guardrail** via BFCL (above) — the headline correctness fix.
89+
- **Real, non-zero, always-runs correctness eval**: BFCL AST scoring needs no
90+
Docker or swebench harness, so it never silently degrades to a proxy and is
91+
never all-zero.
92+
- **BFCL per-sample correctness gate** (`measure._bfcl_correctness_ok`): a
93+
temperature-0 patch may not flip BFCL answers correct→wrong/undecodable
94+
(tolerates `bfcl_max_correctness_regressions` flips for batch-numerics noise).
95+
- **Anti-inflation scoring** (`scoring.paired_speedups`): per-instance speedup
96+
clamped to `[1/cap, cap]` (cap = 8); a patched instance that errored/early-exited
97+
is counted as a regression (`1/cap`), so "fail fast" can no longer inflate the
98+
geomean.
99+
- **Binary-hunk patch-policy bypass closed** (`evaluator.validate_patch`): patches
100+
containing `GIT binary patch` / `Binary files … differ` are rejected (the +line
101+
token scanner can't see a base85 payload; the build is Python-only anyway).
102+
- **Build-timeout mismatch fixed**: `evaluation.build_timeout_seconds` is now 7200,
103+
matching the documented budget and `environment.build_timeout_seconds`.
104+
105+
## Layout / build
106+
107+
The task source was reconstructed (`serving_eval/*.py` recovered from the judge
108+
image; `evaluator.py`/`config.yaml`/`readme`/`harbor/app` from the generated
109+
dataset). `docker/build_images.sh` does an **overlay rebuild** of the agent +
110+
judge images (`experimental-v0.11.1`), replacing `/opt/serving_eval` with the
111+
refreshed harness (incl. `bfcl_data/`) and asserting the BFCL data is present in
112+
the image. Both images carry `/opt/serving_eval`: the judge runs the
113+
authoritative measurement, the agent image runs the same harness for the public
114+
test.
115+
116+
## Known limitations
117+
118+
- The SWE-bench sandbox is `LocalSandbox` (empty dir) unless Docker-in-Docker is
119+
available on the judge; its accuracy stays a proxy. BFCL now carries the real
120+
task-quality guardrail, which is the point of the 50/50 split.
121+
- Latency is still a single sample per build on independently-autoscaled Modal
122+
serves; the cap + per-workload geomean + live guardrail reduce, but do not
123+
eliminate, run-to-run variance. Pinning `max_containers=1` and repeated
124+
sampling remain future hardening.
Lines changed: 139 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,139 @@
1+
tag: systems
2+
runtime:
3+
language: python
4+
timeout_seconds: 21600
5+
environment: "Patched vLLM (v0.11.0) source; Modal H100 GPU serving Qwen3-Coder-30B-A3B-Instruct; two agentic workloads (mini-swe-agent SWE-bench + BFCL memory) scored 50/50; latency-primary judge with a live BFCL accuracy guardrail + greedy/per-sample correctness gates"
6+
apt_packages:
7+
- bash
8+
- ca-certificates
9+
- curl
10+
- git
11+
- python3
12+
- python3-pip
13+
judge_apt_packages:
14+
- bash
15+
- ca-certificates
16+
- curl
17+
- git
18+
- python3
19+
- python3-pip
20+
judge_pip_packages:
21+
- modal
22+
- openai
23+
- datasets
24+
- huggingface-hub
25+
- rank-bm25
26+
docker:
27+
# Experimental local images. Build them with
28+
# 2.0/problems/vllm_llm_serving_optimization/docker/build_images.sh before running a
29+
# local Harbor trial. Both images need a clean upstream vLLM v0.11.0 checkout
30+
# (NOT the continuum fork). The judge image additionally vendors the
31+
# mini-swe-agent harness and the latency/accuracy scorer.
32+
image: frontiercs/vllm-serving-optimization-agent:experimental-v0.11.1
33+
judge_image: frontiercs/vllm-serving-optimization-judge:experimental-v0.11.1
34+
environment:
35+
cpus: 8
36+
memory_mb: 32768
37+
storage_mb: 65536
38+
build_timeout_seconds: 7200
39+
evaluation:
40+
# Model + accelerator served on Modal. Qwen3-Coder-30B-A3B (latest Qwen3 coder,
41+
# MoE: 30B total / ~3.3B active per token) actually resolves SWE-bench (the 8B
42+
# Llama was ~0) and is fast thanks to the sparse MoE. SINGLE H100 on purpose:
43+
# the 30B weights (~61GB bf16) leave only ~12-20GB KV (~5 concurrent 32K seqs),
44+
# which is exactly the contention this scheduling task needs. H100:2 was measured
45+
# to OVER-PROVISION (no queueing -> scheduling can't help -> codex ~1.0x); on one
46+
# card the continuum reference wins (SWE ~1.1-1.6x). vLLM TP size is derived from
47+
# "H100:N". Avoid B200 (sm_100) — the v0.11.0 precompiled wheel has no Blackwell kernels.
48+
model: Qwen/Qwen3-Coder-30B-A3B-Instruct
49+
gpu: "H100"
50+
# Workload 1 (latency-primary): mini-swe-agent on SWE-bench Lite (split test).
51+
# Lite (300, self-contained) is cheaper/cleaner than Verified; instances are
52+
# drawn by a fixed-seed RANDOM sample across all 12 repos (not a sorted prefix,
53+
# which clustered into 1-2 hard repos and was unrepresentative).
54+
dataset: princeton-nlp/SWE-bench_Lite
55+
dataset_split: test
56+
# Fixed seed for the deterministic random instance sample (SWE + BFCL), so the
57+
# evaluated set is reproducible run-to-run.
58+
sample_seed: 20260624
59+
# Iterative (agent-role) public test: a strict subset of the final eval set.
60+
public_slice: "0:5"
61+
# Final (verifier-role) evaluation: 50 instances sampled from Lite with
62+
# sample_seed (set "0:300" for the full Lite split; that is many H100-hours).
63+
eval_slice: "0:50"
64+
# Workload 2 (agentic + real correctness): BFCL **memory** category. Each
65+
# instance is a multi-turn agentic task — the model is given a key-value memory
66+
# tool suite pre-seeded (vendored snapshots) with facts from a prior
67+
# conversation, asked a question, and must issue retrieve/search tool calls
68+
# over several turns, then answer (word-boundary match vs ground truth). Gives a
69+
# genuine multi-step request path AND a non-zero, non-ceilinged accuracy signal
70+
# (a strong model gets ~70-90%, not a pinned 1.0), so the guardrail has range.
71+
bfcl_public_slice: "0:8"
72+
# Full memory verify set: all 155 vendored instances at final.
73+
bfcl_eval_slice: "0:155"
74+
bfcl_max_tokens: 768
75+
memory_max_steps: 20
76+
# BFCL memory arrives as a seeded Poisson process at its OWN rate (bfcl_jps).
77+
# Measured: BFCL's short, 93%-prefix-cacheable requests have NO clean scheduling
78+
# signal at any rate. Below KV saturation (jps<=1.0) the metric is reproducible
79+
# (~1.0x, identical run-to-run) but flat; at/above saturation (jps>=1.5) a real
80+
# ~1.4x job-FCFS effect appears but is swamped by batch-numerics non-determinism
81+
# (the SAME patch measured 0.71x and 1.55x across two clean jps=1.5 runs; an
82+
# identical-build control swung 10x at jps=2.5). So BFCL runs at jps=1.0 — the
83+
# one reproducible point — mainly for its correctness gate + accuracy guardrail,
84+
# and is DOWN-WEIGHTED in the latency blend (see *_weight below). SWE carries the
85+
# latency signal (its many-turn latency is robust to single-token flips).
86+
bfcl_jps: 1.0
87+
# bfcl_workers only caps the fallback burst path (arrival_mode != jps).
88+
bfcl_workers: 64
89+
# Final score = swebench_weight * SWE-bench score + bfcl_weight * BFCL score.
90+
# SWE-heavy: SWE carries the reliable latency signal; BFCL is down-weighted to
91+
# 0.2 because at its one reproducible load (jps=1.0) it is latency-neutral (~1.0x)
92+
# and mainly serves as the correctness gate + accuracy guardrail.
93+
swebench_weight: 0.8
94+
bfcl_weight: 0.2
95+
# Poisson arrival workload (jobs/second). Mirrors a realistic serving load.
96+
arrival_mode: jps
97+
jps: 0.5
98+
workers: 8
99+
step_limit: 50
100+
temperature: 0.0
101+
max_completion_tokens: 2048
102+
# Latency aggregation + scoring.
103+
latency_metric: mean_e2e_seconds
104+
# Per-instance speedup is clamped to [1/cap, cap]; a failed/early-exit patched
105+
# instance is counted as a regression, so "fail fast" cannot inflate the geomean.
106+
max_per_instance_speedup: 8.0
107+
# Accuracy guardrail (per workload). Within `accuracy_tolerance` relative drop
108+
# of baseline => no penalty; beyond it the score decays inverse-proportionally.
109+
# No penalty if the accuracy drop is within EITHER the relative tolerance OR
110+
# the absolute tolerance (resolve_rate over a finite slice is coarse/noisy).
111+
accuracy_tolerance: 0.05
112+
accuracy_abs_tolerance: 0.05
113+
agent_accuracy_mode: patch_validity
114+
final_accuracy_mode: resolve_rate
115+
# Greedy-output correctness smoke (fixed prompts must match the baseline
116+
# token-for-token at temperature 0 before timing is considered).
117+
correctness_smoke_prompts: 8
118+
# BFCL per-sample correctness gate: a temperature-0 patch should not flip BFCL
119+
# answers correct->wrong/undecodable. Allowed flips = max(the count floor below,
120+
# abs_tolerance * n_instances, rel_tolerance * n_baseline_correct) — a 5%/5%
121+
# abs-OR-rel band that absorbs batch-numerics non-determinism (which flips many
122+
# instances run-to-run even between identical builds under concurrency).
123+
bfcl_max_correctness_regressions: 1
124+
bfcl_correctness_abs_tolerance: 0.05
125+
bfcl_correctness_tolerance: 0.05
126+
# Modal serving knobs.
127+
modal_scaledown_seconds: 900
128+
modal_startup_timeout_seconds: 1200
129+
server_health_timeout_seconds: 1800
130+
# Per-phase wall-clock budgets (seconds). Matches the documented build budget.
131+
build_timeout_seconds: 7200
132+
instance_timeout_seconds: 1200
133+
# Use a baseline (vanilla vLLM) cached in the judge image when available,
134+
# otherwise the judge serves vanilla once and caches it for the trial.
135+
baseline_cache_path: /opt/vllm-baseline/baseline_metrics.json
136+
submission:
137+
kind: file
138+
path: /app/solution.patch
139+
max_queue_size: 2
Lines changed: 74 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,74 @@
1+
# Experimental vLLM Serving-Optimization Images
2+
3+
This task needs two images, mirroring the duckdb-e2e split: a public **agent**
4+
image and a private **judge** image. Both bundle a clean upstream vLLM checkout
5+
and the shared `serving_eval` harness. Build them before running a local Harbor
6+
trial:
7+
8+
```bash
9+
bash 2.0/problems/vllm_llm_serving_optimization/docker/build_images.sh
10+
```
11+
12+
Defaults:
13+
14+
```text
15+
VLLM_REF=v0.11.0
16+
AGENT_TAG=frontiercs/vllm-serving-optimization-agent:experimental-v0.11.0
17+
JUDGE_TAG=frontiercs/vllm-serving-optimization-judge:experimental-v0.11.0
18+
```
19+
20+
The agent image contains:
21+
22+
```text
23+
/app/vllm # clean upstream vLLM (no continuum, no reference fix)
24+
/opt/serving_eval # shared harness, used by the async public test
25+
/opt/vllm-baseline # optional precomputed baseline cache
26+
```
27+
28+
The judge image contains:
29+
30+
```text
31+
/opt/vllm-clean # clean upstream vLLM (build + baseline reference)
32+
/opt/serving_eval # shared harness, used by the evaluator
33+
/opt/vllm-baseline # baseline-metrics cache (filled on first measurement)
34+
```
35+
36+
## Runtime requirements (important)
37+
38+
Unlike duckdb-e2e, this task does **not** run the model inside the container.
39+
Both the agent public test and the judge serve `meta-llama/Llama-3.1-8B-Instruct`
40+
on a **Modal L40S** built from the (patched) vLLM source. The containers
41+
therefore need:
42+
43+
- `MODAL_TOKEN_ID` / `MODAL_TOKEN_SECRET` in the environment (Modal auth). Use a
44+
Modal service-user token for unattended runs.
45+
- A Modal Secret named `huggingface-secret` containing `HF_TOKEN` with access to
46+
the gated Llama-3.1 weights (`modal secret create huggingface-secret HF_TOKEN=...`).
47+
The container also reads `HF_TOKEN` for the `datasets` download.
48+
- The judge additionally needs a reachable Docker daemon (mounted socket or
49+
DinD) to run the SWE-bench per-instance testbeds for the final resolve-rate.
50+
When no daemon is reachable, the harness falls back to a local sandbox and the
51+
patch-validity accuracy proxy.
52+
53+
The Modal image build uses `VLLM_USE_PRECOMPILED=1`, so only vLLM's Python layer
54+
is rebuilt from the submitted source (minutes, not a full CUDA compile). This is
55+
why the patch policy is Python-only.
56+
57+
## Baseline cache
58+
59+
The judge measures the vanilla (clean-tree) baseline once per role and caches it
60+
at `/opt/vllm-baseline/baseline_metrics.json`, keyed by role (`agent` / `final`).
61+
Baseline and patched builds are never served simultaneously, so a single L40S is
62+
sufficient per environment. To precompute and bake the baseline into the image
63+
(recommended for faster trials), run the harness against the clean tree offline
64+
and copy the resulting `baseline_metrics.json` into the image at that path.
65+
66+
## Smoke test
67+
68+
```bash
69+
bash 2.0/problems/vllm_llm_serving_optimization/docker/smoke_images.sh
70+
```
71+
72+
This checks that the clean vLLM checkout, the `serving_eval` package, and the
73+
Modal/OpenAI/datasets (and, for the judge, swebench + docker) clients are
74+
importable. It does not exercise Modal or a GPU.

0 commit comments

Comments
 (0)