Skip to content

fix: Treat zombie-only process groups as terminated in E2E teardown - #393

Merged
jiangkuaixue123 merged 2 commits into
vllm-project:mainfrom
ms-llmd:bug/terminate-with-zoombie-process
Sep 24, 2026
Merged

jiangkuaixue123 merged 2 commits into
vllm-project:mainfrom
ms-llmd:bug/terminate-with-zoombie-process

Conversation

@ronenkat

Copy link
Copy Markdown
Contributor

Purpose

E2E teardown can fail with process group <pgid> still alive after SIGKILL even
though vLLM shut down cleanly and GSM8K passed. When the container's PID 1 does
not reap orphans (e.g. exec pytest), orphaned vLLM workers stay zombies, and
os.killpg(pgid, 0) keeps reporting their group as alive. This PR makes the
teardown liveness check ignore zombies.

Issue

Scope

  • In scope: tests/e2e/process_utils.py: new process_group_is_alive(), used by
    both liveness checks in terminate_process_groups; unit tests.
  • Out of scope:
    • The container image. A reaping PID 1 (tini, shareProcessNamespace) also
      avoids the failure, but a pod command: can bypass an image ENTRYPOINT,
      so the fix lives in the harness instead.
    • vLLM's shutdown path, where parents exit before reaping their children.

Implementation Notes

  • process_group_is_alive() keeps os.killpg(pgid, 0) as a fast path, then
    scans /proc/*/stat for a group member whose state is not Z/X.
    State and pgrp are parsed after the last ), because comm can contain ).
  • It returns "gone" only when it finds group members and all of them are
    zombies. If /proc is missing (non-Linux) or shows no member (exited
    mid-scan), it keeps the killpg answer and re-polls.
  • D-state processes still count as alive, so real driver hangs are still
    reported.
  • A new proc_root argument, defaulting to /proc, makes the check testable.

Test Plan

  • Unit: pytest tests/unit/test_e2e_process_utils.py tests/unit/test_e2e_runner.py.
  • E2E: run test_deepseek_v2_lite[afd-graph-2a2f] with pytest as PID 1 (no
    tini, no shareProcessNamespace), 20s timeout, on the node where it failed.

Test Result

  • Unit (Linux, vllm/vllm-openai:v0.26.0): tests/unit/test_e2e_process_utils.py
    and tests/unit/test_e2e_runner.py: 153 passed, 0 skipped. Two new tests
    cover a zombie-only group (no failure) and a zombie beside a D-state member
    (still reported).
  • mypy (repo hook, Python 3.10, Linux): no issues in the changed files.
  • E2E (4x H100, same node and scenario, pytest as PID 1, 20s timeout):
Harness Result
main FAILED: still alive after SIGKILL, also at 120s
this PR PASSED, GSM8K 0.2857; zombies VLLM::Worker_DP with ppid=1 in the checked group were observed and ignored

Docs Impact

  • Files updated: none
  • If none, reason: test-harness internal behavior; no user-facing change.

Essential PR Checklist
  • Purpose is clear and linked to public context when possible.
  • Scope is bounded.
  • Compatibility with vLLM v0.26.0 is considered.
  • No changes are made to the vLLM source checkout.
  • Plugin-owned classes or explicit dotted class paths are preferred over monkey patches.
  • Any compat shim or monkey patch is isolated, idempotent, version-guarded, documented, and tested.
  • Imports remain CPU-safe; CUDA-heavy work is delayed or GPU-gated.
  • Validation evidence is included, including skipped GPU tests when applicable.
  • Documentation impact is stated.

🤖 Generated with Claude Code

os.killpg(pgid, 0) succeeds while any group member is a zombie. When a
container's PID 1 does not reap orphans (e.g. exec pytest), vLLM workers
whose parent exited first stay zombies with ppid 1, so teardown reports
'process group still alive after SIGKILL' even though every real process
has exited and no timeout can help.

Add process_group_is_alive(), which scans /proc for a non-zombie member
of the group and falls back to the killpg answer when /proc is missing
or shows no member, and use it for both liveness checks in
terminate_process_groups.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: ronenkat <16743404+ronenkat@users.noreply.github.com>
@ronenkat ronenkat changed the title Treat zombie-only process groups as terminated in E2E teardown [fix] Treat zombie-only process groups as terminated in E2E teardown Sep 23, 2026
@ronenkat ronenkat changed the title [fix] Treat zombie-only process groups as terminated in E2E teardown fix: Treat zombie-only process groups as terminated in E2E teardown Sep 23, 2026
Comment thread tests/e2e/process_utils.py Outdated
if not process_entry.name.isdecimal():
continue
try:
stat = (process_entry / "stat").read_text()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Parse /proc/*/stat as bytes, or treat decoding failures as an inconclusive scan. Linux permits a non-UTF-8 comm value in any process, including one outside the target group. read_text() then raises UnicodeDecodeError, which this except OSError does not catch. That propagates through both liveness checks and can skip the later SIGKILL/reap phase, failing the E2E teardown. This edge case is valid but has not been reproduced in CI.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, fixed.

  • process_group_is_alive now reads /proc//stat with read_bytes() and parses it as bytes (b")", b"Z", b"X"), matching find_processes_matching_environment. A non-UTF-8 process name can no longer raise an exception.
  • If the process group field can't be parsed, int() raises ValueError, and the docstring now says so. Both liveness checks in terminate_process_groups catch (OSError, ValueError). A parse error is reported as a failure and the group is treated as alive, so SIGKILL and reaping still run.
  • New unit tests cover a non-UTF-8 comm outside the group (ignored), a live member with a non-UTF-8 comm (group still counted as alive), and an unparsable process group (reported, then escalated). All three fail on the previous code.

Signed-off-by: ronenkat <16743404+ronenkat@users.noreply.github.com>
@jiangkuaixue123 jiangkuaixue123 added the ready Used to trigger ready CI in PRs. label Sep 24, 2026
@jiangkuaixue123
jiangkuaixue123 merged commit 6e778e8 into vllm-project:main Sep 24, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready Used to trigger ready CI in PRs.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: E2E teardown fails with "still alive after SIGKILL" because zombie vLLM processes count as alive

2 participants