Skip to content

Fix Linux batch-pool deadlock (force spawn start method) - #8

Merged
tjayasinghe merged 1 commit into
mainfrom
fix/linux-batch-fork-deadlock
Jun 30, 2026
Merged

tjayasinghe merged 1 commit into
mainfrom
fix/linux-batch-fork-deadlock

Conversation

@tjayasinghe

Copy link
Copy Markdown
Owner

Summary

Fixes the Linux-only CI hang: the test suite ran for 3+ hours on ubuntu-latest while
the identical suite passed in ~90 s on Windows. Root cause: the batch ProcessPoolExecutor
used the platform-default multiprocessing start method — fork on Linux — and a forked
worker inherits the parent's already-initialized native thread pools (numba / OpenBLAS /
OpenMP) and CUDA context, which deadlocks. Windows already defaults to spawn, which is
why it passed.

This is also a real bug for users, not just CI: batch_periodograms(device="cpu", workers>1) (and the GPU pool) could deadlock on Linux — the primary platform for astronomy
compute clusters.

Fix

  • Force the spawn start method for the batch pool on every platform
    (runner.py::_run_pool). It's the exact path Windows/macOS already exercise, so workers
    start fresh and thread-pinned (the batch.sizing.pin_worker_threads docstring already
    assumed spawn). No fork → no deadlock; also correct for CUDA (a context can't be forked).
  • Add timeout-minutes: 20 to the CI test job so a future hang fails fast instead of
    running to GitHub's 6-hour limit.

Verification

pytest 123 passed / 2 skipped, ruff + mypy clean. Windows CI already proves the spawn
pool path; this PR's Linux CI run is the definitive test — it should now finish in
~2 min instead of hanging.

🤖 Generated with Claude Code

The batch ProcessPoolExecutor used the platform-default multiprocessing start method,
which is "fork" on Linux. A forked worker inherits the parent's already-initialized
native thread pools (numba / OpenBLAS / OpenMP) and any CUDA context, which deadlocks the
workers -- the test suite hung for hours on Linux CI while passing in ~90s on Windows
(which already defaults to spawn). The thread-pinning in batch.sizing was already written
assuming spawn semantics.

Force the spawn context for the pool on every platform (the path Windows/macOS already
exercise), so CPU and GPU batch pools start fresh, thread-pinned workers and never
deadlock. Also add timeout-minutes: 20 to the CI test job so a future hang fails fast
instead of running to GitHub's 6-hour limit.

Verified: pytest 123 passed / 2 skipped, ruff + mypy clean (the spawn path is the one
Windows already runs green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@tjayasinghe
tjayasinghe merged commit cec7bf0 into main Jun 30, 2026
6 checks passed
@tjayasinghe
tjayasinghe deleted the fix/linux-batch-fork-deadlock branch June 30, 2026 05:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant