Fix Linux batch-pool deadlock (force spawn start method) - #8
Merged
Merged
Conversation
The batch ProcessPoolExecutor used the platform-default multiprocessing start method, which is "fork" on Linux. A forked worker inherits the parent's already-initialized native thread pools (numba / OpenBLAS / OpenMP) and any CUDA context, which deadlocks the workers -- the test suite hung for hours on Linux CI while passing in ~90s on Windows (which already defaults to spawn). The thread-pinning in batch.sizing was already written assuming spawn semantics. Force the spawn context for the pool on every platform (the path Windows/macOS already exercise), so CPU and GPU batch pools start fresh, thread-pinned workers and never deadlock. Also add timeout-minutes: 20 to the CI test job so a future hang fails fast instead of running to GitHub's 6-hour limit. Verified: pytest 123 passed / 2 skipped, ruff + mypy clean (the spawn path is the one Windows already runs green). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes the Linux-only CI hang: the test suite ran for 3+ hours on
ubuntu-latestwhilethe identical suite passed in ~90 s on Windows. Root cause: the batch
ProcessPoolExecutorused the platform-default multiprocessing start method —
forkon Linux — and a forkedworker inherits the parent's already-initialized native thread pools (numba / OpenBLAS /
OpenMP) and CUDA context, which deadlocks. Windows already defaults to
spawn, which iswhy it passed.
This is also a real bug for users, not just CI:
batch_periodograms(device="cpu", workers>1)(and the GPU pool) could deadlock on Linux — the primary platform for astronomycompute clusters.
Fix
spawnstart method for the batch pool on every platform(
runner.py::_run_pool). It's the exact path Windows/macOS already exercise, so workersstart fresh and thread-pinned (the
batch.sizing.pin_worker_threadsdocstring alreadyassumed spawn). No fork → no deadlock; also correct for CUDA (a context can't be forked).
timeout-minutes: 20to the CI test job so a future hang fails fast instead ofrunning to GitHub's 6-hour limit.
Verification
pytest123 passed / 2 skipped, ruff + mypy clean. Windows CI already proves the spawnpool path; this PR's Linux CI run is the definitive test — it should now finish in
~2 min instead of hanging.
🤖 Generated with Claude Code