CrashQueueLab is a deterministic fault-injection and invariant-testing laboratory for durable job queues. It makes worker crash windows observable, replays append-only transition history, and checks whether persisted queue state still satisfies its correctness model.
It is intentionally not another Celery, RQ, Dramatiq, or Huey clone. The SQLite queue is a small, inspectable system under test; the differentiator is the evidence produced around failures.
Run both core experiments and print one concise evidence summary. The output directory must be new or empty so results from separate runs cannot be mixed.
crashqueuelab showcase \
--output-directory reports/showcase \
--seed 20260827The command saves complete databases and histories while keeping terminal output focused on the result. Key fields are:
{
"effect_commit_before_acknowledgement": {
"attempts_per_scenario": {
"idempotent": 2,
"non_idempotent": 2
},
"durable_effects": {
"idempotent": 1,
"non_idempotent": 2
},
"invariants_ok": true,
"ok": true
},
"ok": true,
"process_fault_campaign": {
"crash_mode": "spawned_process_os_exit",
"crash_time_invariants_all_ok": true,
"crash_time_replay_all_ok": true,
"exit_codes": [86],
"failed": 0,
"named_fault_points": 13,
"passed": 13,
"recovery_all_succeeded": true,
"timeouts": 0
},
"queue_delivery_semantics": "at-least-once execution",
"seed": 20260827
}This demonstrates two different forms of evidence:
- Real child-process termination and restart recovery at all 13 named hooks.
- Two executions producing two unprotected effects but one idempotent effect.
The complete terminal object also includes final states and artifact paths.
Full reports are written below reports/showcase/; the terminal object is a
summary, not a replacement for the retained evidence.
A worker commits an external effect and then crashes before acknowledging queue success. The lease expires, recovery schedules another attempt, and a second worker executes the job again. CrashQueueLab runs that exact schedule twice:
- A non-idempotent destination records two durable effects.
- A destination protected by a unique effect key records one durable effect.
Both runs contain two executions. The result demonstrates an effectively-once side effect without claiming exactly-once execution.
flowchart LR
P["Producer"] -->|"submit"| Q["SQLite durable queue"]
Q -->|"leased claim"| W1["Worker A"]
W1 -->|"commit effect"| E["Separate effect ledger"]
E -. "crash before acknowledgement" .-> X["Process exits"]
Q -->|"lease expiry and recovery"| W2["Worker B"]
W2 -->|"repeat effect"| E
W2 -->|"acknowledge"| Q
Q --> H["Append-only history"]
H --> R["Replay and invariant checker"]
- Run the 60-second showcase above.
- Open
reports/showcase/process-fault-campaign/campaign.jsonand inspect any case directory containing its queue, effect ledger, and case report. - Inspect
worker.pyfor the effect-to-acknowledgement window andqueue.pyfor transaction and lease-fencing guards. - Compare the independent
replay.pymodel with the database-wideinvariants.pychecks. - Review the abrupt-exit and competing-worker tests in
test_process_and_concurrency.py.
CrashQueueLab requires Python 3.11 or newer and has no runtime dependencies.
git clone https://github.com/b2ty9t7yhz-source/crashqueuelab.git
cd crashqueuelab
python3 -m venv venv
source venv/bin/activate
python -m pip install -e .
crashqueuelab showcase \
--output-directory reports/showcase \
--seed 20260827Install the development checks only when reviewing the implementation:
python -m pip install -e ".[dev]"
python -m pytestRun only the effect-idempotency comparison:
crashqueuelab demo --output-directory reports/demoThe deterministic summary is written to reports/demo/comparison.json:
{
"fault_point": "AFTER_EFFECT_COMMIT",
"idempotent": {
"attempt_count": 2,
"effect_count": 1,
"final_state": "SUCCEEDED",
"invariants_ok": true
},
"non_idempotent": {
"attempt_count": 2,
"effect_count": 2,
"final_state": "SUCCEEDED",
"invariants_ok": true
},
"queue_delivery_semantics": "at-least-once execution"
}The full output includes two queue databases, two effect databases, both event histories, per-scenario JSON reports, and the compact comparison. Generated databases and reports are ignored by Git.
Run an exhaustive semantic fault campaign in a deterministic, seed-controlled
order. Each case starts a fresh child process and terminates it with os._exit
at the selected fault hook:
crashqueuelab campaign \
--output-directory reports/campaign \
--seed 20260827reports/campaign/campaign.json records the seed, execution order, crash-time
state, replay result, invariant result, and recovered terminal state for all 13
named fault points. The parent records the child exit code, reopens the queue
and effect databases, checks their durable state, and drives recovery to a
terminal result. Each isolated case retains both databases and a
machine-readable report. Reusing the seed reproduces the same case order and
identifiers.
The campaign is exhaustive over the documented semantic hooks, not over every possible operating-system, hardware, network, or instruction-level failure.
| Component | Responsibility | Primary implementation |
|---|---|---|
| Durable queue | Guarded submit, claim, heartbeat, success, retry, and recovery transitions | queue.py |
| SQLite storage | WAL mode, synchronous=FULL, bounded connections, and short BEGIN IMMEDIATE writes |
storage.py |
| Worker boundary | Keeps handler execution and external effects outside queue transactions | worker.py |
| Fault harness | Seeded child-process termination across 13 semantic hooks | campaign.py |
| Correctness evidence | Independent event replay and database-wide invariant checks | replay.py, invariants.py |
| Synthetic destination | Separate idempotent and non-idempotent effect ledgers | effects.py |
The queue row is the current materialized state. The append-only event stream is an independently replayable explanation of how that state was reached. Every state-changing transaction writes both or rolls back both.
An external effect and queue acknowledgement normally occur in separate transactions. A crash between them makes redelivery correct but may duplicate the effect. The project exposes this window instead of claiming exactly-once execution.
A worker may finish after its lease expires. Success, failure, and heartbeat updates therefore require the current worker ID, opaque lease-generation token, and an unexpired deadline. Expiry permits recovery; it does not cancel the old process.
Each queue mutation and exactly one history event share a short SQLite write transaction. Replay reconstructs the state independently, and the invariant checker rejects disagreement between history, attempts, and the materialized job row.
ManualClock, deterministic IDs, explicit hooks, fixed seeds, and process
barriers replace timing-dependent sleeps. The campaign still uses real
os._exit termination, so normal Python cleanup cannot make a crash case pass.
- SQLite keeps transactions and durable artifacts inspectable; it is not a claim of distributed scalability.
- Queue transactions never span handler execution or external effects; this preserves short write locks but requires at-least-once semantics and destination idempotency.
- Named semantic hooks make failures reproducible; they do not enumerate every possible scheduler interleaving or instruction-level failure.
ManualClockmakes lease and retry tests deterministic; production wall-clock behavior and cross-machine skew are outside V1.
- Durable SQLite job, attempt, and event storage in WAL mode
- Idempotent submission using canonical JSON and a payload SHA-256 digest
- Deterministic claim ordering and
BEGIN IMMEDIATEwrite transactions - Worker ID, opaque lease-generation token, heartbeat, expiry, and stale-worker fencing
- Capped exponential retry, maximum attempts, and dead-letter state
- Append-only transition history protected by SQLite triggers
- Independent history replay and database-wide invariant checking
- Thirteen named fault hooks across handler and transaction boundaries
- Seeded spawned-process campaign covering every named fault point exactly once
- Abrupt child-process exit and concurrent multi-process claim tests
- Separate idempotent and non-idempotent synthetic effect ledgers
- Typed Python API, structured JSON CLI errors, and deterministic reports
The five durable states and every legal guard are documented in the state machine.
CrashQueueLab provides at-least-once execution behavior. An expired, unacknowledged lease can be executed again; therefore duplicate execution is an expected outcome, not a hidden edge case.
The project deliberately distinguishes:
| Term | V1 position |
|---|---|
| At-most-once execution | Not used; it can lose work after an early acknowledgement |
| At-least-once execution | Implemented through leases, recovery, and retry |
| Exactly-once execution | Not claimed |
| Idempotent processing | Demonstrated with a stable unique effect key |
| Effectively-once side effect | Demonstrated even though execution occurs twice |
See Delivery semantics for the two-commit crash window and the same-database exception.
The JSON CLI exposes submission, claim, heartbeat, success, failure, retry promotion, expired-lease recovery, history, replay, and invariant checking.
crashqueuelab submit \
--database demo/queue.sqlite \
--idempotency-key request-1 \
--job-type synthetic.effect \
--payload '{"effect_key":"effect-1","value":7}'
crashqueuelab history --database demo/queue.sqlite --job-id JOB_ID
crashqueuelab replay --database demo/queue.sqlite --job-id JOB_ID
crashqueuelab check --database demo/queue.sqliteThe package also exposes immutable typed models and a direct API:
from crashqueuelab import DurableQueue, ManualClock
clock = ManualClock(1_000_000)
queue = DurableQueue("demo/queue.sqlite", clock=clock)
submitted = queue.submit(
idempotency_key="request-1",
job_type="synthetic.effect",
payload={"effect_key": "effect-1", "value": 7},
)
claimed = queue.claim(worker_id="worker-a", lease_duration_us=100_000)
assert claimed is not None
queue.succeed(
job_id=submitted.job.job_id,
worker_id="worker-a",
lease_token=claimed.lease_token or "",
)Handlers are trusted in-process callables. The project never serializes or executes arbitrary Python functions from the database.
Verification snapshot from 2026-08-28 on Python 3.13.12/macOS: 81 tests passed with 95.88% branch coverage. Ruff, strict mypy, isolated package build, and a clean-wheel showcase run also passed. The configured gates below define what CI must enforce rather than relying on this snapshot alone.
CI rejects changes unless all of these gates pass:
| Gate | Evidence enforced by CI |
|---|---|
| Compatibility | Full suite on Python 3.11, 3.12, and 3.13 |
| Fault correctness | Exception rollback tests and 13 spawned-process os._exit cases |
| Concurrency | Spawned workers race through a process barrier, not sleep |
| Coverage | Branch coverage must remain at or above 95% |
| Resource safety | Unclosed SQLite ResourceWarning values fail the suite |
| Static analysis | Ruff lint/format and strict mypy |
| Packaging | Wheel and source distribution build in isolation |
| Installation | A clean virtual environment installs and runs the built wheel |
Run the same checks locally:
python -m pytest --cov --cov-report=term-missing
python -m ruff check .
python -m ruff format --check .
python -m mypy
python -m build- Architecture
- Correctness model and invariants
- State machine
- Fault model and crash points
- Delivery semantics
- Design decisions and tradeoffs
- Testing strategy
- Contributing
V1 uses Python, SQLite, synthetic payloads, and a local filesystem. It does not use Redis, PostgreSQL, Kafka, Docker, Kubernetes, a web dashboard, untrusted jobs, or network filesystems.
The tests do not prove behavior under physical disk corruption, power loss, filesystem durability violations, network partitions, clock skew, or production-scale load. These limits are intentional and documented rather than hidden behind reliability claims.
CrashQueueLab is available under the MIT License.