All notable changes to this project are documented here. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
2.8.1 - 2026-09-18
- An Agent Skill for coding agents. A model trained before 2026 has never
seen interlock, so an agent asked for a circuit breaker reaches for a
consecutive-failure counter and guesses at the API.
skills/interlock-cb/SKILL.mdfollows the open Agent Skills format and installs withnpx skills add bagowix/interlockinto Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI and the other clients that read it. The skill is a procedure: inventory the outbound calls, pick the integration per dependency, sizeConfigfrom observed traffic, roll out inMETRICS_ONLY, map rejections to503 + Retry-After, test with an injected clock, and migrate from pybreaker, circuitbreaker, aiobreaker or purgatory. The reference material stays in the docs, which the skill links. The install command sits in the README quickstart and on the docs landing page.
- The source distribution no longer carries the Hypothesis example
database. The release job runs the test suite before
uv build, and Hypothesis leaves its.hypothesis/cache in the checkout. Git ignores it through the nested.gitignoreHypothesis writes there, while hatchling reads only the root one, so the 2.8.0 sdist shipped 37 opaque cache files. The sdist target now excludes the directory explicitly.
2.8.0 - 2026-09-01
-
The open wait can now grow while probe rounds keep failing.
wait_duration_in_openwas a constant, so a breaker that could not recover retried at full rate indefinitely, each round hammering a dependency already in trouble.Config.wait_duration_backoff_multiplierlengthens the wait after each consecutive failed round andConfig.wait_duration_in_open_maxcaps it; a passing round resets both. The default multiplier of1.0keeps the historical constant wait, so nothing changes until it is raised. The growing interval doubles as a signal: a breaker that is merely waiting out a blip looks nothing like one that has failed ten rounds in a row. -
A Dependabot pull request no longer fails CI on the Codecov upload. A run triggered by Dependabot resolves
secrets.*against the separate Dependabot secret store, soCODECOV_TOKENhas to be maintained in two places — and a drift between them is invisible until the upload is rejected and the requiredCoveragecheck turns red on every open bump PR at once, with nothing wrong in any of the diffs. The coverage gate isfail_under = 100, whichpytestenforces in that same job before the upload runs, so the upload is now non-fatal on bot pull requests only; human pull requests and pushes tomainstill fail hard when Codecov rejects a report.
-
A backoff the coordinated lane cannot honour is now refused instead of ignored. Reopening a breaker with a shared
Storageis the backend's decision, taken fromwait_duration_in_openand its own clock; no failed-round count crosses the wire, becauseSharedStatecarries mechanism rather than policy and has no field for one. Await_duration_backoff_multiplierabove1.0alongside a storage was therefore read, validated and then silently dropped — the option looked enabled while every round waited exactly as long as the last. It now raisesValueErrorat construction, onCircuitBreaker,Registryand the per-breakerRegistry.get(config=...)override alike. -
A probe that never reached the dependency no longer decides the round. A
HALF_OPENprobe asks one question — has the dependency recovered? — and every failure was taken as its answer, including failures that never left the process: no free connection in the local pool, no bulkhead permit. A single such probe could re-open the breaker on evidence it did not have, and while the local cause persisted every round failed the same way, however healthy the dependency had become. Such a call is now recorded through the newunreachableflag onStateMachine.record, which hands the probe's slot back without a verdict instead of counting a failure the probe never observed. The httpx and httpx2 transports passPoolTimeoutthat way out of the box; other integrations take the set throughunreachable_exceptionsonCircuitBreaker,RegistryandEngine.CLOSEDis deliberately untouched — an exhausted pool is a real signal there, and shedding load is the point.
2.7.0 - 2026-08-18
-
A rejected request now lands in its client library's own exception hierarchy. Every HTTP integration raised a bare
CircuitOpenError, a type no handler written against httpx, aiohttp or requests catches. Application code degrades on the idiom its client teaches —except httpx.TransportError,except requests.exceptions.RequestException,except aiohttp.ClientError— so the day a breaker leftMETRICS_ONLYand started rejecting, the rejection flew past every one of those handlers at once and turned a partial outage into a full one. The rejection is nowCircuitOpenTransportError(httpx, httpx2),CircuitOpenClientError(aiohttp) orCircuitOpenRequestError(requests): each is both a native error of the host library and still aCircuitOpenError, carryingbreaker_name,retry_afterandlast_failureplus the library's own request context (.requeston httpx and httpx2,.request/.responseon requests).The host base is deliberately the broadest "the request never completed" type —
httpx.TransportError,aiohttp.ClientConnectionError,requests.exceptions.ConnectionError— and never a leaf such asConnectError,ReadTimeoutorSSLError. A leaf promises a physical cause that never happened, and leaves are exactly what retry predicates key on: inheriting from one would make every rejection retryable, and each retry would hit the same open circuit and burn an attempt for nothing. One choice puts the rejection inside every degradation handler and outside every typical retry predicate. -
The httpx and httpx2 transports retype the other two interlock errors as well.
CallTimeoutTransportError(anhttpx.TimeoutException) andBulkheadFullTransportError(anhttpx.PoolTimeout) surface a pipeline timeout or bulkhead rejection raised by a layer sitting inside the wrapped transport. Both describe transient local conditions, so — unlike a circuit rejection — they sit under httpx's timeout types on purpose, where retry predicates do fire on them. An error raised by the wrapped transport itself is never retyped, and neither is one that already carries a host hierarchy.
except httpx.TransportError— andexcept httpx2.TransportError,except aiohttp.ClientError,except requests.exceptions.RequestException— now catch circuit rejections. That is the point of the change, but it is a behavioural difference for code that already catches those types: a rejection reaches such a handler where it previously escaped to an outerexcept Exception. Nothing that catchesCircuitOpenErrorchanges, and no typical retry predicate starts firing — see the note on base types above.
2.6.1 - 2026-08-17
-
Documentation pages now carry Open Graph and Twitter card tags. A link to any page pasted into Slack, Reddit or a chat rendered as a bare url, because the theme emits no such tags and nothing supplied them. Each page now declares its title, description and canonical url. The card is the text-only
summarykind: Zensical ships no social-card generation and the documentation carries no social-card image. -
The documentation site is verified with Google Search Console. Every generated page carries the ownership tag for the property, so the maintainers can finally see which pages are indexed and which searches reach them — previously a blind spot. Nothing about the library itself changes. Removing the tag silently un-verifies the property.
-
The README opens with the state machine as a diagram. The three states and the condition on every edge had to be assembled from prose, so the shape of the thing being installed was the one thing the landing page never showed.
docs/img/state-machine.svgdraws CLOSED, OPEN and HALF_OPEN with their transitions — including the slow-call rate, the edge no other Python breaker has. One asset, no external fonts and no theme-dependent colours, so it renders the same on GitHub, on PyPI and in either colour scheme.
-
Every link in the PyPI sidebar now goes somewhere different.
HomepageandRepositoryboth pointed at the GitHub repository, so a reader arriving on the package page got two identically-targeted links and no obvious route to the documentation.Homepagenow points at the documentation site and the redundantDocumentationentry is gone; the repository stays reachable throughRepository. -
Page titles now describe the page to someone who has not arrived yet. Titles were written for a reader already inside the site —
Comparison,httpx,Timeout— which tells someone meeting the project in a search result or a pasted link nothing about what they are looking at. Aseo_titlesmap inzensical.tomlgives each page a self-describing title; a page left out of the map keeps the theme default. -
The README now demonstrates what this breaker does differently. Its only configured example set a failure rate and a minimum call count — the two knobs every other library has — while slow-call detection and the single sync/async class, the reasons to choose it, stayed bullet points with no code behind them. The quickstart now configures the slow-call thresholds, explains the dependency they catch (one that answers slowly and never raises, so a consecutive-failure counter never trips), and guards an async callable with the same instance. A new section shows Redis-backed shared state in five lines, and the rollout section dropped the paragraph that restated the states guide.
-
The documentation landing page was titled
interlock - interlock. Zensical does not populatepage.is_homepage, so the theme fell through to the generic "page title - site name" branch and duplicated the project name on the one page that is linked and ranked the most, leaving it without a single word describing what the project is. -
Every safe-rollout example now runs as written. The README and the
httpx,httpx2,aiohttpandrequestspages passedlistener=metrics_listener, a name defined nowhere on the page, so anyone copying the shadow-mode snippet — the one the documentation recommends starting with — got aNameErrorbefore reaching the breaker. They now pass the built-inLoggingEventListener(), which needs no setup and can be swapped for a metrics exporter. Theaiohttpandrequestsblocks also never imported the middleware and the adapter they construct, unlike every other block on those pages. -
Every link in the README now works from PyPI as well. The README is the package's long description, and PyPI resolves a relative link against
pypi.orgrather than the repository — so all 31 of them, the whole ofdocs/plusCONTRIBUTING.md,SECURITY.md, the licence and the examples, answered 404 for a reader who arrived on the package page, which is exactly where the newHomepagelink now sends people. Every link is absolute: documentation goes to the published site rather than to raw Markdown, which also spares a GitHub reader the tab syntax that only renders once built.tests/test_readme.pyfails the build on a relative link, on a documentation url with no page behind it, on a repository url with no file behind it, and on an anchor with no matching heading.
2.6.0 - 2026-08-12
CircuitBreaker.call_sync()andcall_async()protect a call without deciding what it is.call()inspects every callable it is handed — unwrappingfunctools.partial, probing__call__— and the decorator copies metadata onto a freshly allocated closure. A caller that already knows its own nature paid for a decision it could not change. The two new methods run the same protected path with the dispatch removed;call()keeps dispatching and stays the default.call_asyncawaits whatever the callable returns, so a callable that is not a coroutine function (a middleware handler, say) works too;call_syncnever awaits, so a coroutine function passed to it is recorded as an immediate success.Registrycan now be enumerated:names()anditems(). The registry could only be asked about a name you already knew, which is exactly the case that does not hold where it matters most: every HTTP integration creates its breakers lazily, one per host, so the set of names is only known at runtime. Listing them for a diagnostics endpoint during aMETRICS_ONLYrollout, or applying an operator override to all of them before a maintenance window, meant reaching into the privateregistry._breakers. Both methods take a point-in-time copy under the registry lock — a breaker created afterwards is not in it, and the returned tuple never changes.- The guarded transport is reachable without private attributes. The
httpx2/httpx transports kept the transport they wrap in
self._transport, so checking what a wrapper was actually built around — the pool limits, the TLS context, the proxy — meant reading privates through two libraries, and tests ended up asserting constructor kwargs instead of the object. Both the synchronous and asynchronous transports now expose a read-onlywrappedproperty alongsideregistry.
DISABLEDno longer claims to be a metrics no-op.State.DISABLEDanddisable()were documented as "admit all traffic, record nothing", but only the sliding window ever went quiet: every admitted call is still classified, timed and delivered to theEventListenerason_call. So an operator who reached fordisable()to silence a noisy exporter kept seeing its metrics, and one who switched a rollout fromMETRICS_ONLYtoDISABLEDhad no documented promise that listener-fed dashboards would survive it. The docstrings and the states guide now separate the two surfaces:DISABLEDrecords no outcome — thresholds are never evaluated andsnapshot()gets nothing new — and leaves listener events flowing. Behaviour is unchanged.- A bug in your own code no longer opens the circuit of a healthy
dependency. The HTTP integrations counted every exception raised inside the
guarded call as the dependency failing, including the ones the client library
raises for the caller's mistake: a scheme-less URL (
UnsupportedProtocol) or a local protocol violation (LocalProtocolError) in httpx2/httpx, a URL or proxy URL with no host (InvalidURL) in requests. A burst of them was enough to trip the breaker and start rejecting real traffic to a host that was answering fine — and inMETRICS_ONLYthey polluted the very baseline used to pick thresholds. Those exceptions now count as successes and still propagate unchanged.HttpStatusClassifier(excluded_exceptions=...)replaces the set: pass()for the old behaviour, or addPoolTimeout(kept a failure by default — an exhausted pool usually means the dependency is holding connections open) when the local pool is sized below your own burst. aiohttp excludes nothing by default, since it rejects malformed URLs before the middleware chain runs; the knob is there for the exceptions of middlewares of your own. The set is validated at construction — an entry that is not anExceptionsubclass raisesTypeErrorthere instead of crashing the classifier mid-incident, when the first real failure arrives.
- The HTTP integrations no longer rebuild their guarded wrapper on every
request. Each request through the httpx2/httpx transports, the requests
adapter and the aiohttp middleware re-ran the breaker's sync/async detection
and built a fresh
functools.wrapsclosure before the request could start — per-request overhead with no behavioural value, paid exactly when a service is busiest. They now call the breaker directly (call_sync/call_async), and the aiohttp middleware no longer needs its per-request coroutine wrapper. Against a stubbed transport that answers from memory, an httpx request through the wrapper falls from ~10.5 µs to ~7.5 µs (sync) and from ~9.9 µs to ~8.0 µs (async); ~3.5 µs of that is the stub itself, so the integration's own overhead drops by roughly 40%. Behaviour is unchanged. - The pipeline's breaker step stopped rebuilding its wrapper as well.
CircuitBreakerStrategydecorated the next layer on every call for the same reason, although it statically knows whether that layer is sync or async. It now routes throughcall_sync/call_async: the same protected path, one detection and one closure fewer per pipeline call. - A
Registrycache hit no longer takes the registry lock. Because the transport integrations create one breaker per host lazily, every request paid a lock acquisition to find a breaker that was already there. The lookup now reads the cache first and falls back to the locked, double-checked creation path only for a name it has not seen — one name still yields exactly one breaker, and a lost creation race returns the winner. FailureClassifier.is_failurenow declaresexception: Exception | None. The engine never passed aBaseException— cancellation and shutdown are released without being classified — so the wider annotation invited classifier authors to write cancellation handling that could never run. A classifier still annotatedBaseException | Nonekeeps type-checking.- A caller-owned
Registryneeds its own HTTP classifier. Handing one to the httpx2/httpx transports, the aiohttp middleware or the requests adapter moves the failure policy to the registry, so a returned503counts as a success unless the registry is built withclassifier=HttpStatusClassifier(). Theregistryargument now states that where the registry is configured. - The opposite trap is documented as well. A registry configured with
HttpStatusClassifierreads the status off every result it records, so borrowing a breaker from it for non-HTTP work raisedAttributeErroron the first returned value. The four shared-registry guides now say to keep such a registry for HTTP clients only. - Rejecting an option the injected registry already owns explains itself. The
ValueErrorused to name the conflicting options and stop; it now says the registry owns them and to configure them there.
2.5.0 - 2026-08-08
-
Windows and macOS regressions are now caught before release. A lightweight CI matrix exercises the platform-sensitive runtime paths on the oldest and newest supported Python versions while the full quality, coverage and Redis matrix stays on Ubuntu.
-
Transport integrations can key breakers by logical dependency instead of raw host. Service-discovery suffixes and shared gateway hosts previously made the httpx2/httpx transports, aiohttp middleware, and requests adapter merge unrelated dependencies or split one dependency across several breakers. Their new
name_resolver=callback receives the native request and supplies the single name used by the registry, open-circuit errors, and listener events. Host-based naming remains the default, while empty custom names and non-string results fail before any network I/O. -
Listener contracts now match the events a component actually emits.
CoreEventListener,StorageEventListenerandPipelineEventListenerlet a breaker-only, coordination-only or strategy-only sink pass strict type checking without inheriting unrelated hooks. The existingEventListenerremains the complete ten-hook contract and continues to satisfy every narrowed annotation, so existing listeners require no migration. Listener documentation now makes thenamenamespace explicit: breaker for core/storage events, strategy for pipeline events. -
Partial
EventListenerimplementations now pass strict type checking. A listener that only needs one hook previously had to stub all ten before mypy, pyright and pyrefly would accept it, despite runtime dispatch already treating every hook as optional.EventListenernow provides inherited no-op hooks, so subclasses override only what they observe and remain compatible when later releases add hooks; structurally complete listeners continue to work without inheritance. -
Transport integrations can share a caller-owned
Registryacross clients. Previously each httpx2/httpx transport, aiohttp middleware, and requests adapter built an isolated registry, so clients calling the same host accumulated separate windows and could disagree about dependency health. The newregistry=option gives them one breaker per host and one policy source, rejects conflicting construction options, and leaves teardown to the registry's owner while still closing each wrapped connection pool. -
Direct-call baseline and threaded-contention benchmarks, closing out the two gaps left in #84. The CodSpeed suite reported the absolute cost of every protected path but not the number a prospective user actually wants — the breaker's overhead relative to calling the function directly.
test_baseline_direct_calland its async twin make that ratio derivable from any run.benchmarks/test_contention.pydrives one breaker from four worker threads so the work under the shared lock is tracked as a trend; its docstring andCONTRIBUTING.mdspell out that instruction counting runs threads one at a time, so the number is lock-path work, not wall-clock contention.
- Lowered the
otelextra's floor fromopentelemetry-api>=1.43.0back to>=1.20.0. The higher floor was never a requirement ofOTelEventListener— it arrived as a side effect of a routine Dependabot bump that moved the dev pin and the extra's floor together.opentelemetry-distro/SDK releases pin the whole OTel stack to one API version, so the stale floor forced adopters on an older distro to choose between upgrading their entire OTel stack and droppinginterlock-cb[otel]— for a listener whose calls (get_meter,create_histogram,create_counter) have been stable across that range..github/workflows/ci.yml'sextras-minjob now pins and tests againstopentelemetry-api==1.20.0; the dev dependency group stays on the current version.
2.4.0 - 2026-08-06
-
Breakers can now start in a safe operator state before serving traffic.
CircuitBreakerandRegistryacceptinitial_state;CLOSED,FORCED_OPEN,DISABLEDandMETRICS_ONLYare valid, while transitionalOPEN/HALF_OPENfail fast. Lazy registry creation applies the state before publishing a breaker, so per-host HTTP integrations can deploy in shadow mode without a first-call race. The httpx2 and httpx transports, aiohttp middleware and requests adapter expose their registry for diagnostics and accept the same option. -
httpx now has a first-class transport integration (#82). Install
interlock-cb[httpx]and wraphttpx.HTTPTransportorhttpx.AsyncHTTPTransportto apply one breaker per request host without decorators. The sync and async wrappers preserve streaming responses, delegate the wrapped transport's full lifecycle (context entry and exit,close()/aclose()), reject host-less URLs before I/O, and use the same configurable429, 500, 502, 503, 504failure policy as the existing HTTP integrations. Unit and real-loopback tests run against both the minimum supported httpx 0.27.0 and the locked latest version in CI.
- The README now gets readers from installation to their first guarded call faster. The core API appears before project comparisons, the feature language is less brittle, and optional integrations are presented in one compact, linked overview instead of several partial examples.
-
The httpx2 transports now delegate context-manager entry and exit to the wrapped transport. Entering
with client:(orasync with client:) previously never entered the wrapped transport, so a custom transport that acquires resources in__enter__/__aenter__was not ready before its first request. Exit now also releases every per-host breaker, matchingclose()/aclose(). -
Closing an HTTP integration now also releases its per-host breakers. The httpx2 transports and requests adapter close both their native connection resources and the registry; the aiohttp middleware exposes
aclose()for application shutdown.
2.3.0 - 2026-08-05
- The coordinator's write queue is now bounded (#99). The queue feeding a
coordinated breaker's background lane had no upper bound: a lane that stopped
draining — a storage client blocking without a timeout, an async lane whose
event loop is gone — grew it for as long as the process lived. It now holds
at most
write_queue_sizewrites (a newRedisStorage/AsyncRedisStoragekeyword, default 128, read off the storage object like the other coordination knobs). Over capacity the arriving write is dropped rather than queued: it never blocks and never raises inside the protected path, and it is reported through the new optionalon_storage_write_droppedlistener hook (LoggingEventListenerlogs it atWARNING,OTelEventListenercounts it oninterlock.storage.events). Coordinated writes were already best-effort — a dropped one is reconciled by the next poll and bystate_ttl— and healthy lanes never reach the bound, since writes are per transition rather than per call. Shutdown keeps working with a full queue: one slot stays reserved for the stop sentinel.
2.2.0 - 2026-08-03
-
The Redis integration guide now documents the coordinated-mode contract (#101). Fencing (
expected_version), thestate_ttl-bounded probe-lease leak on an interrupted probe, the "tallies only whileHALF_OPEN" rule forrecord_probe, and the need for explicit teardown were previously only implicit in the code and internal notes.docs/integrations/redis.mdgains a "Coordinated mode contract" section plus a checklist for anyone implementingStorage/AsyncStorageagainst a backend other than Redis, and theStorage/AsyncStorageprotocol docstrings point at it. -
The degraded-storage retry policy is now configurable (#100). Every release before this one retried a downed
Storagebackend on a single fixed cadence (retry_backoff) — with many instances sharing that backend, they all retried in lockstep, so a recovering backend immediately faced a synchronised retry wave.RedisStorage/AsyncRedisStoragegain three new keyword-only knobs:retry_backoff_multipliergrows the delay geometrically with each further consecutive failure,retry_backoff_maxcaps it, andretry_jitterspreads it with a proportional random draw (deterministic under an injected clock, so it stays reproducible in tests). The attempt counter resets on recovery. Defaults (multiplier=1.0,jitter=0.0) exactly reproduce the fixed-delay behavior of earlier releases — nothing changes unless you opt in. See the Redis integration guide's Tuning section for the new knobs. -
CI now catches two release-only failure modes before they ship (#85). A
packaging-smokejob builds the wheel, runstwine checkandcheck-wheel-contentsagainst it, then installs it withpip-equivalent semantics (--no-deps, no dev group, no extras) into a clean Python 3.11 environment and importsinterlockthere. It assertsinterlock.__version__matches the installed distribution's metadata, thatinterlock/py.typedsurvived the build, and that every module underinterlock/integrations/made it into the wheel — an accidental import of an optional dependency from the core, or a packaging config that silently dropped a subpackage, would pass the dev-env test suite and only break for a plainpip install interlock-cb. Separately,release.ymlnow verifies the pushed tag is exactlyv{interlock.__version__}as its first step, beforeuv sync, tests, or the build run — the cheapest possible place to fail, and the tag reaches the shell throughenvrather than direct interpolation.
-
The model-based state-machine test now reaches probe rounds. The
RuleBasedStateMachinefrom #106 exists for the order-dependent half of the machine — the probe budget, generation fencing, out-of-order settles — and it was not exercising any of it: measured over 20 seeds, 18 finished a whole run without enteringHALF_OPENonce, and the best seed managed 7 entries. The rule mix was the cause. Four of nine rules were operator controls, three of them leading into an absorbing override state that onlyresetleaves — andresetwipes the window on the way out — while tripping a window took an uninterrupted run of2 x minimum_number_of_callssteps, since every call cost one step to admit and another to settle. The walk spent its examples bouncing between the overrides and CLOSED. Three changes fix it: acallrule that admits and settles in one step, the wayEngine.call_*actually works; the four operator controls collapsed into one sampled rule; and ateardownthat reports trips and probe rounds throughtarget()so hypothesis's search steers toward them.advancealso draws from a strategy that reaches the end of the open wait, which plainst.floats— biased toward0.0— almost never did. 17 of 20 seeds now reachHALF_OPENon roughly half of their examples, and every override state is still entered, so the #79 invariant keeps its coverage.max_examplesdrops from 200 to 120: the walk no longer needs the extra examples to stumble into a probe round, which keeps the file's runtime in the same range as before. -
The
Clockprotocol now documents thatmonotonic()must be non-negative.TimeBasedSlidingWindowreserves a negative second as its "never written" bucket sentinel, so a custom clock returning negative values would report a permanently empty window and the breaker would never trip. Every stdlib monotonic clock already satisfies this; the contract simply said "monotonically increasing" and left the rest implied.
2.1.4 - 2026-08-03
-
"Correctness and testing" docs page (
docs/correctness.md), linked from the docs nav and fromREADME.md.README.mdwas honest that interlock-cb is young and that pybreaker and circuitbreaker are proven, but left the strongest counter-argument unmade: the project already holds a bar most libraries in this space do not — 100% branch coverage, three strict type checkers, property- and model-based tests for the state machine, mutation testing on the state machine and engine, CI on free-threaded CPython, tested (not guessed) lower-bound dependency versions, and an auditable supply chain. The new page has one section per mechanism, each linking to the enforcing config, workflow or test file, plus a "known limits" section stating what is not covered. -
OpenSSF Best Practices badge, earned at the passing level (https://www.bestpractices.dev/projects/13932) and added to
README.md, closing out #107. -
Public-API breakage detection (
griffe check) on every pull request. v2.0 shipped without breaking changes and the standalone breaker surface stays untouched by the pipeline layer, but until now nothing mechanically verified either promise — mypy, pyright and pyrefly check that the code is internally consistent, not that its public surface is still compatible with the previous release, andtests/typing_surface.pyonly pins the shapes somebody remembered to write down..github/workflows/api-compatibility.ymldiffs the working tree against the latest release tag: removed or renamed objects, changed parameter kinds, order or defaults, narrowed return types. It coversinterlock/integrations/*too, since that surface is public even though it is not re-exported from__init__. A breaking change is not automatically wrong — a future major release will make one on purpose — so the job is failable but overridable: label the pull requestbreaking-changeto acknowledge it. The one finding it ignores is the release bump ofinterlock.VERSION, which griffe reports as a changed attribute value — that is the release mechanism working, and every release would otherwise have to be labelled a breaking change.griffeis a dev/CI-only dependency; the core stays at zero. -
Mutation testing (
mutmut) overinterlock/_state_machine.pyandinterlock/_engine.py— the two modules where a surviving mutant is a real bug. The 100% coverage gate proves every branch executes, not that anything would notice if it changed, and the first run made that concrete: 105 of 549 mutants survived a fully covered suite. Killing them added tests for what the suite was taking on trust — that probe accounting frees exactly one slot, that a probe round is judged on rates rather than counts, that an era counter is never reused, thatretry_afteris clamped and measured from the moment the breaker opened, that a coordinated rejection still names its breaker and its last failure, and that theauto_transitiontimer is armed and cancelled exactly at the two moments it should be and by nothing else. The score is now 526 of 549 (95.8%); the 23 survivors are equivalent mutants, enumerated by class inCONTRIBUTING.md. It runs weekly out of band (.github/workflows/mutation.yml), never as a pull-request gate, and against the deterministic suite only — a randomised test that reaches a branch some of the time makes the score irreproducible. -
The test suite now runs on free-threaded CPython (3.14t) in CI, alongside 3.11–3.14. interlock has always documented the breaker as thread-safe, but every interpreter in the matrix had a GIL, so nothing could falsify that claim. A new
tests/test_concurrency.pydrives a single breaker from many threads at once and asserts the four properties the lock exists to provide: window counts add up under concurrent recording,snapshot()never returns a torn view, the HALF_OPEN caps (max_concurrent_probes,permitted_calls_in_half_open) are never exceeded, and aRegistryhands out exactly one breaker per name. No races were found; this verifies an existing claim rather than making a new one. The package is pure Python, so one wheel already serves both build flavours and the only distribution-level signal left is metadata: it now declaresProgramming Language :: Python :: Free Threading :: 3 - Stable. -
The state machine is now also tested as a model (
tests/test_state_machine_model.py): a hypothesisRuleBasedStateMachinegenerates the sequence of steps — interleaved outcomes, clock advances, admissions, probe releases and operator overrides — rather than replaying a hand-written one, and checks the machine against an independently predicted state, generation, window aggregates and probe budget after every step. The existing properties intests/test_state_machine_properties.pytarget documented boundaries directly and stay; this one looks for the transition orders nobody thought to write down (an override during a probe round, a probe settling an era late). No counterexample was found against the current implementation. The one sequence it did shrink pointed at the model instead: a failed probe returns toOPENwithout clearing the probe counters, which is stale bookkeeping rather than a leak — nothing reads them outsideHALF_OPEN, and entering it starts the round over. The budget invariant is now scoped toHALF_OPEN, where the contract defines it, and the sequence is pinned as a regression test. -
CircuitBreaker.close()/aclose()andRegistry.close_all()/aclose_all()— a deterministic way to release a breaker's background work. Until now a coordinated breaker's lane (a daemon thread for a sync storage, an asyncio task for an async one) and theauto_transitiontimer ended only when the object was garbage collected, so a service could not flush queued shared writes, could not know when its threads were gone, and an async lane outlivingasyncio.run()produced "task was destroyed but it is pending".close()drains the queued writes in order, wakes the lane instead of waiting outpoll_interval, joins it, and cancels the timer without arming another. It is idempotent, safe to call from any thread, and terminal: the lane never restarts, and afterwards the breaker keeps protecting calls on local state while shared writes are dropped — the same behaviour as a degraded storage. The cached shared view is dropped with the lane, since nothing would refresh it and a peer'sOPENwould otherwise never expire. This is teardown, not a state change:close()does not close the circuit (reset()does). The weakref-based collection path stays as the safety net for abandoned breakers. -
pyrefly joins mypy and pyright as a third strict type checker in CI and in the pre-commit hooks. Three independent implementations of the same type system disagree in the corners, and the corners are exactly where a signature-preserving decorator lives; a check the other two miss should fail before a release, not in a user's editor. The whole package is clean under its
strictpreset.missing-override-decoratoris the one error kind turned off:typing.overrideis 3.12+ and the core carries notyping_extensionsdependency to backport it.
- Coverage reporting moved to Codecov.
pytest --covwritescoverage.xmlplus a JUnit report and CI uploads both, so a PR now shows the per-file coverage diff and the failing tests themselves rather than a single total. The 100% gate is unchanged and still enforced by pytest (fail_underinpyproject.toml); the Codecov statuses mirror it. This replacespy-cov-action/python-coverage-comment-actionand the badge branch it maintained — the CI job no longer needscontents: writeorpull-requests: write.
- A coordinator lane that exited while an op was in flight left the work queue
permanently unfinished: the op was dequeued before the weak reference was
resolved, so a lane stopping on a collected coordinator never called
task_done()and a join on that queue could never return. Both lanes now release the dequeued op on the drop path. - An
EventListenerthat raises can no longer damage the breaker it observes. Hooks were invoked directly at every call site, so a bug in a listener — a metrics exporter, a logging handler, a custom sink — could replace a successful protected result with an observability exception or mask the dependency's own error. In coordinated (storage-backed) mode the damage was worse and silent: a raisingon_state_changepropagated out of the background lane's poll tick and terminated the lane for good, after which the breaker never refreshed the shared view or flushed queued writes again; and the same exception raised during a queued write was reported to the application as a storage failure throughon_storage_degraded. Every hook now goes through one dispatcher: anExceptionis logged to theinterlocklogger atERRORwith its traceback and then ignored, whileBaseException— cancellation, shutdown — still propagates untouched. Hooks are dispatched by name and only if defined, so a listener may implement just the ones it needs. User-supplied policy callbacks keep their previous behaviour and still raise: aFailureClassifier, a pipeline fallback function, a tenacitybefore_sleephook.
-
The CI supply chain is now auditable from the outside. Workflows are the most privileged code in the repository and were the least checked part of it: every
uses:resolved to a mutable tag, so a compromised action could have changed what a release publishes without a single commit here. Three changes close that, and they only work together — a Scorecard grade over unpinned actions would have been a badge that says less than it looks like it does:- every action is pinned to a full commit SHA with the version in a trailing
comment (Dependabot updates both, and
zizmor --fix=allwrites the pin for a newly added action). Pinning surfaced two actions sitting a major behind their upstream, so they were bumped at the same time:astral-sh/setup-uv7 → 9 andcodecov/codecov-action5 → 7; - zizmor audits
.github/workflows/on every pull request and locally through the pre-commit hook. CI hands it the job's own token so the four audits that resolve a pin against its upstream repository —impostor-commit,ref-confusion,known-vulnerable-actions,stale-action-refs— actually run; without one they are silently skipped and a SHA is only as trustworthy as the person who typed it. Its findings are fixed rather than muted:actions/checkoutno longer leaves the job's credentials in.git/config(persist-credentials: false),docs.ymlgrantspages: writeandid-token: writeto the deploy job instead of to the whole workflow, and the release build no longer restores a dependency cache that a pull-request run could have written. The one suppression, with its reason, lives in.github/zizmor.yml; - OpenSSF Scorecard
runs weekly and on every push to
main, publishing a per-check score to the code-scanning dashboard and to a README badge.
Nothing about the release flow changed: it still builds once and publishes that artefact through PyPI's OIDC trusted publisher. The PEP 740 attestations it already produced are now requested explicitly rather than inherited from the action's default, and
SECURITY.mddocuments where to fetch and verify them. - every action is pinned to a full commit SHA with the version in a trailing
comment (Dependabot updates both, and
- Migration-guide links now use the anchors generated by Zensical, and the docs workflow runs in strict mode so broken links fail CI.
2.1.3 - 2026-07-31
CircuitBreaker.snapshot()now reads the sliding window under the engine lock, so concurrent call settlement cannot expose mixed counters from a partially completed update. The count-based snapshot path remains O(1).- Local manual controls now take precedence over shared state for coordinated
breakers.
force_open()rejects locally, whiledisable()andmetrics_only()admit locally without consuming a sharedHALF_OPENprobe.reset()clears only local control and metrics, then resumes the cached shared state rather than resetting the cluster.
2.1.2 - 2026-07-14
- The httpx transports now reject a request whose URL carries no host with an
eager
ValueError, matching the aiohttp and requests integrations. Previously such a request silently created (and shared) a breaker keyed on the empty string. - Overlapping
with breaker:/async with breaker:blocks on one breaker instance no longer mix up each other's timing. The guarded-block bookkeeping used a single instance-level stack, so when blocks from different threads or interleaved asyncio tasks exited out of LIFO order, each exit settled with the other block's start time and admission — corrupting durations (and so slow-call classification) and probe attribution. The stack now lives in aContextVar: every thread and every asyncio task keeps its own, and nested blocks on one breaker keep working. The decorator andcall()surfaces were never affected. - A
HALF_OPENprobe interrupted by aBaseException— anasyncio.CancelledError(client disconnect, task cancellation, a timeout composed outside the breaker),KeyboardInterrupt, or any other non-Exception— now returns its probe slot instead of leaking it. Previously each such interruption permanently shrank theHALF_OPENbudget; once every slot had leaked the breaker wedged inHALF_OPEN, rejecting all traffic until a manualreset(). The interruption is still never recorded as an outcome (the v1 invariant stands: cancellation says nothing about the dependency). In coordinated (storage-backed) mode an interrupted leased probe is not returned to the shared budget — the storage protocol has no un-lease operation — and the backend TTL bounds that leak.
2.1.1 - 2026-07-14
CircuitBreaker.call()andPipeline.call()are now overloaded on the callable's sync/async nature, so strict type checkers infer the exact result type at the call site. Previously the plainAwaitable[R] | Rreturn type madeawait breaker.call(async_fn)an error under both mypy--strictand pyright strict, and left sync results as a union needing a cast. Runtime behaviour is unchanged. The user-facing typing surface is now locked by static assertions (tests/typing_surface.py) checked by both type checkers in CI.
2.1.0 - 2026-07-11
- Litestar integration via the
litestarextra (interlock.integrations.litestar, requires Litestar ≥ 2.23):breaker_dependency(name, *, registry)— aProvidefactory injecting a shared breaker (annotate handlers withNamedDependency[CircuitBreaker]) — andcircuit_open_handler, mappingCircuitOpenErrorto503 Service Unavailablewith aRetry-Afterheader. examples/pipeline.py: a third runnable demo — timeout + breaker + fallback composed around a dependency that hangs instead of erroring; deterministic output, walked through on the demo page.
- PyPI metadata now mentions the resilience pipeline (package description,
keywords,
Framework :: AsyncIOclassifier). - Docs navigation restructured: Comparison moved to the top-level block, the API reference got its own Reference section (both used to render as children of Integrations), and integration pages are ordered by demand — FastAPI and Litestar first.
- Security policy updated for the 2.x line.
2.0.0 - 2026-07-10
The resilience-pipeline milestone: interlock grows from one pattern into a composable resilience framework, while the breaker stays a breaker.
Backwards compatibility: there are no breaking changes. The entire v1
public API — CircuitBreaker, Registry, Config, the timeout primitives,
every integration — is untouched; the whole v1 test suite passes unmodified.
The major version marks the scope of what is added, not a migration burden.
- Resilience pipeline core (
interlock.pipeline): aStrategyprotocol (sync + async in one class, mirroring the v1 breaker contract), aPipelineexecutor applying strategies in declaration order (first = outermost, Polly semantics), and adapters for the existing primitives —CircuitBreakerStrategy(wraps a standaloneCircuitBreakerunchanged) andTimeoutStrategy(bounds every attempt viatimeout/sync_timeout).BaseExceptionpasses through every layer unswallowed; the standalone breaker API is untouched. RetryStrategy(interlock.integrations.tenacity, requires thetenacityextra): a bounded retry layer for the pipeline delegating all policy to tenacity — attempts always capped, the original exception re-raised when the budget runs out,CircuitOpenErrornot retried by default (retry_unless_open), patient mode viawait_probe, sync/async sleep injectable,before_sleephook passed through. Importing the module without tenacity installed now raises an error pointing at the extra.BulkheadStrategy(interlock.pipeline): caps concurrent calls per runtime (athreading.Semaphorefor sync, anasyncio.Semaphorefor async, one config for both). With no free slot it rejects immediately by default or waits up tomax_waitseconds, raising the newBulkheadFullError(exported frominterlock) — deliberately distinct fromCircuitOpenError: local saturation is not dependency failure.FallbackStrategy(interlock.pipeline): substitutes an explicit fallback value for selected failures only — thefallbackcallable receives the caught exception,onacceptsExceptionsubclasses exclusively (so cancellation always propagates), and the strategy's result is typed as the honest unionT | F, notAny. Works outermost overCircuitOpenError/BulkheadFullError/CallTimeoutError, and never masks shadow-mode (metrics_only) statistics.- Pipeline DSL:
Pipelineis now usable as a signature-preserving decorator (@pipeline,ParamSpec-typed like the breaker's), andPipeline.builder()assembles strategies step by step —.fallback(...),.retry(...)(lazy tenacity import),.circuit_breaker(...),.bulkhead(...),.timeout(...),.add(custom),.build(). The pipeline surface (Pipeline,PipelineBuilder,Strategyand the four shipped strategies) is re-exported frominterlock. There is deliberately no context manager: awithblock cannot be re-run, so retry inside it is semantically impossible. - Pipeline observability:
EventListenergains three optional hooks —on_retry(name, attempt, delay),on_bulkhead_rejected(name)andon_fallback(name, error)— dispatched via safegetattr(the v1.2 pattern), so pre-2.0 listeners keep working unchanged.RetryStrategy,BulkheadStrategyandFallbackStrategy(and the matching builder steps) acceptname=andlistener=.LoggingEventListenerlogs retries at INFO and bulkhead rejections / fallbacks at WARNING;OTelEventListenercounts all three in a newinterlock.pipeline.eventscounter. - New docs: a resilience pipeline guide (strategies, recommended ordering with rationale, builder DSL, migration from v1, custom strategies); pipeline sections in the API reference, retries guide, observability guide, tenacity integration page, README and the comparison page (fallback is now shipped).
- Docs: a comparison page — interlock-cb vs pybreaker, circuitbreaker, aiobreaker and purgatory (feature table, honest trade-offs).
- Runnable examples (
examples/):lifecycle.pywalks one breaker through CLOSED → OPEN → HALF_OPEN → CLOSED around a flaky gateway;two_clients.pyshows two independently guarded clients in one asyncio loop — one dependency fails and falls back while the other keeps serving. Zero dependencies, no network, deterministic output; kept green by a CI smoke test and explained line by line on the new demo docs page.
- Docs: integration page titles no longer repeat the word "integration" under the Integrations nav section (e.g. "httpx2 integration" → "httpx2").
1.3.0 - 2026-07-08
- tenacity integration via the
tenacityextra (interlock.integrations.tenacity):retry_unless_open(*transient)— a retry predicate that retries transient exceptions but stops as soon as the circuit opens — andwait_probe(fallback, *, jitter=0.1)— a wait strategy that sleeps exactlyCircuitOpenError.retry_after(plus jitter) after a rejection and delegates to the fallback strategy otherwise. - aiohttp integration via the
aiohttpextra (interlock.integrations.aiohttp, requires aiohttp ≥ 3.12):CircuitBreakerMiddleware— a client middleware applying one breaker per request host. - requests integration via the
requestsextra (interlock.integrations.requests):CircuitBreakerAdapter— anHTTPAdapterforsession.mount(...)applying one breaker per request host. HttpStatusClassifier(httpx2, aiohttp, requests variants) now acceptsfailure_statusesto override the canonical retryable set (429, 500, 502, 503, 504).- New docs: integrations overview, "Retries and circuit breakers" guide, and recipes for LLM SDKs (OpenAI/Anthropic) and Flask/Django 503 handlers.
- Integration modules moved into the
interlock.integrationssubpackage:interlock.integrations.httpx2,interlock.integrations.otel,interlock.integrations.fastapi,interlock.integrations.redis. The old top-level import paths (interlock.httpx2,interlock.otel,interlock.fastapi,interlock.redis) are removed. Update imports accordingly; extras names and all public classes are unchanged.
1.2.0 - 2026-07-07
- Distributed shared state via the
redisextra (interlock.redis):RedisStorage(sync) andAsyncRedisStorage(async) coordinate breaker state across processes and machines through one Redis hash per breaker. Every transition runs as a Lua script (atomic across racing instances, version-fenced against stale decisions), elapse checks use the Redis server'sTIME, and keys carry a TTL so abandoned state self-expires. Works against Redis 5.0+, Valkey, or any RESP-compatible server. CircuitBreakerandRegistryaccept an optionalstorage(Storage/AsyncStorage). Without one, behaviour is unchanged and purely local. With one, a shared OPEN gates admission on every instance, and HALF_OPEN recovery probes are budgeted globally (permitted_calls_in_half_openin total across the fleet) via an atomic probe lease — the single inline storage operation on the protected path; everything else is a locally cached view refreshed by a background poller plus fire-and-forget writes. A coordinated breaker matches its storage's runtime: a sync storage serves the sync API, an async storage the async one; mixing styles raisesInterlockError.- Graceful degradation: a storage failure never reaches the protected
path. The breaker falls back to its local state, leaves the backend alone
for
retry_backoffseconds, and resynchronises (shared view authoritative again) once the backend recovers. EventListenergainson_storage_degraded/on_storage_recovered, implemented byLoggingEventListener(WARNING/INFO) andOTelEventListener(newinterlock.storage.eventscounter). The engine dispatches the two new hooks only if present, so existing listeners keep working unchanged.- Reworked
Storageprotocol (plus newAsyncStorage) as atomic intent operations —read,trip_open,begin_half_open_if_elapsed,lease_probe,record_probe,close— with new public DTOsSharedStateandProbeLease. The previousStorageshape (load/save) was declared but never consumed by the engine; this release gives it its first functional form.
- Outcomes are now recorded into the state-machine era that admitted them: a call admitted in CLOSED can no longer settle as a HALF_OPEN probe, and a probe settling after a close or reset no longer pollutes the fresh window.
reset()clears the remembered last failure, so aCircuitOpenErrorraised after a reset no longer reports a pre-reset exception.
1.1.0 - 2026-06-28
sync_timeout(seconds)decorator: a synchronous counterpart totimeout. It runs the wrapped callable in a daemon worker thread joined with a deadline and raisesCallTimeoutErroron overrun. Documents the worker-thread limitation: Python cannot kill a thread, so the worker keeps running in the background after a timeout.Config.auto_transition(defaultFalse): opt into a timer that proactively moves a breakerOPEN → HALF_OPENoncewait_duration_in_openelapses, emitting the state change without waiting for the next call. The lazy transition stays authoritative; the timer admits no probe and is cancelled onreset(),force_open(), or when a call makes the move first.- FastAPI integration via the
fastapiextra (interlock.fastapi):breaker_dependency(name, *, registry)injects a shared breaker withDepends, andinstall_exception_handler(app)mapsCircuitOpenErrorto503 Service Unavailablewith aRetry-Afterheader.
1.0.0 - 2026-06-27
- Core state machine:
CLOSED/OPEN/HALF_OPENplus the operator overridesFORCED_OPEN,DISABLEDandMETRICS_ONLY(shadow mode). - Sliding windows behind a
SlidingWindowprotocol, with count-based and time-based implementations selected viaConfig.window_type. - Failure-rate trigger with
failure_rate_thresholdandminimum_number_of_calls, and slow-call detection viaslow_call_duration_thresholdandslow_call_rate_threshold. - Lazy
OPEN → HALF_OPENtransition with a probe limit and a concurrency cap. - Single public
CircuitBreakerfor sync and async, usable as a decorator, a sync/async context manager, andbreaker.call(fn, ...). Decorators preserve the signature and sync/async nature viaParamSpec+@overload. - Manual control:
reset(),force_open(),disable(),metrics_only(). Registryof named breakers with a shared default config and per-name overrides.- Immutable
Config(frozen dataclass) with eager validation. FailureClassifierprotocol with a default policy (any raised exception is a failure); classification by result is supported by custom classifiers.CircuitOpenErrorcarrying the breaker name, an estimatedretry_after, and the last recorded failure.- Async-first
timeoutprimitive that turns a hang intoCallTimeoutError. - Observability:
EventListenerprotocol, a zero-dependencyLoggingEventListener, and anOTelEventListener(extrainterlock-cb[otel]). - httpx2 transport integration (extra
interlock-cb[httpx2]):CircuitBreakerTransportandAsyncCircuitBreakerTransportapply a breaker per host, with anHttpStatusClassifiertreating429, 500, 502, 503, 504and transport exceptions as failures. InterlockDeprecationWarning(subclassesUserWarning, visible by default).py.typed; strict mypy and pyright; 100% test coverage.