Turning connection threads into fibers changes when handlers can interleave. That is the source of every correctness question here, and it is worth being precise about, because the honest answer is neither "always safe" nor "unsafe".
Under OS threads, a handler that blocks in an offload just stops. Under the runtime, it yields, and a peer handler runs on the same OS thread in the gap. Two consequences follow.
All fibers on a carrier share one OS thread, so they share its errno and its
OpenSSL error queue. Handler A sets errno, yields at the offload, handler B
overwrites it, A resumes and reads B's value.
This is not theoretical: measured at 100% corruption without mitigation,
0 with save/restore. The runtime saves and restores per-fiber errno and
the OpenSSL ERR queue at every yield, and shadows pthread_getspecific /
pthread_setspecific per fiber. Evidence:
experiments/correctness/ (coop_err.c).
long v = shared; /* read */
accel_run(buf, n); /* yields — a peer handler runs here */
shared = v + 1; /* write — based on a stale read */Between OS threads this is racy too, but the window is tiny and the thread
blocks rather than handing the core to a peer that touches the same state.
Under fibers the yield puts a peer in the gap every time. examples/servers/hostile_server.c
does exactly this, and the test suite asserts the hazard is real before
asserting that the mitigations work — a recent run lost 41,408 updates
unprotected.
If the shared state is protected by a pthread_mutex, the runtime's
fiber-aware mutex respects it: a fiber that finds the lock held parks and
retries instead of deadlocking the carrier. A naive interposer would deadlock
here — TOFFLOAD_REALMUTEX=1 exists to demonstrate that it does. Fiber-aware
condition variables and reader-writer locks work the same way.
Measured: unlocked loses updates, LOCKED=1 gives 0 lost with no deadlock,
naive mutex deadlocks. Tests: lock-respect, cond-rwlock.
No application cooperation needed. The detector write-protects the executable's writable data segments, snapshots a global version clock when a fiber parks at the offload, and treats a post-offload write fault to a page that changed during the park as a conflict.
| Application | Conflicts detected |
|---|---|
conn_server (no cross-connection writes) |
0 — full overlap preserved |
hostile_server (unlocked RMW) |
125,039 — hazard flagged automatically |
Honest costs:
- Overhead. Always-on page protection costs roughly 40% on safe
applications: one
mprotectper fiber resume. A production version would reprotect only dirtied pages. - Page granularity. State on the same page as unrelated state gets flagged
too. In
hostile_server,sharedandtotalshare a page, so a benigntotal++counts as a conflict. Byte granularity needs a compiler pass. - Scope. Watches the executable's globals and BSS. Heap-shared state needs
toffload_watch_region()from<toffload/toffload.h>.
TOFFLOAD_ENFORCE=1 takes a handler lock from the request read to the response
write, so a conflicting handler's read → offload → write critical section runs
alone. Measured: 0 lost updates, no application locks required.
The cost is exactly what it sounds like: conflicting handlers no longer overlap.
On hostile_server throughput drops by roughly two orders of magnitude, because
every request conflicts there. That is the correct trade for a truly hostile
workload and a bad one for a workload that rarely conflicts, which is why the
recommended pattern is two-phase: run with the detector, look at
det_conflicts, and enable enforcement only if it is nonzero.
The adaptive single-run mode (enable enforcement on the first detected conflict) loses a small number of updates before it engages — 28 out of 28,866 in one run. Two-phase is sound; single-pass soundness would need checkpoint and rollback.
Sub-libc blocking. A fiber that enters a raw futex or io_uring_enter
inside a storage engine blocks the carrier and every other fiber on it. The
runtime cannot yield through what it cannot interpose. MariaDB/InnoDB is the
measured example; see bench/results/mariadb.md.
TOFFLOAD_RT=1 in production. It pins the carrier and runs it SCHED_FIFO.
If a fiber waits on a mutex held by a real OS thread, the realtime carrier can
starve that thread and you get priority-inversion deadlock. It reproduces the
paper's microbenchmark setup and belongs nowhere near a production server.
Versioned symbol interposition. Interposing pthread_cond_* without
matching glibc's GLIBC_2.3.2 symbol versions crashes MariaDB at startup
(1 of 5 stock binaries tested). The runtime does this correctly with a linker
version script and dlvsym. If you fork this code, keep that.
Thread-pool servers. A fixed worker pool created before the first
accept() collapses onto the carrier under TOFFLOAD_POOL=1. Sometimes that
is what you want; often it changes the server's concurrency model in ways it
does not expect. Test it.
make check runs on any machine with no GPU and no root:
| Test | Asserts |
|---|---|
results-match |
identical verified AES output with and without the runtime |
fibers-engage |
the mechanism engaged rather than silently passing through |
passthrough |
a non-offloading server stays correct under the runtime |
cond-rwlock |
fiber-aware cond and rwlock complete without deadlock |
hazard-real |
unlocked RMW across the offload really does lose updates |
lock-respect |
an application mutex across the offload gives 0 lost |
enforce |
enforce mode gives 0 lost with no application locks |
stats-quiet |
no output on a host process's stderr unless asked |
config-file, env-overrides-file |
knob precedence |
Every load run is bounded by CLIENT_TIMEOUT (30 s), so a runtime stall fails
the suite loudly instead of hanging it. That guard exists because a leaked
handler lock in enforce mode did deadlock the runtime during development.
A response that fails verification (bad > 0), a lost update where the
application holds a lock, or a stall is a bug, not a tuning issue. Please open
an issue with the output of TOFFLOAD_STATS=stderr:1 and the knobs you set.
See SECURITY.md if the issue has security impact.