Skip to content

Latest commit

 

History

History
141 lines (108 loc) · 6.71 KB

File metadata and controls

141 lines (108 loc) · 6.71 KB

Safety and correctness

Turning connection threads into fibers changes when handlers can interleave. That is the source of every correctness question here, and it is worth being precise about, because the honest answer is neither "always safe" nor "unsafe".

What changes

Under OS threads, a handler that blocks in an offload just stops. Under the runtime, it yields, and a peer handler runs on the same OS thread in the gap. Two consequences follow.

Thread-local state becomes shared

All fibers on a carrier share one OS thread, so they share its errno and its OpenSSL error queue. Handler A sets errno, yields at the offload, handler B overwrites it, A resumes and reads B's value.

This is not theoretical: measured at 100% corruption without mitigation, 0 with save/restore. The runtime saves and restores per-fiber errno and the OpenSSL ERR queue at every yield, and shadows pthread_getspecific / pthread_setspecific per fiber. Evidence: experiments/correctness/ (coop_err.c).

Read-modify-write across the offload stops being atomic

long v = shared;        /* read  */
accel_run(buf, n);      /* yields — a peer handler runs here */
shared = v + 1;         /* write — based on a stale read */

Between OS threads this is racy too, but the window is tiny and the thread blocks rather than handing the core to a peer that touches the same state. Under fibers the yield puts a peer in the gap every time. examples/servers/hostile_server.c does exactly this, and the test suite asserts the hazard is real before asserting that the mitigations work — a recent run lost 41,408 updates unprotected.

Three mitigations, in order of preference

1. The application already locks it

If the shared state is protected by a pthread_mutex, the runtime's fiber-aware mutex respects it: a fiber that finds the lock held parks and retries instead of deadlocking the carrier. A naive interposer would deadlock here — TOFFLOAD_REALMUTEX=1 exists to demonstrate that it does. Fiber-aware condition variables and reader-writer locks work the same way.

Measured: unlocked loses updates, LOCKED=1 gives 0 lost with no deadlock, naive mutex deadlocks. Tests: lock-respect, cond-rwlock.

2. The conflict detector finds it for you

No application cooperation needed. The detector write-protects the executable's writable data segments, snapshots a global version clock when a fiber parks at the offload, and treats a post-offload write fault to a page that changed during the park as a conflict.

Application Conflicts detected
conn_server (no cross-connection writes) 0 — full overlap preserved
hostile_server (unlocked RMW) 125,039 — hazard flagged automatically

Honest costs:

  • Overhead. Always-on page protection costs roughly 40% on safe applications: one mprotect per fiber resume. A production version would reprotect only dirtied pages.
  • Page granularity. State on the same page as unrelated state gets flagged too. In hostile_server, shared and total share a page, so a benign total++ counts as a conflict. Byte granularity needs a compiler pass.
  • Scope. Watches the executable's globals and BSS. Heap-shared state needs toffload_watch_region() from <toffload/toffload.h>.

3. Enforce mode serializes the conflict

TOFFLOAD_ENFORCE=1 takes a handler lock from the request read to the response write, so a conflicting handler's read → offload → write critical section runs alone. Measured: 0 lost updates, no application locks required.

The cost is exactly what it sounds like: conflicting handlers no longer overlap. On hostile_server throughput drops by roughly two orders of magnitude, because every request conflicts there. That is the correct trade for a truly hostile workload and a bad one for a workload that rarely conflicts, which is why the recommended pattern is two-phase: run with the detector, look at det_conflicts, and enable enforcement only if it is nonzero.

The adaptive single-run mode (enable enforcement on the first detected conflict) loses a small number of updates before it engages — 28 out of 28,866 in one run. Two-phase is sound; single-pass soundness would need checkpoint and rollback.

Where it is unsafe, plainly

Sub-libc blocking. A fiber that enters a raw futex or io_uring_enter inside a storage engine blocks the carrier and every other fiber on it. The runtime cannot yield through what it cannot interpose. MariaDB/InnoDB is the measured example; see bench/results/mariadb.md.

TOFFLOAD_RT=1 in production. It pins the carrier and runs it SCHED_FIFO. If a fiber waits on a mutex held by a real OS thread, the realtime carrier can starve that thread and you get priority-inversion deadlock. It reproduces the paper's microbenchmark setup and belongs nowhere near a production server.

Versioned symbol interposition. Interposing pthread_cond_* without matching glibc's GLIBC_2.3.2 symbol versions crashes MariaDB at startup (1 of 5 stock binaries tested). The runtime does this correctly with a linker version script and dlvsym. If you fork this code, keep that.

Thread-pool servers. A fixed worker pool created before the first accept() collapses onto the carrier under TOFFLOAD_POOL=1. Sometimes that is what you want; often it changes the server's concurrency model in ways it does not expect. Test it.

What the test suite actually checks

make check runs on any machine with no GPU and no root:

Test Asserts
results-match identical verified AES output with and without the runtime
fibers-engage the mechanism engaged rather than silently passing through
passthrough a non-offloading server stays correct under the runtime
cond-rwlock fiber-aware cond and rwlock complete without deadlock
hazard-real unlocked RMW across the offload really does lose updates
lock-respect an application mutex across the offload gives 0 lost
enforce enforce mode gives 0 lost with no application locks
stats-quiet no output on a host process's stderr unless asked
config-file, env-overrides-file knob precedence

Every load run is bounded by CLIENT_TIMEOUT (30 s), so a runtime stall fails the suite loudly instead of hanging it. That guard exists because a leaked handler lock in enforce mode did deadlock the runtime during development.

Reporting a correctness bug

A response that fails verification (bad > 0), a lost update where the application holds a lock, or a stall is a bug, not a tuning issue. Please open an issue with the output of TOFFLOAD_STATS=stderr:1 and the knobs you set. See SECURITY.md if the issue has security impact.