Problem / motivation
Once bulk bytes use sendfile, protocol scheduling and shared userspace bookkeeping determine the next ceiling. Several operations occur around every resident hit and may constrain multi-ring scaling:
- Whole/paged sendfile hits emit
INFO logs (runtime.rs); worker and load-test defaults are RUST_LOG=info (main.rs, dataplane_loadtest.sh). At 100K+ requests/s this can dominate user work or contaminate benchmarks.
BlockIndex has one process-wide RwLock (index.rs).
- LRU touch/pin/unpin uses one process-wide
Mutex; eviction scans tracked units while holding it (eviction.rs).
- The version cache uses one
Mutex; headers, frame rejoin, and some byte paths allocate/copy per request.
These are hypotheses, not a reason to add sharding or lock-free complexity without profiles. Talon's previous request-affinity experiment regressed performance.
Proposed work
- Profile all-hit workloads with logging disabled, default logging, and realistic sinks.
- Move per-hit success events to
DEBUG/TRACE or sampling while retaining counters and slow/error logs.
- Measure lock wait/hold time, allocations, cycles/request, cache misses, context switches, and ring scaling.
- Optimize only measured bottlenecks. Candidates include deferred/approximate LRU touches, shorter eviction critical sections, sharded metadata, reusable frame buffers, and removal of payload-sized copies.
- Preserve fd pinning, eviction, version freshness, recovery, and cardinality bounds.
Acceptance criteria
Related: #285, #291, #303, #569
Problem / motivation
Once bulk bytes use
sendfile, protocol scheduling and shared userspace bookkeeping determine the next ceiling. Several operations occur around every resident hit and may constrain multi-ring scaling:INFOlogs (runtime.rs); worker and load-test defaults areRUST_LOG=info(main.rs,dataplane_loadtest.sh). At 100K+ requests/s this can dominate user work or contaminate benchmarks.BlockIndexhas one process-wideRwLock(index.rs).Mutex; eviction scans tracked units while holding it (eviction.rs).Mutex; headers, frame rejoin, and some byte paths allocate/copy per request.These are hypotheses, not a reason to add sharding or lock-free complexity without profiles. Talon's previous request-affinity experiment regressed performance.
Proposed work
DEBUG/TRACEor sampling while retaining counters and slow/error logs.Acceptance criteria
Related: #285, #291, #303, #569