Skip to content

perf(worker): profile and reduce cache-hit logging, allocation, and lock overhead #575

Description

@somfornot

Problem / motivation

Once bulk bytes use sendfile, protocol scheduling and shared userspace bookkeeping determine the next ceiling. Several operations occur around every resident hit and may constrain multi-ring scaling:

  • Whole/paged sendfile hits emit INFO logs (runtime.rs); worker and load-test defaults are RUST_LOG=info (main.rs, dataplane_loadtest.sh). At 100K+ requests/s this can dominate user work or contaminate benchmarks.
  • BlockIndex has one process-wide RwLock (index.rs).
  • LRU touch/pin/unpin uses one process-wide Mutex; eviction scans tracked units while holding it (eviction.rs).
  • The version cache uses one Mutex; headers, frame rejoin, and some byte paths allocate/copy per request.

These are hypotheses, not a reason to add sharding or lock-free complexity without profiles. Talon's previous request-affinity experiment regressed performance.

Proposed work

  1. Profile all-hit workloads with logging disabled, default logging, and realistic sinks.
  2. Move per-hit success events to DEBUG/TRACE or sampling while retaining counters and slow/error logs.
  3. Measure lock wait/hold time, allocations, cycles/request, cache misses, context switches, and ring scaling.
  4. Optimize only measured bottlenecks. Candidates include deferred/approximate LRU touches, shorter eviction critical sections, sharded metadata, reusable frame buffers, and removal of payload-sized copies.
  5. Preserve fd pinning, eviction, version freshness, recovery, and cardinality bounds.

Acceptance criteria

  • A reproducible profile attributes CPU/allocation/lock cost to logging, BlockIndex, LRU, version cache, framing, and metrics.
  • Default INFO logging no longer emits one event per successful cache hit.
  • Benchmark scripts set and record logging configuration explicitly.
  • Any data-structure rewrite has contention tests and preserves eviction/pinning/recovery semantics.
  • Before/after results cover 1/physical/logical ring counts and report throughput, user/system CPU, allocations, contention, and p99/p999.
  • Unmeasurable or negative changes do not become default.

Related: #285, #291, #303, #569

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions