perf(internal_logs source): decouple broadcast drain from downstream send and batch events - #26518
Open
thomasqueirozb wants to merge 3 commits into
Open
thomasqueirozb wants to merge 3 commits into
thomasqueirozb wants to merge 3 commits into
Conversation
Open
3 of 9 tasks
…queue to one batch
# Conflicts: # src/sources/internal_logs.rs
thomasqueirozb
added this pull request to stack #26519
September 30, 2026 19:14
thomasqueirozb
marked this pull request as ready for review
September 30, 2026 19:15
There was a problem hiding this comment.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #25218, which makes
internal_logsbroadcast lag visible throughcomponent_discarded_events_total{intentional="false"}. This PR reduces how often that lag occurs.recv_manyand callssend_batch. This keeps the broadcast receiver drained while the sink is backpressured, and amortizes per-event overhead downstream.MAX_BATCH_SIZE, 1024 events). A 10,000-event queue lowered the console drop rate by only 0.2 percentage points (benchmark) and raises worst-case memory 10x.ShutdownSignalfor the fullrun()scope, so the source does not report completion while batches are still in flight.send_batchfails, the drain task is aborted before the source returns, so it cannot keep itsShutdownSignalclone alive.Moved from #25218. The review threads there about
INTERMEDIATE_QUEUE_CAPACITYand drain task cleanup whensend_batchfails are addressed in 5604fed.Vector configuration
Minimal config (from the issue, console sink):
Blackhole sink (isolates the source path from stdout/JSON costs):
Benchmark
The prometheus_exporter is scraped at the end of each 20s run to read
component_received_events_totalandcomponent_discarded_events_total{intentional="false"}for theinternal_logssource."Single-task loop" and "master (patched)" rows use master's source loop with lag counted in the drop metric, which is the behavior after #25218.
The tables below were measured with an intermediate queue capacity of 10,000. The final capacity of 1024 changed the console drop rate by 0.2 percentage points in a separate run (details).
Design comparison, console sink,
VECTOR_LOG=trace, 20sBuffer size is the broadcast capacity in
src/trace.rs.Buffer size made almost no difference (3-6% fewer drops) with either design, so the original 99 is retained.
Sink comparison,
VECTOR_LOG=trace, 20sInterpretation:
send_eventand broadcast consumption being coupled cap throughput at ~5k events/sec delivered and drop ~86% of events even when the sink is free (blackhole).tracebecause stdout + JSON encoding caps at ~20k events/sec.Under
VECTOR_LOG=debug(normal load), both configs show zero drops.How did you test this PR?
cargo nextest run --no-default-features --features sources-internal_logs --lib sources::internal_logs::(all tests pass, includingbroadcast_lag_increments_discarded_metricfrom fix(internal_logs source): report broadcast lag drops in component_discarded_events_total #25218)cargo clippy --no-default-features --features sources-internal_logs --lib --tests -- -D warningscargo vdev check eventsVECTOR_LOG=debugandVECTOR_LOG=trace, comparingcomponent_received_events_totalandcomponent_discarded_events_total(see Benchmark).Change Type
Is this a breaking change?
Does this PR include user facing changes?
no-changeloglabel to this PR.References
internal_logssource silently drops logs under high load #24220