Skip to content

perf: even faster trie - #327

Open
lczyk wants to merge 2 commits into
canonical:mainfrom
lczyk:perf/tree-perf-improvements
Open

lczyk wants to merge 2 commits into
canonical:mainfrom
lczyk:perf/tree-perf-improvements

Conversation

@lczyk

@lczyk lczyk commented Sep 29, 2026

Copy link
Copy Markdown
Contributor
  • Have you signed the CLA?

since #302 got merged, here is a followup which was originally designed as letFunny#29


two optimisations for pathConflictTree, on top of the trie from this branch. no behaviour change intended: same conflicts detected, same error type. tests pass unchanged. checked with my fuzz tests too.

1. reuse the traversal queues (f449d7a)

pathHasConflict used to grow currentQueue and nextQueue from nil on every path insertion, and seeded the first level with slices.Collect(maps.Values(...)). with a few thousand paths this dominated allocation: the queues were the single largest source of gc work in the whole conflict check.

i've changed currentQueue and nextQueue to fields on pathConflictTree, reset with [:0] per call and restored through a defer, so the backing arrays are grown once and reused.

proof of gc work (drop somewhere in internal/setup):

func BenchmarkConflictAlloc(b *testing.B) {
      paths := map[string][]*setup.Slice{}
      for i := 0; i < 1000; i++ {
              pkg := fmt.Sprintf("pkg%04d", i)
              p := "/usr/bin/" + pkg + "-*"
              paths[p] = []*setup.Slice{{Package: pkg, Name: "s",
                      Contents: map[string]setup.PathInfo{p: {Kind: setup.GlobPath}}}}
      }
      for i := 0; i < b.N; i++ {
              tree := setup.NewConflictTree(paths)
              if err := tree.HasConflict(); err != nil {
                      b.Fatal(err)
              }
      }
}
$ go test ./internal/setup/ -run '^$' -bench ConflictAlloc -benchtime 20x \
  -memprofile /tmp/mem.prof -o /tmp/base.test
$ go tool pprof -list 'pathHasConflict$' -sample_index=alloc_space /tmp/base.test /tmp/mem.prof | grep -E '(nextQueue =|of Total)'
  218.43MB   218.43MB (flat, cum) 90.63% of Total
  218.43MB   218.43MB    156:								nextQueue = append(nextQueue, child)
$ GODEBUG=gctrace=1 go test ./internal/setup/ -run '^$' -bench ConflictAlloc -benchtime 20x 2>&1 | grep '^gc ' | wc -l
     149

after:

$ go tool pprof -list 'pathHasConflict$' -sample_index=alloc_space /tmp/base.test /tmp/mem.prof | grep -E '(nextQueue =|of Total)'
  514.38kB   514.38kB (flat, cum)  1.72% of Total
         .          .    111:		g.currentQueue, g.nextQueue = currentQueue, nextQueue
  514.38kB   514.38kB    166:								nextQueue = append(nextQueue, child)
$ GODEBUG=gctrace=1 go test ./internal/setup/ -run '^$' -bench ConflictAlloc -benchtime 20x 2>&1 | grep '^gc ' | wc -l
      40

2. exact-match child lookup (3c097a7)

when a segment matched, the old traversal enqueued every child of the matched node and compared each one against the next segment on the following round. but Children is already keyed by segment text, so for a literal next-segment the only children that can possibly match are the one with the identical text and any child whose segment holds a wildcard.

now appendCandidates does and exact map lookup, plus a scan of the new node.GlobChildren list. a wildcard next-segment still considers every child, since it can match anything. glob-vs-glob comparisons are unchanged.

this turns the linear sibling scan in wide directories into a map lookup. real chisel-releases paths are at present aboout ~60% literal directories and this speeds this up, especially because the wide nodes tend to also be mostly literal so the quadratic scan hurts a lot there).

proof, with fix 1 above applied, otherwise GC noise swamps the signal:

func BenchmarkConflictWideDir(b *testing.B) {
      for _, n := range []int{1250, 2500, 5000} {
              paths := map[string][]*setup.Slice{}
              for i := 0; i < n; i++ {
                      pkg := fmt.Sprintf("pkg%04d", i)
                      p := fmt.Sprintf("/usr/share/doc/file%05d", i)
                      paths[p] = []*setup.Slice{{Package: pkg, Name: "s",
                              Contents: map[string]setup.PathInfo{p: {Kind: setup.CopyPath}}}}
              }
              b.Run(fmt.Sprintf("files=%d", n), func(b *testing.B) {
                      for i := 0; i < b.N; i++ {
                              tree := setup.NewConflictTree(paths)
                              if err := tree.HasConflict(); err != nil {
                                      b.Fatal(err)
                              }
                      }
              })
      }
}
$ go test ./internal/setup/ -run '^$' -bench ConflictWideDir -count=1 | grep 'BenchmarkConflictWideDir'
BenchmarkConflictWideDir/files=1250-8         	      68	  14806423 ns/op
BenchmarkConflictWideDir/files=2500-8         	      16	  69988763 ns/op
BenchmarkConflictWideDir/files=5000-8         	       4	 273993969 ns/op

after:

go test ./internal/setup/ -run '^$' -bench ConflictWideDir -count=1 | grep 'BenchmarkConflictWideDir'
BenchmarkConflictWideDir/files=1250-8         	    1753	    679446 ns/op
BenchmarkConflictWideDir/files=2500-8         	     903	   1294642 ns/op
BenchmarkConflictWideDir/files=5000-8         	     466	   2629180 ns/op

The traversal queues regrew from nil on every path insertion, which
accounted for ~90% of allocated bytes. BenchmarkConflictTree/pkgs=1000:
-91% B/op, -37% allocs/op, -17% sec/op.
A literal segment scanned every sibling node linearly even though
Children is keyed by segment text. Look up the exact child directly and
scan only wildcard children, tracked in a new GlobChildren list.
BenchmarkConflictTree/pkgs=1000: -70% sec/op; DeepPaths: -85% sec/op.
@lczyk
lczyk requested a review from letFunny September 29, 2026 07:19
@lczyk lczyk changed the title perf: even faster tree perf: even faster trie Sep 29, 2026

@letFunny letFunny left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @lczyk thank you for the interest. I see in the CI the speedup is 40%, nice! From my point of view the changes look correct, and they are quite simple.

I defer the decision on whether to prioritize the PR, the style of the code and the approval to @upils as the maintainer :)

type pathConflictTree struct {
Root *node
PathToSlices map[string][]*Slice
// currentQueue and nextQueue are kept across pathHasConflict calls so

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is okay, the downside is that we are adding state and keeping the queues in memory for as long as the object lives. I think it is a good tradeoff as the graph will also be in memory for the same amount of time.

Comment on lines +111 to +112
currentQueue := g.currentQueue[:0]
nextQueue := g.nextQueue[:0]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Even though we now have state we are clearing it before each run which makes it safe.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: With the current approach both queues are non-empty. So one could think that the stored values are meaningful and may be used by another function/method between executions of pathHasConflict. What about also clearing the queues in the deferred call?

Segment segment
SegmentSlices []*segmentSlice
Children map[string]*node
// GlobChildren lists the children whose segment contains a wildcard, so

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This also looks okay to me. There are many optimizations optimizations we can do when looking for Children to reduce the amount of comparisons. GlobChildren seems a good compromise between being quite simple and having an impact on performance. On that note, what is the effect? I remember you said ~20%.

Nit: I find the comment describes the usage with regard to the other field in a bit of a convoluted way. I would instead say this is an optimization that has the subset of children with globs and it is used in appendCandidates...

@upils upils left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @lczyk. This is nice to gain even more perf with a reasonable added complexity.

Once the suggested changes are done I think I will still keep this PR in Group Review as I would like to keep the reviews focused on other more pressing matters for now.

currentQueue := g.currentQueue[:0]
nextQueue := g.nextQueue[:0]
defer func() {
// Keep the grown queues for the next call.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe it should be very clear that these are kept only to avoid more allocations, and it does not serve any other purpose in the algorithm. So far the recently added complexity was essentially due to implementing the logic itself. The propose change here is a pure optimization and being able to "ignore" it when reasoning about the logic might help.

Comment on lines +198 to +201
for _, child := range parent.Children {
queue = append(queue, child)
}
return queue

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
for _, child := range parent.Children {
queue = append(queue, child)
}
return queue
return slices.AppendSeq(queue, maps.Values(parent.Children))

Comment on lines +111 to +112
currentQueue := g.currentQueue[:0]
nextQueue := g.nextQueue[:0]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: With the current approach both queues are non-empty. So one could think that the stored values are meaningful and may be used by another function/method between executions of pathHasConflict. What about also clearing the queues in the deferred call?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants