Repository navigation
Conversation
Code reviewNo correctness bug found. The fast path is equivalent to the general path at both ends of the range:
Perf reproduces on an M4 Max, baseline being this tree with only the 9-line hunk removed:
A +1..12% regression on the bulk Two notes on the benchmark1. 2. It is a new file for a single benchmark. |
b.Loop already keeps loop-body results alive on go1.24.
FoundRedundant sink. Branch was 4 commits behind master, so its CI legs never ran the iOS SIGILL detection fix in FixedSink dropped, master merged. 2100219 KeptThe
It is equivalent to the general path for every |
Summary
copyand slice bookkeeping in byte-at-a-time callersWhy
Some streaming users feed trie or protocol prefixes one byte at a time. The general buffering path is correct, but its slice and
copysetup dominates such small writes. The fast path writes directly into the existing sponge buffer and still permutes immediately when the rate block fills.This is a performance-only refactor, so TDD is not applicable. The existing seeded fuzz test already checks byte-by-byte writes at, before, and after the 136-byte boundary against
x/crypto/sha3.Benchmark
Apple M2 Max, Go 1.25.7, median of five runs:
All cases remain at zero allocations.
Verification
go test ./...go test -race ./...go test -tags purego ./...go vet ./...go mod tidy -diffgolangci-lint run(twice)