Skip to content

encode/zstd: support compressing with a trained dictionary - #8075

Closed
sorena-paydar wants to merge 1 commit into
caddyserver:masterfrom
sorena-paydar:feat/zstd-dictionary
Closed

sorena-paydar wants to merge 1 commit into
caddyserver:masterfrom
sorena-paydar:feat/zstd-dictionary

Conversation

@sorena-paydar

Copy link
Copy Markdown

Implements the dictionary option discussed in #6204.

@mholt wrote there:

An option to load a dictionary file should be pretty simple, as mentioned above. I'd welcome a PR to review!

This is that option, scoped to a single dictionary loaded from disk — not the full Compression Dictionary Transport negotiation the issue opens with. See Limitations below for why that part can't be done from an encoder module today.

What it does

Adds a dictionary option to the zstd encoder, taking a path to a dictionary produced by zstd --train:

encode {
	zstd {
		dictionary /etc/caddy/api.dict
	}
}

JSON:

{"encodings": {"zstd": {"dictionary": "/etc/caddy/api.dict"}}, "handler": "encode"}

Why it helps

Small, structurally similar payloads are the case that plain zstd cannot do much with, because frame overhead swamps the content. Measured on a 108-byte JSON body against a dictionary trained on 60 similar responses:

bytes
uncompressed 108
encode zstd 121 (larger than the input)
encode zstd + dictionary 39

That is the shape of a lot of API traffic behind a reverse proxy.

Limitations, and a question

A dictionary frame is only decodable by a client holding that dictionary. The dictionary is not carried in the frame; a decoder without it fails with unknown dictionary rather than degrading. So this option is for traffic between parties that share a dictionary out of band — service-to-service, internal API clients — and is not safe to turn on for browser traffic.

The web-facing answer to that is Compression Dictionary Transport: the client advertises Available-Dictionary: :<sha-256>: and the server replies with Content-Encoding: dcz only when it holds the matching dictionary. That cannot be implemented from an encoder module as things stand, because the interface is:

type Encoding interface {
	AcceptEncoding() string
	NewEncoder() Encoder
}

NewEncoder() receives no request, so the encoder cannot vary per client or inspect Available-Dictionary without changing a shared interface. That felt like the wrong thing to fold into this PR unprompted.

So my question: is a plain dictionary option the right surface, or would you rather this were gated behind a distinct encoding name so it can never be selected by a client that merely sent Accept-Encoding: zstd? I'm happy to rework it either way, including taking on the request-aware interface change if you want the negotiation properly.

The doc comment on the field states the constraint so it is visible from the JSON docs.

Implementation notes

The dictionary is read and validated in Provision, not per request, for two reasons: NewEncoder() has no way to report an error, and passing raw sample data where a trained dictionary belongs is an easy mistake — zstd.WithEncoderDict rejects it with magic number mismatch, which is not obvious. The error names zstd --train instead:

invalid zstd dictionary /etc/caddy/api.dict (expected the output of `zstd --train`): ...

Testing

New modules/caddyhttp/encode/zstd/zstd_test.go covers:

  • round trip through a dictionary-aware decoder, and an assertion that the frame header carries a non-zero DictionaryID
  • a decoder without the dictionary failing, which is the property that makes this unsafe for browsers
  • the no-dictionary path still producing plain, universally decodable zstd
  • Provision rejecting a missing file and a file that is not a trained dictionary, with the error naming zstd --train
  • Caddyfile parsing: valid, missing argument, repeated
  • the dictionary actually compressing better than plain zstd on a small payload

Plus a Caddyfile adapter test at caddytest/integration/caddyfile_adapt/encode_zstd_dictionary.caddyfiletest.

I checked the tests fail without the change, which caught a weak one: the round-trip test originally passed even with the dictionary wiring disabled, because a decoder holding dictionaries also reads plain frames. The DictionaryID assertion is what fixed it. With the wiring disabled three tests now fail; with it, all pass.

go test ./modules/caddyhttp/encode/... -count=1     ok
go test ./caddytest/integration/ -run TestCaddyfileAdaptToJSON  ok
go vet ./modules/caddyhttp/encode/zstd/             ok

testdata/sample.dict is a 2.8 KB dictionary trained on generated JSON samples; testdata/invalid.dict is plain text, for the rejection test.

Assistance Disclosure

AI-assisted. I directed the design and the scoping decisions, and verified the behaviour rather than taking it on trust: the compression figures above are measured, the round-trip and frame-header assertions are in the test file, and I confirmed the tests fail when the implementation is disabled. The Available-Dictionary limitation above came from reading the Encoding interface and finding it takes no request.

Adds a `dictionary` option to the zstd encoder that loads a dictionary
produced by `zstd --train` and compresses responses against it.

Small payloads are where this matters. A 108 byte JSON body grows to 121
bytes under plain zstd, since the frame overhead exceeds what there is to
compress, but falls to 39 bytes against a dictionary trained on similar
responses.

The dictionary is not carried in the frame, so only clients holding the
same dictionary can decode the response. That makes this useful between
services that share a dictionary out of band, and unsuitable for browser
traffic, which needs the Available-Dictionary negotiation from
Compression Dictionary Transport. That negotiation cannot be expressed
here: Encoding.NewEncoder() takes no request, so the encoder cannot vary
per client without changing a shared interface.

The dictionary is read and validated when the module is provisioned
rather than per request, because NewEncoder() has no way to report an
error and passing raw sample data instead of a trained dictionary is an
easy mistake to make. The error names `zstd --train` so it is clear what
the file should be.

Signed-off-by: Sorena Paydar <sorenapaydar81@gmail.com>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@steadytao

steadytao commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

You (your AI) do not understand compression and dictionaries enough to make this change.

@steadytao steadytao closed this Sep 28, 2026
@steadytao steadytao added the not reviewable ⛔ Fails basic standards or reasonably appears autonomous without demonstrated author understanding label Sep 28, 2026
@steadytao

Copy link
Copy Markdown
Member

Fails to comply with RFC 8878 and 9842

@sorena-paydar
sorena-paydar deleted the feat/zstd-dictionary branch September 28, 2026 12:39
@sorena-paydar

Copy link
Copy Markdown
Author

Very helpful comment, Thanks a lot !!

@steadytao

Copy link
Copy Markdown
Member

As stated. RFC 8878 and 9842.

@steadytao

Copy link
Copy Markdown
Member

I am more than happy to review a change but do not submit AI-written code that you do not understand the implications of to me and then expect me to explain to you why it is wrong. Be respectful in the future. For such a large change, one would generally wish to discuss further in the issue rather then submit a PR that does not work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

not reviewable ⛔ Fails basic standards or reasonably appears autonomous without demonstrated author understanding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants