Skip to content

Use one codegen unit for release builds - #4692

Closed
charliermarsh wants to merge 3 commits into
mainfrom
charlie/codex-ty-pgo-codegen
Closed

charliermarsh wants to merge 3 commits into
mainfrom
charlie/codex-ty-pgo-codegen

Conversation

@charliermarsh

@charliermarsh charliermarsh commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Summary

One codegen unit produced smaller binaries and reduced peak compiler memory, but increased combined PGO compilation time by 24–27%. Runtime benefits were modest.

We compared a release default of 16 → 1 codegen units using Rust 1.99.0, fat LTO, and PGO on the same source revision:

Platform Binary size PGO compilation time Sampled peak process memory
Linux x86_64 28.29 → 27.66 MiB 8m 04s → 10m 06s 5.26 → 3.91 GiB
macOS ARM64 24.90 → 23.58 MiB 6m 09s → 7m 37s 11.52 → 8.63 GiB
Windows x86_64 30.28 → 30.04 MiB 9m 40s → 12m 15s 4.81 → 3.71 GiB

Binary size fell by 0.8–5.3% and peak process memory by 23–26%. Compilation time includes the instrumented and final PGO builds. Each configuration was built once from a clean target directory on the same runner; build order and OS caches were not controlled. Linux measurements used native builds rather than the manylinux packaging container.

Runtime performance: We checked 11 pinned project source subsets, including two source trees excluded from training, across two rounds of 30 alternating pairs. Both binaries produced identical diagnostics. The PGO corpus uses sparse checkouts, so imports of omitted sibling packages remain unresolved; some optional dependencies are also absent. Runtime timings include rendering those diagnostics. The astropy/units subset was consistently 0.7–2.4% faster. Other cases showed no repeatable regression; Windows results were noisier and sometimes changed direction between rounds. Single-thread Black and Warehouse medians were 0.2–1.3% faster on Linux and macOS; Windows results were inconclusive.

In a two-file fixture with cross-file edits, both binaries returned identical language-server replies. Session latency changed little, and the edit-latency measurements did not establish a difference. Session timing includes startup and protocol transport.

Detailed evidence is available in the build and runtime comparison and the single-thread follow-up.

@charliermarsh charliermarsh added the internal An internal refactor or improvement label Oct 8, 2026
@charliermarsh

Copy link
Copy Markdown
Member Author

Closing this because the measured gains do not justify changing the release default.

One codegen unit reduced binary size by 0.8–5.3% and sampled peak process memory by 23–26%, but combined PGO compilation took 24–27% longer in these runs. Runtime gains were modest in the tested source subsets, and the incremental language-server measurements did not establish an improvement. We have not identified a build-memory constraint that would make this tradeoff worthwhile.

We'll keep the existing release configuration. The measurements and their limits remain in the description for reference.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

internal An internal refactor or improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant