Repository navigation
perf(lookups): pack the code tables as base 36 differences - #605
Conversation
The lookup tables shipped their codes as fixed-width digits written back to back. They now ship the width and each code's difference from the one before it, in base 36 (packCodes in scripts/lookup-table.ts), which findCodeIndex unpacks once per table. A dense table takes about a third of the bytes: isValidNcm drops from 82.6 KB to 29.6 KB, isValidCbo, isValidCnae, isValidNbs and isValidCest by more than half, and the full import by 79.9 KB (gzip 356.1 KB to 327.3 KB). The generators write the new form, and the committed tables were repacked from the old ones without a new download (cbo.ts, the one generator that runs offline, reproduces its table byte for byte). isValidNcm now uses findCodeIndex too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RLkm9YrtAifc6XCLFVEsdH
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
commit: |
Tree-shaking report✅ No size regression. 1 grew, 14 shrank out of 194 exports.
What changed (15)
All exports (194)
How this is measuredEvery export is imported alone into an esbuild consumer bundle (minified, tree-shaken) built from the head and from the base of this pull request; the sizes are the resulting bundles, gzip is their gzipped size. 🔴 marks a regression: a pre-existing export that grew more than 20% and more than 256 B, or the bundle importing every pre-existing export growing more than 5%. 🟡 is growth under the threshold, 🟢 a decrease, ⚪ no change, 🆕 an export that does not exist on the base (never a regression), 🗑️ an export that was removed. An intentional increase is accepted with the |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## claude/municipalities-state-case #605 +/- ##
====================================================================
Coverage ? 100.00%
====================================================================
Files ? 237
Lines ? 2427
Branches ? 713
====================================================================
Hits ? 2427
Misses ? 0
Partials ? 0
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
indexOf scanned the whole unpacked table, which made isValidNcm about 60 times slower than the Set it used before (100k lookups: 20 ms before the packing, 1.2 s with indexOf, 30 ms now). The codes are ascending, so findCodeIndex bisects them. packCodes also rejects a repeated first code, which its order check let through. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RLkm9YrtAifc6XCLFVEsdH
…eady makes findCodeIndex only finds a code of exactly the table width, so the bank code and legal nature length checks before it could never change a result, and Stryker flagged them as survivors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RLkm9YrtAifc6XCLFVEsdH
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RLkm9YrtAifc6XCLFVEsdH
docs(contributors): add kwy404 for the RENAVAM number fix
What does this PR do?
Stacks on #604. The generated lookup tables (NCM, CBO, CNAE, NBS, CEST, CFOP, the LC 116 service list, legal natures and COMPE) used to ship their codes as fixed-width digits written back to back. They now ship the code width, then each code's difference from the previous one, written in base 36 (
"3:1,2,1"is001,003,004).packCodesinscripts/lookup-table.tsproduces this format.findCodeIndexunpacks a table once, on its first lookup, and finds codes in it by binary search.isValidNcmnow usesfindCodeIndexinstead of its ownSet.In a dense table the difference between consecutive codes is one or two characters, while a code is six to nine digits, so each table ships at about a third of its size:
isValidNcmisValidCboisValidNbsisValidCnaeisValidCestisValidCfopgetCbo/getCnae/getNbs/getCest/getCfopLookups stay fast. The first version used
indexOfand madeisValidNcmabout 60 times slower than its oldSet(100k lookups: 20 ms before, 1.2 s withindexOf). With the binary search the same run takes 30 ms, andisValidCbois faster than before (25 ms, previously a linear scan of the string).No public behaviour changes. A code is only found when it has exactly the table width and is made of digits, so
" 03","01 "and"1e0"are not found. The data is the same:scripts/cbo.ts, the one generator that runs offline, reproduces its table byte for byte.The generators, including
scripts/banks.tsforCOMPE_CODES, now write the packed form, andpackCodesrejects codes that are out of order or repeated. Tests read the tables through an independentunpackCodestest helper. The bundle size rows in both getting-started pages are updated.This PR also carries #595's browser-retry workflow change (see #604), merged in from its base.
Checklist
docs/utilities.mdanddocs/pt-br/utilities.md. Only the getting-started size tables change; the utilities docs are unaffected.npm run checkpasses locally. Also passing: coverage (100%, 8003 tests), Stryker onfind-code-index.tsandis-valid-ncm.ts(100%), Deno, Bun,check:unused,check:duplicationandcheck:api(no breaking change against 2.4.0).npm run build:llms.🤖 Generated with Claude Code
https://claude.ai/code/session_01RLkm9YrtAifc6XCLFVEsdH