perf(rules): optimize ruleset scope partitioning - #3183
Open
tryhard-26 wants to merge 1 commit into
Open
tryhard-26 wants to merge 1 commit into
tryhard-26 wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Please add bug fixes, new features, breaking changes and anything else you think is worthwhile mentioning to the master (unreleased) section of CHANGELOG.md. If no CHANGELOG update is needed add the following to the PR description: [x] No CHANGELOG update needed
github-actions
Bot
dismissed
their stale review
October 2, 2026 07:10
CHANGELOG updated or no update needed, thanks! 😄
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
during ruleset initialization in
capa.rules,_get_rules_for_scopewas called separately for all 8 scopes. inside that loop, it iterated over every non-subscope rule and invokedget_rules_and_dependencies, which rebuiltindex_rules_by_namespaceand the rule dictionary from scratch on every call. across the default 1,054 rule files (1,397 rules in memory), this triggered 8,440 redundant namespace index builds and repeated full dependency traversals, followed by 8 separatetopologically_order_rulessorts on the same rule set. this created an o(scopes * rules^2) bottleneck that added roughly 3.5 seconds of overhead to every ruleset instantiation.this change adds
_get_rules_by_scopeto traverse the dependency graph and topologically sort reachable rules in a single pass across all scopes, then slices each scope from that pre-ordered list.get_rules_and_dependenciesandtopologically_order_rulesnow accept optional pre-indexed namespace and name mappings to avoid redundant dictionary allocations.benchmarked across the default 1,054 rule files (1,397 rules in memory) in an identical harness:
a new regression test verifies that per-scope rule membership and topological dependency order invariants hold across the full ruleset.