Replies: 3 comments
|
Follow-up questions, posted to AI, as these PRs focused on producing a functional equivalent to what I had done 3 years ago:
Good questions — I went and read the actual interface family, the list-merge algorithm, and cross-checked against the current 1. Coverage of new entity types (last ~3 years)Partially covered, and there's a real gap. The But two other additions are not covered at all, silently:
Neither list is even copied into the merged output — not by my new code, and not by the original pre-existing plain 2. Will future classes "just work"?Not automatically — two things are still manual, by design (the interfaces deliberately avoid reflection/a shared base class):
3. Do list merges reconcile, or just append?They reconcile, by default. Right now only One loose end I noticed while checking this: the doc comment above |
|
|
FWIW, looking at existing issue lists, I've asked around and got more AI analysis for current status:
I actually built the full PR stack and ran several of these issues' exact reproduction cases against it rather than just reading code, so here's what's real vs. not. Vulnerability-data merge issues — not fixed
If vulnerability-aware merging matters to you, that's a real gap this PR stack doesn't close — Metadata-block issues — mostly not fixed
What I actually tested (adjacent, but not what you asked about) — mixed results, one new bug found
Want me to open a follow-up issue (or draft PR) for the dangling-dependency-ref bug I found in CycloneDX/cyclonedx-cli#399, and/or scope a |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem overview
For various reasons, including anticipated customer demand and actual legal compliance, as well as for internal benefit for our developers to plug security holes in a timely manner (especially in old but supported releases of the large product my dayjob is making), we wanted to produce a "Release Bundle SBOM" -- single document which lists all components that end up in the customer's environment when some or all components of our product get installed there. One aspect to this is honestly listing the FOSS components and their license/author attributions, another is knowing their versions deployed (eventually the potential vulnerabilities to investigate), yet another is having a dependency graph so we can do this for any subset of our product.
To clarify, an installation of our product delivers roughly 20 to 50 separate programs (each with a dependency tree into Java, DotNet, NodeJS and other ecosystems). Some editions are used as plain programs, others can be delivered as containers for easier runs in arbitrary environments - but so add dependency trees from the base OS images involved. While each component's dependency tree is strictly-speaking unique, they largely overlap per ecosystem, especially starting a couple of levels below our import statements (e.g. "using Spring" may mean certain variability in our Java
pom.xmlfiles to pick the components our program uses, but then it is the Spring sources' curated choice of what gets pulled in - with little control from our side). Each incoming document is a complete SBOM of the dependency or product it describes - not just that one component entry, but also its dependency tree (and so the other components which this tree mentions).At the scale of our project, a completely honest merge - which is this toolkit's default - deals with about 30000 components and their unique 1:1 dependency links (with a LOT of repeating names). When I began with this 3-4 years ago, it was not what our team expected, to the point of considering the retention of all duplicates a bug (what sort of a "merge" is that?), but during the review discussions I've learned it is done for quite a valid reason - the same potentially-vulnerable component can be a security issue when used in one codepath, and not a problem in another (e.g. exposed to network or not, or the component has many features that are enabled in some configurations and not called in others). So the vulnerability investigations, such as verdicts maintained in DependencyTrack server, should take each such link as a separate use-case that may be or not be vulnerable independently.
This is nice in theory, but we are not gonna get anyone looking at, and clicking through in DT, the 30k use-cases and ultimately verdicts - neither several times a week as the team iterates the codebase and issues patches to older releases, nor even once in a quarter of a year for new initial public releases. Just the sheer amount of this work would mean that we have to pull our most experienced and expensive team members, who know the project and various implications for different customers' unique deployments like a chess gross-meister playing a dozen games at once from memory, into man-weeks worth of this review instead of actual work on, you know, developing the product.
Given that the majority of our dependency links are outside our control anyway (the dependency webs of frameworks and ecosystems we build upon), the reasonable trade-off is to have a de-duplicated merged document, where each instance of a component is named once, and has multiple pointers (graph links) from whoever touches it. In our case, the merged document with 30k incoming components ends up with 700-1000 components, and a much smaller amount of vulnerability instances to investigate when new relevant CVEs come into our DT server.
Looking at the issue tracker here, using a "lossy" merge strategy is a reasonable trade-off for many other askers too :)
While my effort started with a single goal way back when, I found that actually "reasonable expectations" do differ from project to project, so I tried to introduce a concept of merge strategies which impact decisions in otherwise similar code structure, so the mostly-same iterations do not have to be written again and again once somebody has a bright idea of doing something differently.
Below I will list preceding tickets from the currently proposed solution, from the older iteration, and mention other issues that asked about a similar feature or voiced concerns about the original full or flat merge implementations (if only to have GitHub references from them, so original askers are notified that their wishes may have been heard).
Problems solved
As I followed this windy road, several problems were identified and addressed:
bom-refidentifier for some component (and inversereffrom dependency declarations) in its original document is usually the same string as itspurl. When you add and keep dozens of copies of essentially the same component during such merges of related dependency trees, each copy (and corresponding back-references) should use an unique name.scopevalues in different copies of dependency branches (e.g. used in tests=>excludedfrom the deliverable vs. used in production=>required). This led to introduction of different strategy options: a simpler merge conflates a few meanings to keep exactly one mention of a component (e.g. "optional+required" = "required") but aborts production of a document in case of conflicts that can not be reconciled ("required+excluded = ?"), and a recently completed optional strategy allows to rename such components so theirbom-refadds ascope=...suffix for each different copy and merges the other information into both, grouping the back-references reasonably (all excluded go here, all required go there).--input-file NAMEthey exceed the OS-dictated CLI limits. An option was added to pass an--input-file-listfor this (additionally/instead).Ready to go for a spin?
For smaller review scopes, the project changes were split into multiple pull requests listed further below. Here are my recent development branches which include all of these changes, and maybe something on top (like internal README notes about getting the custom-versioned builds to play together) so brave souls can just try it out like I do:
New proposal:
cyclonedx-dotnet-library -- 5 new stacked PRs, each build+test verified standalone, in proposed merge order:
cyclonedx-cli -- 8 new stacked PRs, each build+test verified against a packed library snapshot, in proposed merge order:
Old proposal
Most PRs are closed as superseded by those above:
cyclonedx-dotnet-library:
cyclonedx-cli:
Related issues in the tracker
Listing to show that not only our team saw this as a problem, and for GitHub to add backlinks in those tickets to this epic as a possible solution :)
merge-d polyglot SBOM loses dependency graph information cyclonedx-cli#179metadata.componentfrom original SBOMs cyclonedx-cli#218 - can't say OTOH if that situation is addressed here and now, or not; still worth a mentionmetadata.toolsnot merged correctly when one SBOM uses legacy format and the other uses the newer format cyclonedx-cli#408 - maybe addressedHonourable mentions
Kudos to everyone who participated in discussions about the initial development, whether in GitHub reviews or in the Slack and other channels. After 3 years I probably would not tag everyone who helped move this forward, but at least gotta tip the hat to those I can: @nscuro @andreas-hilti @jkowalleck @coderpatros @fnxpt @stevespringett @pombredanne ...
Special thanks to @claude for making sense from my original ideas and reflection-heavy implementation into (more) proper code :)
Also thanks to my dayjob for putting up with the man-weeks or even months sunk into this work, back in the day and recently, and yet letting it be out in the open.
All reactions