Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
4 changes: 2 additions & 2 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -12,5 +12,5 @@ abstract: >-
case and fixture digests, analyzer identities, normalized reports, and
retained raw evidence.
license: MIT
version: v0.3.0
date-released: "2026-08-24"
version: v0.4.0
date-released: "2026-08-26"
406 changes: 246 additions & 160 deletions adapters/bifrost/README.md

Large diffs are not rendered by default.

343 changes: 195 additions & 148 deletions adapters/codeql/README.md

Large diffs are not rendered by default.

30 changes: 10 additions & 20 deletions adapters/joern/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -242,25 +242,15 @@ zero error documents, zero `inconclusive`, `unsupported`, or `runner-error`
outcomes.

None of the six still stands at the pre-expansion fixture revision
`sha256:aee59a14f96633cf5798df6d211525ea0d10748800ba9c9ac0a3787406bd19ea`.
The Python, JavaScript, Java, Rust, PHP, and Ruby kernels were each re-run whole
after that
language's challenge-tier row was rolled out, and each carries the expanded
corpus revision current when it ran —
`sha256:3e7a8de5e1eefb18e8166af0ccdf309bccf1d5c26026893a4513f1943926ab1f` for
Python,
`sha256:64ef139f452fd296bb26463bc552e5e5998ca4bb4584d45565d858424814bde9` for
JavaScript,
`sha256:f476894a41d283e3bcaaf5188ee08abe7886ce8e3919257403b0aa853ef718e2` for
Java,
`sha256:88ad35289ae465278b95fd436532132118a6b6aa681adb3d266d67766c8770c5` for
Rust,
`sha256:f74647fe824ca9f6900c48aa9d403f0e9f59230e4193e0b02bd65e29a9e4e660` for
PHP, and
`sha256:020d0d8f79360af6e74064a692e2d65ffa31cd97f9971f9dad8bec065d862043` for
Ruby, the last wave to land. `fixture_revision` digests the whole case corpus,
so each wave's fixtures
moved it for every run after it. Reports at different fixture revisions are not
`sha256:aee59a14f96633cf5798df6d211525ea0d10748800ba9c9ac0a3787406bd19ea`, and
none stands at an intermediate wave revision either. The Python, JavaScript,
Java, Rust, PHP, and Ruby kernels were all re-run whole for the v0.4.0 freeze
after every challenge-tier row had rolled out, so all six carry the single
revision
`sha256:13a11ff48f26dba889f76aeb9ef60213a129abe5ebcfcb966da3a2418c12807e`.
`fixture_revision` digests the whole case corpus, so each wave's fixtures
moved it for every run after it; re-running the whole adapter at one revision
is what removes that skew. Reports at different fixture revisions are not
pooled, and each language's expanded assertions are a different population from
the 32 (30 for Rust) it reported in v0.3.0, not a movement within one.

Expand All @@ -286,7 +276,7 @@ expansion introduced no drift — and 16/24 on its challenge twelve**.
templates: the sixteen v0.3.0 templates plus the thirteen preregistered
challenge templates ([the challenge tier](../../docs/challenge-tier.md)). Each
report was re-run whole — a whole-population replacement, not an append — and
each carries the expanded corpus revision current when it ran; no Joern kernel
all six carry the one v0.4.0 fixture revision; no Joern kernel
is left at `sha256:aee59a14f96633cf5798df6d211525ea0d10748800ba9c9ac0a3787406bd19ea`.
Split by stratum, JavaScript is **26/32 on the classic sixteen — identical case
for case to its v0.3.0 snapshot, so the expansion introduced no drift — and
Expand Down
28 changes: 9 additions & 19 deletions adapters/semgrep/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -390,29 +390,19 @@ does not require the marker's own line.
## Observed results

Semgrep CE 1.174.0. No kernel is left at the pre-expansion fixture revision
`sha256:aee59a14f96633cf5798df6d211525ea0d10748800ba9c9ac0a3787406bd19ea`. All
`sha256:aee59a14f96633cf5798df6d211525ea0d10748800ba9c9ac0a3787406bd19ea`, and
none is left at an intermediate wave revision either. All
eleven — Python, JavaScript, Java, TypeScript, Kotlin, Go, C++, C, Rust, PHP,
and Ruby — were each re-run whole after that language's challenge-tier row was
rolled out and carry the expanded corpus revision current when each ran —

| Kernel | `fixture_revision` |
| --- | --- |
| Python | `sha256:3e7a8de5e1eefb18e8166af0ccdf309bccf1d5c26026893a4513f1943926ab1f` |
| JavaScript | `sha256:61c06a78b95b86764d3c220cfefd7af37373db64b15ae0b76c6ebf924217ab2e` |
| Java | `sha256:cf571f29e434030019d5e8f8361319b0bb3b4d6c4c752bd65860e07bfcf26bbc` |
| TypeScript | `sha256:2c906faeb98b48d1aba7da7bc80a78c4084051b84efac6ac3a1b74f54c843fd2` |
| Kotlin | `sha256:7ac23321e5d0974ed9087b9642ee3c88b3f3af014ba507330131da30fbb9b4d7` |
| Go | `sha256:7f37b99ddab7764a8536112c09ff7c8d77e0b02f7786abde65dfbaf3654d9949` |
| C++ | `sha256:a1570fc74526f0088488e3fba0941a7da47244635d7ceecf6787f1f76200b4ee` |
| C | `sha256:75f631ca05df2609055972622faaf3946331f7537140b08ba7ec6648bd0e077c` |
| Rust | `sha256:88ad35289ae465278b95fd436532132118a6b6aa681adb3d266d67766c8770c5` |
| PHP | `sha256:f74647fe824ca9f6900c48aa9d403f0e9f59230e4193e0b02bd65e29a9e4e660` |
| Ruby | `sha256:020d0d8f79360af6e74064a692e2d65ffa31cd97f9971f9dad8bec065d862043` |
and Ruby — were re-run whole for the v0.4.0 freeze, after every
challenge-tier row had rolled out, so all eleven carry the single revision
`sha256:13a11ff48f26dba889f76aeb9ef60213a129abe5ebcfcb966da3a2418c12807e`.

`fixture_revision` digests the whole case corpus, so each wave's fixtures moved
it for every run after it, and reports at different fixture revisions are not
pooled. The configuration hash is unchanged across all eleven: no rule file was
touched.
pooled; re-running the whole adapter at one revision is what removes that skew.
The configuration hash
`865d0bd2989f9ddd0b90f2d6675584e86706b109a033d4a1ac00bd21a617b100` is the same
across all eleven: no rule file was touched.

All eleven kernels ran. 622 assertions: 154 executed against Semgrep, 468
excluded by declared capability. Zero `inconclusive` and zero `runner-error`
Expand Down
15 changes: 13 additions & 2 deletions docs/astro.config.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ export default defineConfig({
base: '/dataflowbench',
redirects: {
// Explicit current-snapshot pointer alongside versioned snapshot URLs.
'/current': '/dataflowbench/snapshots/v0-3-0/',
'/current': '/dataflowbench/snapshots/v0-4-0/',
},
integrations: [
starlight({
Expand All @@ -34,7 +34,18 @@ export default defineConfig({
items: [
{ label: 'All snapshots', slug: 'snapshots' },
{
label: 'v0.3.0 (current)',
label: 'v0.4.0 (current)',
items: [
{ label: 'Snapshot overview', slug: 'snapshots/v0-4-0' },
{ label: 'Analyzers', slug: 'snapshots/v0-4-0/analyzers' },
{ label: 'Languages', slug: 'snapshots/v0-4-0/languages' },
{ label: 'Semantic templates', slug: 'snapshots/v0-4-0/templates' },
{ label: 'Case evidence', slug: 'snapshots/v0-4-0/evidence' },
],
},
{
label: 'v0.3.0 (archived)',
collapsed: true,
items: [
{ label: 'Snapshot overview', slug: 'snapshots/v0-3-0' },
{ label: 'Analyzers', slug: 'snapshots/v0-3-0/analyzers' },
Expand Down
24 changes: 24 additions & 0 deletions docs/milestones.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,30 @@ v0.4.0 cores are different populations and are never compared number-to-number;
the frozen v0.3.0 evidence is unaffected because freeze validation is
manifest-scoped.

### Current expanded-breadth status

Release v0.4.0 is that first expanded-breadth freeze: 744 cases and 42 bound
reports at one fixture revision, with all thirteen kernels carrying their
expanded core denominators and four analyzer populations — Bifrost v0.10.6
(build `18d09c57`) over the pinned 118-case breadth slice plus thirteen
kernels, CodeQL 2.26.3 over eleven, Joern 4.0.610 over six, and Semgrep CE
1.174.0 over eleven. Coverage differs per analyzer; a language with no report
for an analyzer is coverage, not a score.

The tier did what it was preregistered to do: no analyzer answers a whole
expanded core correctly in any language, so the core no longer saturates.
Bifrost decides 115 of the 118 breadth-slice cases and 227 of its 738 core
assertions, 222 of those decisions correct, declining the rest as
`inconclusive`; its `element-object` `internal_invariant` failures and its
fully inconclusive Ruby kernel are published as retained execution and
capability coverage and tracked upstream. CodeQL produces a definitive answer
for all 626 of its bound assertions, 509 correct, with dynamic dispatch its
systematic miss. Joern answers all 344 of its assertions, 270 correct, and its
depth-6 behavior tracks the call-depth bound the preregistration verified in
advance. Semgrep CE scores the 154-assertion intraprocedural partition of its
populations and declines the other 468 by declared capability. See
[`releases/v0.4.0.md`](releases/v0.4.0.md) for the bound evidence.

## M3: taint modeling

Add balanced categories for sources and sinks, propagators, sanitizers, opaque
Expand Down
Loading
Loading