You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Part of #117. Depends on #121; it should also incorporate #118 when the real-project-v2 release is available.
Goal
Publish an expanded, evidence-derived website view that answers “how good is
Bifrost?” using all reliable evidence without concealing the difference between
prospective evaluation and retrospectively reviewed legacy cases.
Required presentation
Prospective evaluation: v1 and v2 shown with their own frozen
per-profile denominators and release identities.
Balanced reviewed legacy core: its own selection provenance, review
tier, per-language denominators, exclusions, replacements, and immutable
artifact identity.
Remaining development corpus: retained as regression evidence and never
silently included in a publication-grade headline.
The homepage should show corpus breadth and confidence together—for example,
“36 published v1 cases, 36 preregistered v2 cases, and N×11 independently
re-reviewed legacy cases”—rather than making 36 look like the entire evidence
base.
Execution and release
Freeze the reviewed legacy ground truth before new analyzer execution.
Rerun Bifrost and each exact registered reference profile against the frozen
cohort with candidate identity and environment provenance.
Generate result tables from hash-verified reports; do not hand-copy scores.
Keep candidate/profile, language, slice, and selection-regime denominators
available in machine-readable release metadata.
Use reproducible candidate environments, but do not introduce a two-host
evidence requirement.
Acceptance criteria
All public numbers are generated from immutable, checksum-verified report
artifacts and fail CI on drift.
V1, v2, legacy-reviewed, controls, overflow, and remaining development
cases cannot be mixed accidentally by the result generator.
The homepage answers the headline question quickly and shows the total
evidence breadth without flattening trust tiers.
Per-language/profile tables precede any aggregate, and every aggregate
names its weighting and included slices.
Prospectively selected and retrospectively reviewed results remain
visually and textually distinct.
Historical development pages and raw review evidence remain accessible.
Public limitations exclude language-wide/ecosystem-wide superiority,
causal defect attribution, and latency/memory/startup claims not measured
by the release.
Documentation and release audit expose selection provenance, agent/model
provenance, human accountability, source locks, candidate identities,
report hashes, and exact UsageBench revision.
cargo test, full corpus/evaluation/promotion validation, docs checks,
production build, internal links, and rendered desktop/mobile QA pass.
Part of #117. Depends on #121; it should also incorporate #118 when the
real-project-v2release is available.Goal
Publish an expanded, evidence-derived website view that answers “how good is
Bifrost?” using all reliable evidence without concealing the difference between
prospective evaluation and retrospectively reviewed legacy cases.
Required presentation
per-profile denominators and release identities.
tier, per-language denominators, exclusions, replacements, and immutable
artifact identity.
silently included in a publication-grade headline.
The homepage should show corpus breadth and confidence together—for example,
“36 published v1 cases, 36 preregistered v2 cases, and N×11 independently
re-reviewed legacy cases”—rather than making 36 look like the entire evidence
base.
Execution and release
cohort with candidate identity and environment provenance.
available in machine-readable release metadata.
evidence requirement.
Acceptance criteria
artifacts and fail CI on drift.
cases cannot be mixed accidentally by the result generator.
evidence breadth without flattening trust tiers.
names its weighting and included slices.
visually and textually distinct.
causal defect attribution, and latency/memory/startup claims not measured
by the release.
provenance, human accountability, source locks, candidate identities,
report hashes, and exact UsageBench revision.
cargo test, full corpus/evaluation/promotion validation, docs checks,production build, internal links, and rendered desktop/mobile QA pass.