Skip to content

Publish stratified v1, v2, and reviewed-legacy results on the website #122

Description

@DavidBakerEffendi

Part of #117. Depends on #121; it should also incorporate #118 when the
real-project-v2 release is available.

Goal

Publish an expanded, evidence-derived website view that answers “how good is
Bifrost?” using all reliable evidence without concealing the difference between
prospective evaluation and retrospectively reviewed legacy cases.

Required presentation

  1. Prospective evaluation: v1 and v2 shown with their own frozen
    per-profile denominators and release identities.
  2. Balanced reviewed legacy core: its own selection provenance, review
    tier, per-language denominators, exclusions, replacements, and immutable
    artifact identity.
  3. Remaining development corpus: retained as regression evidence and never
    silently included in a publication-grade headline.

The homepage should show corpus breadth and confidence together—for example,
“36 published v1 cases, 36 preregistered v2 cases, and N×11 independently
re-reviewed legacy cases”—rather than making 36 look like the entire evidence
base.

Execution and release

  • Freeze the reviewed legacy ground truth before new analyzer execution.
  • Rerun Bifrost and each exact registered reference profile against the frozen
    cohort with candidate identity and environment provenance.
  • Generate result tables from hash-verified reports; do not hand-copy scores.
  • Keep candidate/profile, language, slice, and selection-regime denominators
    available in machine-readable release metadata.
  • Use reproducible candidate environments, but do not introduce a two-host
    evidence requirement.

Acceptance criteria

  • All public numbers are generated from immutable, checksum-verified report
    artifacts and fail CI on drift.
  • V1, v2, legacy-reviewed, controls, overflow, and remaining development
    cases cannot be mixed accidentally by the result generator.
  • The homepage answers the headline question quickly and shows the total
    evidence breadth without flattening trust tiers.
  • Per-language/profile tables precede any aggregate, and every aggregate
    names its weighting and included slices.
  • Prospectively selected and retrospectively reviewed results remain
    visually and textually distinct.
  • Historical development pages and raw review evidence remain accessible.
  • Public limitations exclude language-wide/ecosystem-wide superiority,
    causal defect attribution, and latency/memory/startup claims not measured
    by the release.
  • Documentation and release audit expose selection provenance, agent/model
    provenance, human accountability, source locks, candidate identities,
    report hashes, and exact UsageBench revision.
  • cargo test, full corpus/evaluation/promotion validation, docs checks,
    production build, internal links, and rendered desktop/mobile QA pass.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions