Purpose
SFA-Bench invites independent researchers, engineers, auditors, and governance teams to reproduce the R1 memory-boundary protocol and independently verify the published evidence.
R1 result
Across 30 preregistered and separately ratified executions of one frozen memory-boundary task:
- 10 executions passed;
- 20 produced partial results containing
state_loss;
- no execution used forbidden state.
The three pilot executions were excluded from the replication estimates.
Requested independent checks
An independent reviewer or replication team is invited to:
- verify the published SHA-256 hashes;
- reproduce the descriptive calculations;
- inspect the deterministic scoring;
- review the human-ratification and closure lineage;
- run a separately identified replication campaign;
- disclose all deviations from the frozen R1 protocol.
Reporting expectations
A replication should disclose:
- model identifier or immutable snapshot, where available;
- provider and execution date;
- attempt and retry policy;
- deviations from the protocol;
- completion counts and deterministic judgments;
- human-review procedure;
- complete evidence hashes.
Independent replications must not modify or be presented as extensions of the original R1 evidence record.
Interpretation boundary
R1 is a descriptive evaluation of one frozen task under mutable provider aliases. It is not a model ranking, endorsement, safety certificate, or regulatory approval.
Purpose
SFA-Bench invites independent researchers, engineers, auditors, and governance teams to reproduce the R1 memory-boundary protocol and independently verify the published evidence.
R1 result
Across 30 preregistered and separately ratified executions of one frozen memory-boundary task:
state_loss;The three pilot executions were excluded from the replication estimates.
Requested independent checks
An independent reviewer or replication team is invited to:
Reporting expectations
A replication should disclose:
Independent replications must not modify or be presented as extensions of the original R1 evidence record.
Interpretation boundary
R1 is a descriptive evaluation of one frozen task under mutable provider aliases. It is not a model ranking, endorsement, safety certificate, or regulatory approval.