You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
RESULTS.md at the repo root is the single source of truth for all experiment results. Every table currently has -- placeholders and NOT RUN status. This issue tracks updating those placeholders as runs complete.
How to update
When a run finishes, edit RESULTS.md and:
Replace the NOT RUN status cell with the run date (YYYY-MM-DD).
Fill in result cells with actual numbers (mean ± std where applicable).
Add a bootstrap 95% CI where the table has a CI column.
Reference the eval JSON path in a comment on this issue so results are traceable.
Tables to fill (in priority order)
Section 1 — Main Evaluation: rows Base, Base+BOHDI, LoRA no wrapper, LoRA+BOHDI. Requires eval pipeline to complete on the 200-example holdout.
Do not report a number in RESULTS.md without also referencing the eval JSON that produced it (e.g. eval/lora_no_wrapper.json @ commit abc1234). This keeps the table auditable.
What this is
RESULTS.mdat the repo root is the single source of truth for all experiment results. Every table currently has--placeholders andNOT RUNstatus. This issue tracks updating those placeholders as runs complete.How to update
When a run finishes, edit
RESULTS.mdand:NOT RUNstatus cell with the run date (YYYY-MM-DD).Tables to fill (in priority order)
Rule
Do not report a number in RESULTS.md without also referencing the eval JSON that produced it (e.g.
eval/lora_no_wrapper.json @ commit abc1234). This keeps the table auditable.