Skip to content

Milestones

List view

  • Scale from Pilot v1 to a full cross-lingual safety study: hybrid LLM-judge refusal classifier, 300–500 prompts, resolved model coverage, human validation subset. See docs/methodology.md §9 for the detailed plan.

    No due date
    0/5 issues closed