Skip to content

ecosystem: publish suspect-zones benchmark corpus as a HuggingFace dataset #212

Description

@devswha

Why it matters

README headlines (91% / 76% / 13–25%) are unverifiable without the corpus. Publishing as HF dataset is the NLP-research reproducibility move and the prerequisite for a citation / arXiv paper.

Acceptance criteria

  • huggingface.co/datasets/devswha/patina-suspect-zones with current tests/fixtures/suspect-zones/ contents
  • Data card: per-source licensing, language coverage, how 91/76 were computed, known biases
  • CI step (manual trigger ok for v1) re-uploads on tagged release
  • Cross-linked from README, FAQ, research notes, latest.md
  • Licensing gate: review each fixture for redistribution rights before upload

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkCorpus, metrics, calibration, or quality gate workecosystemPackage, dataset, marketplace, or third-party surface workenhancementNew feature or requestpriority: lowTriage: longer-horizon research or large build

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions