Skip to content

Optional EvalPort interop for evaluate_classifications/evaluate_detections results #8321

Description

@adhabnr-ux

Hi FiftyOne team — I maintain EvalPort, an open, framework-agnostic JSON spec for portable LLM eval datasets and results (a TestCase/Suite/ResultSet schema with a validator, so a dataset or a graded run can move between tools without hand-writing a converter each time). Filing this as an issue first per CONTRIBUTING.md before writing any code.

I installed fiftyone (1.21.0) and ran the real evaluation API rather than guessing:

results = dataset.evaluate_classifications("predictions", gt_field="ground_truth", eval_key="eval_simple")
dataset.first()["eval_simple"]   # -> True/False, written onto every Sample
results.report()                 # -> {"cat": {"precision":..., "recall":..., "f1-score":..., "support":...}, ..., "accuracy": ...}

evaluate_classifications/evaluate_detections write a per-Sample pass/fail field and return a ClassificationResults/DetectionResults object (with .report(), .metrics(), .confusion_matrix(), .to_str()/.from_str(), .save()) holding the aggregate metrics. That maps onto EvalPort's two halves directly: each Sample is a TestCase, the per-sample eval-key boolean is a Grader pass/fail, and results.report()'s per-class precision/recall/f1/support is exactly the aggregate shape of an EvalPort ResultSet.

Two ways I could see this landing, and I don't have a strong preference:

  1. A standalone fiftyone-openeval-adapter package in the EvalPort repo, depending on fiftyone as a normal dependency. Zero footprint on this repo.
  2. A small optional module inside this repo if you'd rather it live here.

Either way, real tests would validate against EvalPort's actual JSON Schema, not a mock.

Let me know which direction you'd prefer, or if this isn't a fit for your roadmap right now — no worries either way.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions