Hi FiftyOne team — I maintain EvalPort, an open, framework-agnostic JSON spec for portable LLM eval datasets and results (a TestCase/Suite/ResultSet schema with a validator, so a dataset or a graded run can move between tools without hand-writing a converter each time). Filing this as an issue first per CONTRIBUTING.md before writing any code.
I installed fiftyone (1.21.0) and ran the real evaluation API rather than guessing:
results = dataset.evaluate_classifications("predictions", gt_field="ground_truth", eval_key="eval_simple")
dataset.first()["eval_simple"] # -> True/False, written onto every Sample
results.report() # -> {"cat": {"precision":..., "recall":..., "f1-score":..., "support":...}, ..., "accuracy": ...}
evaluate_classifications/evaluate_detections write a per-Sample pass/fail field and return a ClassificationResults/DetectionResults object (with .report(), .metrics(), .confusion_matrix(), .to_str()/.from_str(), .save()) holding the aggregate metrics. That maps onto EvalPort's two halves directly: each Sample is a TestCase, the per-sample eval-key boolean is a Grader pass/fail, and results.report()'s per-class precision/recall/f1/support is exactly the aggregate shape of an EvalPort ResultSet.
Two ways I could see this landing, and I don't have a strong preference:
- A standalone
fiftyone-openeval-adapter package in the EvalPort repo, depending on fiftyone as a normal dependency. Zero footprint on this repo.
- A small optional module inside this repo if you'd rather it live here.
Either way, real tests would validate against EvalPort's actual JSON Schema, not a mock.
Let me know which direction you'd prefer, or if this isn't a fit for your roadmap right now — no worries either way.
Hi FiftyOne team — I maintain EvalPort, an open, framework-agnostic JSON spec for portable LLM eval datasets and results (a
TestCase/Suite/ResultSetschema with a validator, so a dataset or a graded run can move between tools without hand-writing a converter each time). Filing this as an issue first per CONTRIBUTING.md before writing any code.I installed
fiftyone(1.21.0) and ran the real evaluation API rather than guessing:evaluate_classifications/evaluate_detectionswrite a per-Samplepass/fail field and return aClassificationResults/DetectionResultsobject (with.report(),.metrics(),.confusion_matrix(),.to_str()/.from_str(),.save()) holding the aggregate metrics. That maps onto EvalPort's two halves directly: eachSampleis aTestCase, the per-sample eval-key boolean is aGraderpass/fail, andresults.report()'s per-class precision/recall/f1/support is exactly the aggregate shape of an EvalPortResultSet.Two ways I could see this landing, and I don't have a strong preference:
fiftyone-openeval-adapterpackage in the EvalPort repo, depending onfiftyoneas a normal dependency. Zero footprint on this repo.Either way, real tests would validate against EvalPort's actual JSON Schema, not a mock.
Let me know which direction you'd prefer, or if this isn't a fit for your roadmap right now — no worries either way.