Deterministic evaluator toolkit for QGIS and geospatial AI trainer workflows.
This repo is built for GIS AI trainer/evaluator work: writing geospatial prompts, defining gold-standard expectations, spotting flawed AI reasoning, and giving clear reviewer feedback. It pairs with the companion QGIS asset project:
https://github.com/Taz33m/qgis-ai-geospatial-assets
The companion asset repo provides the actual Lower Manhattan spatial data surface used for realistic evaluation examples.
Many AI answers to GIS questions sound confident while skipping the details that matter in production: CRS choice, geometry repair, source licensing, topology, export formats, uncertainty flags, and human review. This project turns those requirements into a small evaluation system.
It is not a benchmark of model intelligence. It is a reviewer toolkit for deciding whether an AI answer is useful, incomplete, risky, or wrong in real QGIS workflows.
An AI answer says to export NYC road data directly to GeoJSON without checking CRS, geometry validity, topology, source/license constraints, or review flags. Grade the answer and identify the missing GIS review steps.
Expected review focus: the answer should mention working CRS vs export CRS, geometry validation, source/license attribution, schema/provenance preservation, and whether features need human review before being treated as ground truth.
- A structured task bank with QGIS/GIS prompts, expected answer points, red flags, and scoring focus.
- Expected answer points and red flags for each task.
- A scoring rubric for CRS, geometry QA, topology, schema/provenance, exports, and AI response review.
- A failure taxonomy for common GIS AI mistakes.
- Reviewer feedback examples showing how to critique weak answers.
- A deterministic grading CLI for sample responses.
- A QGIS-ready GeoPackage and
.qgzproject so tasks can be inspected in the actual QGIS application. - Unit tests for the scoring behavior.
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
python -m gis_ai_evaluation_lab grade --tasks data/task_bank.json --responses data/sample_ai_responses.json
python -m unittest discover -s testsNo API keys are required. The grader is intentionally deterministic so the repo can be reviewed, tested, and run offline.
Open this file in QGIS:
qgis/gis_ai_evaluation_lab.qgz
It loads:
Evaluation Task Anchors: 25 spatial task markers grouped by evaluation category.Evaluation Tasks: prompt/category/max-score table.Expected Concepts: gold-standard concepts and keyword evidence.Red Flags: risky answer patterns reviewers should penalize.Sample AI Responses: example answers for grading demonstrations.
Regenerate the QGIS package with:
export PROJ_LIB=/Applications/QGIS.app/Contents/Resources/qgis/proj
export PROJ_DATA=/Applications/QGIS.app/Contents/Resources/qgis/proj
export GDAL_DATA=/Applications/QGIS.app/Contents/Resources/qgis/gdal
export QGIS_PREFIX_PATH=/Applications/QGIS.app/Contents/MacOS
/Applications/QGIS.app/Contents/MacOS/python scripts/build_qgis_package.pypython -m gis_ai_evaluation_lab grade \
--tasks data/task_bank.json \
--responses data/sample_ai_responses.json \
--format markdownThe CLI emits a concise reviewer report with:
- score
- pass/fail band
- matched expected concepts
- triggered red flags
- feedback notes
data/
processed/gis_ai_evaluation_lab.gpkg
task_bank.json
sample_ai_responses.json
docs/
evaluation_rubric.md
failure_taxonomy.md
reviewer_feedback_examples.md
src/gis_ai_evaluation_lab/
evaluator.py
cli.py
qgis/
gis_ai_evaluation_lab.qgz
tests/
test_evaluator.py
Good GIS answers should be:
- Spatially grounded: mention CRS, measurement implications, geometry type, and layer context when relevant.
- Operational: describe concrete QGIS tools or workflow steps, not vague advice.
- Source-aware: distinguish official data, OSM-derived data, and manual interpretation.
- QA-oriented: validate geometry/topology and document uncertainty.
- Reviewer-friendly: call out assumptions, limits, and what should be checked before treating data as ground truth.
This is intended as the second flagship project in a three-project portfolio:
- AI-Ready Lower Manhattan QGIS Portfolio: production-style GIS asset package.
- GIS AI Evaluation Lab: AI answer evaluation toolkit for QGIS/GIS workflows.
- RPInSight: tangential product showing geospatial app UX, campus search, and Mapbox integration.