Skip to content

Repository files navigation

GeoLens

GeoLens is an experimental spatial evidence engine. It combines real environmental observations, terrain and infrastructure to derive physical results that remain traceable to their sources.

GeoLens is being rebuilt around a simple rule:

A result is useful only when we can explain where it came from, how it was transformed and what is still unknown.

GeoLens institutional interface showing the spatial evidence mission and programme status

The public GeoLens interface presents the research programme, its current proof and the rule that missing evidence must never become a valid-looking zero.

The first application is water moving from rainfall, across the surface and, where authoritative relationships are available, into a stormwater network.

GeoLens in one minute

Rainfall, terrain, land cover and drainage infrastructure are usually published by different organisations, at different resolutions and in different formats. A map can make these layers look connected even when the underlying relationships have not been demonstrated.

GeoLens tries to build that connection explicitly:

  1. acquire real rainfall, terrain and land-cover evidence for a bounded place and time;
  2. retain the provider, dataset version, timestamps and original resolution;
  3. derive an inspectable runoff quantity;
  4. aggregate it over a defined surface or catchment;
  5. connect it to observed infrastructure only when a source-backed attachment exists;
  6. propagate it through known network directions without inventing missing links;
  7. return the result together with its provenance, assumptions and missing-data state.

If an input is unavailable, GeoLens reports it as unavailable. It does not silently turn missing rainfall, elevation or land cover into zero.

GeoLens is not currently a flood forecast, a hydraulic sewer simulator or a generic risk-score dashboard.

What we are trying to prove

The refoundation has one central question:

Can real environmental evidence be transformed into an inspectable physical state without hiding uncertainty or inventing data?

For stormwater, the target chain is:

real rainfall + real terrain + real land cover
                       |
                       v
             spatial evidence bundle
                       |
                       v
             inspectable runoff model
                       |
                       v
       contributing surface or catchment
                       |
                       v
          observed stormwater topology
                       |
                       v
          known downstream network state
                       |
                       v
        result + provenance + missing state

H3 is used to connect evidence spatially. It is an indexing and representation choice, not a claim that the original datasets have H3-native precision. The source resolution of IMERG, GLO-30, CLC, AHN and BGT remains visible.

Where the project is today

GeoLens has two complementary operational proofs and two historical benchmark programmes.

Case Plain-language outcome What is proven What remains unresolved
Trento Proof 0 The complete software chain works from real environmental inputs to a downstream result Evidence composition, deterministic runoff, catchment aggregation, network propagation, provenance and mass balance The small drainage network is a deterministic fixture, not surveyed municipal infrastructure
Amsterdam observed proof GeoLens can read real Waternet pipes and nodes and derive a real, non-zero surface runoff source Observed topology, elevation-based direction states, real rainfall/terrain/land cover and an inspectable surface contribution No owner-published surface-to-pipe attachment has yet been found, so sewer propagation is deliberately blocked
Emilia-Romagna 2023 A simple terrain-only concentration hypothesis was tested against an independent observed flood extent and did not perform better than chance A reproducible historical benchmark, withheld evaluation data and honest negative evidence A conditioned replay requires discharge, boundary, breach and terrain/channel evidence that is not currently available
Cumbria 2015 Nine frozen Storm Desmond scenarios were tested once against three independently opened flood references and showed low agreement with substantial overprediction A restart-safe real-data run, mass balance, frozen prediction, blind evaluation, diagnostic map and retained negative result Four physical hypotheses remain evidence-blocked; Carlisle cannot be retuned or rerun for validation, and a new holdout is required for any future validation claim

Why development is deliberately constrained now

GeoLens is not paused. Its expansion is frozen while two independent external-evidence paths remain open: owner-supplied Waternet attachment evidence for Amsterdam and an Environment Agency model package for Carlisle.

The annotated Git tag pre-external-evidence-baseline-v1 identifies commit 938b18fb66925e36236ea04a49eefdb2ca9826cb. It records what GeoLens was before either requested package arrived. The tag is intentionally separate from the public release line.

The declared Cumbria execution and blind-evaluation gates are complete. Work can continue on package integrity checks, minimal format/CRS/unit/schema adapters, tests, security, reproducibility and interface inspection. Scientific revisions remain blocked until event-valid evidence and an independently merged deterministic fixture exist.

What cannot happen is equally important. GeoLens will not add a new country or benchmark merely because data are available, anticipate the contents of an agency delivery, or select parameters and thresholds after seeing the expected answer. If the frozen baseline performs poorly, that result remains part of the record. A scientifically revised model may follow only as a separately versioned experiment shown beside the original result.

GeoLens Case 02 spatial inspector showing the Emilia-Romagna event runoff layer and explicit withheld evidence

Case 02 exposes the derived event-runoff concentration together with native resolution, transformation and publication state. Restricted or unavailable spatial evidence remains visibly withheld instead of being rendered as zero.

This distinction matters. GeoLens does not describe a software pipeline as scientifically validated merely because it runs.

Results so far

Case 00 — Trento Proof 0

The verified live window observed 9.24 mm of rainfall. The model derived 2.957 m³ of runoff contribution and delivered the same volume to the fixture outfall with zero mass-balance difference.

In simple terms: the complete transformation chain is operational and numerically inspectable.

The environmental evidence is real. The network geometry is a small deterministic test fixture, so this result does not claim to represent the real Trento drainage system.

Case 01 — Amsterdam urban drainage proof

GeoLens currently exposes:

  • 47 observed Waternet nodes;
  • 47 active stormwater pipes;
  • 4 explicitly typed rainwater outfalls;
  • 26 known and 21 ambiguous pipe directions;
  • one known upstream path containing 5 nodes and 4 pipes;
  • 696 H3 r13 surface cells classified from BGT and conditioned with AHN4 terrain;
  • 100 contributing cells representing 3,676.73 m²;
  • 3.835 mm of real IMERG rainfall for the selected window;
  • 11.4145 m³ of derived surface runoff.

The 11.4145 m³ value is a traceable environmental source term. GeoLens does not present it as observed sewer inflow because the authoritative relationship between the contributing surface and an exact Waternet asset is missing.

The API therefore keeps propagation blocked. The experimental BGT/AHN outlet remains labelled as conditioned and not observed.

Case 02 — Emilia-Romagna 2023 historical benchmark

The Forlì benchmark reconstructs the 16–18 May 2023 event using inputs that are kept separate from the official post-event flood extent.

The model uses:

  • all 96 expected IMERG Final Run V07 half-hour granules;
  • a bounded native rainfall grid with a mean 48-hour accumulation of 93.982 mm;
  • real GLO-30 elevation and slope;
  • real CLC land-cover classes;
  • official DBTR geometry for water, wet areas, riverbeds, embankments and buildings;
  • a frozen 30 m evaluation grid containing 130,307 eligible cells.

The first runoff-and-routing baseline derived 6,176,691.50 m³ over 129,841 source cells and conserved that volume to floating-point precision.

Only after the prediction protocol was frozen did GeoLens compare the result with the independent regional flood extent. The terrain-only concentration score returned:

  • ROC AUC: 0.491624;
  • average precision: 0.277679;
  • observed flooded-cell prevalence: 0.286815.

That is near-random discrimination. The result is retained as useful negative evidence: raw GLO-30 D8 concentration without depression conditioning, river stage, discharge, breach behaviour, embankment hydraulics or downstream boundary conditions does not reconstruct the observed flood footprint.

GeoLens makes no inundation-depth, probability or operational-forecast claim from this result.

Case 03 — Cumbria 2015 protocol-qualified replay

The Carlisle candidate is no longer just a list of promising portals. GeoLens has frozen a public-only downstream-reach domain, a fail-closed hydraulic input protocol, an explicit experimental replacement-solver contract and a sealed blind-evaluation boundary for Storm Desmond from 2015-12-04T00:00:00Z to 2015-12-07T00:00:00Z. The Environment Agency model request is useful, but it is no longer on the critical path.

The audit verifies:

  • all 144 expected NASA IMERG V07 Final Run half-hour granules acquired through the canonical Python earthaccess + xarray path and reduced to a content-addressed 3 × 4 native-grid accumulation outside Git;
  • all 288 expected 15-minute qualified flow observations and all 288 level observations at Sheepmount on the River Eden; flow is input-only for the public baseline, while level remains comparison evidence;
  • all 288 local rainfall observations at Willow Holme, retained as station comparison rather than basin-wide rainfall;
  • four complete candidate upstream hydrographs: 288 qualified flow values each at Great Corby on the Eden, Greenholme on the Irthing, Cummersdale on the Caldew and Newbiggin Bridge on the Petteril;
  • a local protocol envelope containing those four stations, with British National Grid required for future solver geometry and Ordnance Datum Newlyn required for vertical evidence;
  • the four wider-domain hydrographs remain native-only boundary candidates; for the public Sheepmount reach, the replacement contract separately fixes a left-constant 15-minute transform with no gap filling and a versioned positive-excess discharge proxy;
  • rainfall/runoff forcing restricted to the future local domain downstream of those inflows, so upstream catchments already represented by the hydrographs cannot be counted twice;
  • a content-addressed Environment Agency Recorded Flood Outline for event group 4175 and both official Copernicus EMSR147 Carlisle vector products, acquired only after prediction freeze and structurally restricted to evaluation;
  • a content-addressed blind-hindcast protocol that keeps both observed geometries sealed until the prediction artifact, code revision, physical wetness criterion and evaluation domain are frozen;
  • official time-stamped Environment Agency LiDAR metadata: 550 source records over 241 intersecting OS grid references, with a deterministic selection of 231 pre-event records;
  • all 231 selected 1 km records mapped to 30 official time-stamped DTM ZIP identities through the current survey-search contract, with source-to-archive and archive-inventory SHA-256 identities;
  • a header-only probe of the mapped 2009 / 1 m / NY3555 identity returned application/zip and filename lidar_tiles_dtm-2009-1-NY35ne.zip while reading and storing zero archive bytes;
  • a bounded DTM materialization path that downloaded six official pre-event archives, content-addressed every archive and source raster, rejected unsafe ZIP entries, preserved native British National Grid resolution and wrote only explicitly georeferenced 1 km windows outside Git;
  • ten effectively missing pre-event terrain grid references after archive inspection: NY3256, NY3257, NY3258, NY3259, NY3357, NY3358, NY3359, NY3859, NY3959 and NY3960;
  • 16 event-valid WFD Cycle 1 river-context features intersecting the audit AOI, pinned independently from current OS Open Rivers;
  • the 2011 official main report and appendices: the historical model began in 1999, was expanded with surveyed sections in 2003, calibrated against January 2005, and ran from four named upstream watercourse limits to Old Sandsfield; the reports expose none of the runnable cross-section, model or boundary files;
  • 19 bounded Environment Agency Flood Model Locations records and six exact Carlisle model-group identities; pre-event groups 1313, 1314, 1797 and 8323 are request lineage only, while groups 2039 and 9458 are post-event and excluded from input and calibration;
  • an Environment Agency Products 5, 6 and 7 request sent on 2 September 2026 for the four pre-event groups, covering reports, outputs, native model inputs, survey sections, boundary definitions, roughness, defence state, software, datum and reuse conditions; Product 4 and the two post-event model groups remain deliberately excluded, and the response is an optional comparison rather than an open-data replay dependency;
  • a fail-closed cumbria-model intake contract that content-addresses any future delivery outside Git, inventories ten requested components and requires a separate scientific review before any component can become a candidate for physical-gate assessment;
  • 291 current AIMS defence records in the bounded query, pinned as current context only: 114 have no asset start date, 56 start on or after the event and four report refurbishment after 2015;
  • 349 current AIMS Channel records, also context only: 272 have no start date, 17 dated records start on or after the event, and the schema contains no hydraulic cross-sections, bed levels or roughness;
  • a bounded CLC 2012 window on the official aligned 100 m EPSG:3035 grid: all 7,553 cells are available across 12 observed classes, with the 2011–2012 reference period and retrospective corrected-release identity both retained;
  • an input-selected 56 km² British National Grid domain from Sheepmount to the documented Old Sandsfield limit, fixed without loading either observed flood extent;
  • six official pre-event 1 m DTM archives totalling 280,161,858 bytes and 40 unique source GeoTIFFs, all retained outside Git with content hashes;
  • 48 georeferenced 1 km windows from the 52 catalogue-selected references, including six explicit fallbacks to older pre-event source rasters already present in the selected archive set; 46 windows contain observed terrain and two contain only provider-declared NoData;
  • 10 effectively missing 1 km references after real archive inspection: four absent from the catalogue selection, four absent from the selected and already materialized older archive rasters, and two represented only by NoData pixels;
  • 30,678,994 valid elevation pixels and 17,321,006 NoData pixels across the 48 mapped windows, compressed to 36,148,351 physical bytes without replacing missing values with zero; these terrain artifacts remained execution-blocked until the later clean-revision authorization.
  • a complete 72-hour IMERG accumulation over 12 finite native cells, ranging from 81.705 to 115.900 mm with a 102.971 mm cell mean; this remains coarse 0.1-degree evidence and is not sharpened by H3.
  • a reproducible real-source composition for H3 cell 8a1955d9535ffff, selected solely as the cell containing the frozen public-domain centre: 13,528 intersecting 1 m DTM pixels produce a mean elevation of 8.097 m, native CLC class 231 covers the cell, and two coarse IMERG footprints produce an area-weighted 72-hour accumulation of 98.784 mm. This is an inspection result, not routing or hydraulic state.
  • replacement-solver protocol cumbria-public-surface-flow-replacement-v0, pinned as SHA-256 b9db1fbc10cc0aeff5d3a4e24bf1afc5d66226902ecad92ce111f1fcfb60c89b: a primary 20 m grid, predeclared 10 m and 40 m resolution sensitivities, adaptive CFL integration, explicit CLC runoff/Manning ranges, a dry-surface initial assumption, free outer outflow, nine mandatory scenarios, mass-balance accounting and a primary 0.05 m maximum-depth wetness criterion. It authorizes neither a run nor evaluation access.
  • three content-addressed static computation grids materialized directly from the native DTM and CLC footprints, without H3 routing or categorical interpolation. On the primary 20 m mesh, 75,786 of 140,000 cells have complete terrain coverage, 64,214 remain unavailable, and a one-cell missing-data halo leaves 73,502 cells eligible for a future prediction. CLC covers every cell. The 10 m and 40 m sensitivity grids retain the same fail-closed rules.
  • content-addressed time-varying forcing for the exact 72-hour window: 144 native half-hour IMERG amount grids reproduce the pinned accumulation exactly, and 288 Sheepmount 15-minute observations retain both measured flow and the frozen positive-excess proxy. No interval is interpolated, gap-filled or replaced with zero, and no solver or evaluation run occurred.
  • a content-addressed event-input binding for all nine frozen scenarios. Exact projected native-footprint overlap maps IMERG onto every one of the 399,523 valid cells across the three meshes; explicit 900-second indices pair each Sheepmount sample with the correct half-hour rainfall amount, and frozen river footprints distribute discharge only by intersected valid area. The binding created no prediction and opened no evaluation geometry.
  • a frozen event-runner contract that fixes clean-revision authorization, lazy one-interval forcing construction, streaming maximum-depth aggregation, exact 900-second outputs, adaptive-CFL bounds, per-step and final mass balance, the all-nine-scenarios completion rule, NoData encodings and physical wetness thresholds. Runner v0.2 recovered from the earlier operational interruptions by writing an authorization-bound checkpoint after every complete scenario. The final run completed all nine predeclared scenarios with 288 outputs and a passing mass balance each, then created prediction receipt SHA-256 f2a3a7489699a70a6d5c770633bdc9f789184ca26cb8190c85a1d28d40495dc6 while both observed flood references remained sealed.

This proves data access and narrows the missing physics; it is not yet a flood reconstruction. The four upstream series now have fixed identities, units, windows and sampling semantics, but their station coordinates are not asserted to be historical model cross-sections. Old Sandsfield (NY332617) is retained as the documented downstream model limit, not as an observed boundary: the nearby station search finds no observation at that limit and the historical boundary values remain missing. Sheepmount level remains an observation for comparison, not a downstream boundary. A separately screened station named Rockcliffe publishes a qualified groundwater-dip measure rather than a surface-water boundary, so it was explicitly rejected.

The public replacement makes its compromises visible instead of filling those gaps silently. It represents only 2D surface flow over available terrain, omits unavailable event-valid channel sections, defences and floodgates, treats the outer grid as free outflow without an imposed stage, starts the surface dry as a model assumption, and injects only the positive Sheepmount discharge above the first window sample through a declared footprint. The three footprint sizes, three grid resolutions and runoff/roughness ranges are fixed before either observed extent is opened; all nine scenarios must be reported, so no best-looking result can be selected afterwards.

The WFD layer contains only designated 1:50,000 river stretches. The current AIMS defence and channel inventories are updated daily, so present geometry, crest, condition and channel lines cannot masquerade as the December 2015 hydraulic state even when an asset has a pre-event start date. The official 2011 documents and Flood Model Locations catalogue now tell us which historical domain and model groups to request, but they do not provide the cited ISIS/TUFLOW package, cross-sections, roughness or boundary files. The 2015 investigation report is post-event context: it reports overtopping and bypass, with no defence breach, but no narrative location becomes model geometry.

The selected LiDAR catalogue rows and their download mapping are now pinned independently. The earlier DTM / 2009 / 1M / NY3957 probe was invalid: it used a legacy route and a 1 km source reference where the active delivery service expects product id lidar_tiles_dtm, numeric resolution 1 and the containing 5 km tile NY3555. The current bounded search returns 590 survey products, including 123 time-stamped DTM identities; deterministic source-year and 5 km containment mapping resolves the 231 selected source rows to 30 archives. Two identities are labelled 2015, but only their individually selected 1 km source areas have exact pre-event survey dates. GeoLens must therefore mask every archive back to its mapped 1 km references and may not treat the whole 5 km ZIP as event-valid.

Manifest v0.33.0 preserves the original content-addressed public-baseline protocol, records what each acquisition and the first real composition contained, and links the canonically hashed replacement-solver contract to corrected static grids, time-varying forcing, numerical-kernel verification, event-input binding, restart-safe execution, the frozen prediction, independent evaluation-reference acquisition, the single completed blind evaluation, its non-evaluative diagnostic map, a descriptive failure inventory and a post-evaluation physics-revision gate. The prediction retains its source manifest v0.25.0 byte identity, commit and tree, so later gates cannot rewrite the run retrospectively. The 8 km × 7 km British National Grid domain was selected from Sheepmount and the documented Old Sandsfield limit before either observed flood extent was opened. The six ZIP archives occupy 280,161,858 bytes. Their 40 unique source GeoTIFFs yielded 48 exact one-kilometre windows: 46 carry at least one real elevation sample, while NY3859 and NY3960 contain only the provider's NoData value. Four further catalogue-selected cells have no covering raster in the selected or already materialized older archive for the same 5 km tile. Together with the four original catalogue gaps, the effective missing set contains ten cells. Six windows use fully traceable older pre-event source rasters where the newer selected archive lacked usable coverage. The retained Float32 windows decode to 192,000,000 bytes and are stored as 36,148,351 bytes of deterministic gzip-compressed artifacts outside Git. CLC and IMERG add only bounded, native-grid artifacts and receipts; the original NASA granules are not copied. Missing pixels remain missing.

The corrected static-grid receipt is SHA-256 63fa37941c13cc6a18bb3fed9f9ab1690ae7da34da9963cead2299adc879800c. It identifies 36 unique compressed artifacts occupying 2,342,503 bytes and representing 25,725,000 decoded bytes. During binding, fail-closed validation found that the superseded receipt had valid CLC masks but NoData runoff and Manning arrays because its parameter accumulators began at NaN. Version 0.2.0 initializes those accumulators correctly and proves finite parameters on every land-cover-valid cell. The old receipt remains outside Git for traceability but is forbidden as solver input. A solver cell still receives elevation only when its entire native DTM footprint is available; otherwise it remains invalid.

The forcing receipt is SHA-256 782eb107578fe34bccba9bc0202d90d06d504ddd63bc03ba301b1ccb3781c38b. It retains the native 0.1-degree IMERG source resolution and all 144 half-hour amounts; their Float32 sum matches the earlier 72-hour accumulation with a maximum difference of exactly 0 mm. Sheepmount contributes 288 observed 15-minute values, kept separately from the versioned positive-excess transformation above the first 422.716 m3/s sample. The resulting excess volume is 118,482,930.9 m3. These are forcing artifacts, not a flood prediction.

The Python/NumPy kernel cumbria-local-inertial-surface-flow-v0.2.0 now passes a content-addressed, event-isolated fixture suite. The numerical formulation is unchanged; v0.2 adds lazy forcing and streamed output observation so a 10 m event run does not retain the full forcing cube or 288 depth snapshots in memory. The tests cover lake-at-rest equilibrium over variable terrain, exact closed-cell rainfall storage, non-negative sloped free outflow, exact forcing/output time boundaries, streamed maximum-depth output and mandatory failure when CFL stability would require a timestep below 0.05 s. Missing forcing is rejected. The fixture result SHA-256 is b7ac171c9b28ab6bf69ff6cb8d3c43d07c102444483d8e883de82739fcd6a423; the suite reads no Storm Desmond inputs or observed flood geometry, performs no network request or external write and does not create a prediction.

The current event-input binding receipt is SHA-256 502a1ecad80e0f1877f853967524c13a54d4ce52cddfbefc149873f8591134d4. Its 19 unique artifacts occupy 1,097,674 compressed bytes and represent 6,555,387 decoded bytes. It supersedes the v0.1 receipt only to bind the verified v0.2 kernel identities; the persisted mapping artifacts are unchanged and the old receipt remains preserved. CSR mappings retain the 12 native IMERG cell identities and cover every valid solver cell within the frozen tolerance; the 60 m, 100 m and 140 m Sheepmount footprint variants retain their exact valid-area weights. npm run materialize:cumbria-event-input-binding -- --data-root <external-directory> --check verifies the grids, masks, parameter ranges, forcing arrays, units, timing, mappings and receipt with zero writes. It still authorizes zero event runs and zero evaluation access.

Event-runner contract SHA-256 b9f2b9d9f91557b4cd62f35875fa2b16d3ebf8fe5e536416cf2d9bcb0dae23b9 fixes the execution and prediction semantics independently of process lifetime. The completed run is bound to merge commit 1cab9bfccbfecd0818ed7f8e7b8735bd4b437bb1, source manifest SHA-256 f43632afaca8d9a86a4c3a7cdb955519b6ca89f170f27dcd0f772f75bcf19fab and authorization SHA-256 7feda0564923260c2996f292ec4ae82b52a6cd7a17b32aaca805a90816a4ab44. Each scenario checkpoint was re-read against its output, stability, mass-balance and compressed/decoded artifact identities. The final receipt contains all nine scenarios and 72 output descriptors. No partial checkpoint became a prediction, no best scenario was selected and no evaluation geometry was opened.

The same manifest now draws a hard spatial boundary between evidence and simulation. Terrain remains on its native 0.5 m, 1 m or 2 m British National Grid cells; CLC 2012 remains a 100 m categorical product in EPSG:3035; IMERG remains an approximately 0.1-degree, 30-minute observation in EPSG:4326. None is silently resampled onto another source's grid. EPSG:27700 is the declared exchange coordinate frame for overlap calculations, not a common raster grid.

H3 resolution 10 provides a reproducible catalogue and inspection index of 24,230 cells over the frozen hydraulic-protocol envelope, with an approximate mean cell area of 13,199 m2 near Carlisle. Each index value retains the source resolution and overlap semantics that produced it: terrain coverage and NoData statistics, CLC area fractions and IMERG native-cell overlap. H3 does not sharpen IMERG, retain sub-metre terrain detail, route water or store hydraulic state. The separate replacement contract now fixes a 400 × 350 primary grid at 20 m and 10 m / 40 m sensitivity grids in EPSG:27700; these are computation choices, not new source resolution claims.

The generic spatial-evidence-index-v0.1.0 composer now makes that boundary executable. It accepts native source-footprint intersections rather than H3-labelled source measurements, requires an explicit area CRS and measurement method, preserves source identities, versions, acquisition times and resolutions, and rejects overlapping footprints or inconsistent precipitation windows. For Cumbria, coverage fractions use the H3 boundary projected into EPSG:27700 as their denominator; they are not compared against the different spherical area returned by the H3 catalogue library. Complete coverage yields terrain statistics, CLC area fractions and area-weighted rainfall. Partial or unavailable coverage yields null evidence plus explicit coverage diagnostics—never a partial valid-looking value. Observed 0 mm remains zero. The synthetic single-cell fixture still verifies transformation semantics, and a separate 5.65 MB decoded / 96 KB compressed real-evidence artifact proves the same contract against byte-verified DTM, CLC and IMERG inputs. Its receipt and result hashes reproduce exactly. It covers one inspection cell only and does not create a solver state.

The blind evaluation boundary is frozen around the completed prediction under protocol SHA-256 1a135785bef1121e542952fd8ee90d6eed86908d19864d381985fbfd2f8a1dd0. It pins the primary primary-20m maximum-depth mask at >= 0.05 m, its compressed and decoded identities, and the 73,502-cell valid-prediction mask on the 400 × 350 EPSG:27700 grid. Six metrics remain fixed—intersection over union, area precision, area recall, false-positive area, false-negative area and symmetric 95th-percentile boundary distance. Missing observed coverage is excluded and reported rather than treated as dry; missing prediction coverage blocks evaluation. Their geometry cannot select the domain, mesh, threshold or model parameters. npm run verify:cumbria-blind-evaluation-protocol -- --data-root <external-directory> still verifies the historical source commit and manifest, final receipt, all nine checkpoints and all 72 compressed/decoded prediction artifacts independently from the later reference acquisition.

The single authorized blind evaluation is now complete and retained under external receipt SHA-256 610bae9d7e31978e3565a5e04e779bde9e1dc5236e200a0699cf8db9c118ffa2. It is a negative baseline, not a validation success. The frozen prediction marks 14,357 cells (5.743 km²) wet. Against the Environment Agency recorded outline, IoU is 0.0547, precision 0.0585 and recall 0.4536; against EMSR147 initial they are 0.0140, 0.0141 and 0.5833; against EMSR147 monitoring-01 they are 0.0303, 0.0327 and 0.2982. False-positive area remains between 5.407 and 5.662 km² in all three comparisons. In plain language: the model reaches some genuinely flooded places, but spreads water across far too much of the domain and places much of it poorly. No threshold, scenario or parameter was changed after opening the references. Any improved physics must be a separately versioned experiment shown beside this result, never instead of it. npm run evaluate:cumbria-blind-prediction -- --data-root <external-directory> --check verifies the recorded receipt without recomputing the metrics.

Cumbria blind-evaluation spatial error map

The diagnostic map makes the failure spatially inspectable: orange false positives dominate branching flow paths, blue false negatives identify observed flooding the model missed, teal marks limited agreement and dark grey is outside the frozen evaluation domain because required prediction evidence was unavailable—not observed dry ground. It is derived only from the already frozen masks under diagnostic receipt SHA-256 edccd01df0467fb6e3a9d0302ae4869f1d46269aa48901a866f854ca7d0dc721; it performs zero evaluation runs, recomputes no metric and changes neither threshold nor model. The publication PNG is content-addressed as 4d374162fed93626ec33b6e77396f9378a78759cbad42dbc28e00b70cd10fc82.

Failure-inventory receipt SHA-256 73d4eb08f14e7fb6e7fcee7c7ebda9a2bc768f521979076f77d2fc42dc00d508 stratifies the same frozen cells using only depth thresholds declared before reference access and the pre-event CLC 2012 class already present on the 20 m input grid. The finding is stronger than “the 5 cm threshold is too permissive”: 78.4–79.3% of false-positive cells remain at or above 10 cm, and 34.0–36.7% remain at or above 30 cm across all three references. For the Environment Agency comparison the largest false-positive counts occur in pastures (5,890 cells), non-irrigated arable land (3,334), salt marshes (1,226) and discontinuous urban fabric (1,124). These are descriptive counts, not causal attribution: they neither rank scenarios nor authorize retuning. They show that a threshold-only correction cannot explain the retained failure and keep missing channel conveyance, boundary conditions and 2015 defence state as live physical hypotheses.

The next-model boundary is now explicit. Four hypotheses remain unconfirmed: channel conveyance, downstream/initial hydraulic state, December 2015 defences and controls, and source-term placement. None authorizes code changes or a Carlisle rerun until its event-valid evidence is content-addressed and an independent deterministic fixture is merged first. Because the Carlisle flood outlines are now open, they may support failure diagnosis and transparent non-blind comparison only—not calibration, threshold/scenario selection, acceptance or a new blind-validation claim. Any future validation claim requires a new holdout event or area, and the original negative baseline must remain visible beside every revision.

After merge commit df33838ba7774a736ebe17aec7e6c9aee01e1827, the evaluation references were opened and content-addressed under receipt SHA-256 b9bf772af4a356de533edb26c005dd318e36d889ad44fdaeb51586c989adffbf. The fixed domain and provider event group 4175 select exactly Environment Agency feature Recorded_Flood_Outlines.26355 (53,114 bytes, SHA-256 e2ad395a39441a077cf3585e3cce09a49454922798dba483a31fe71fabef78c0). The official EMSR147 initial and monitoring archives retain separate SHA-256 identities 5524efba986082b901bf27fc7ecde0cf6af91fea393b36bfa6bcf6a959448049 and b58b6e047d8e1066fffcf3aa09fafa4b975cd7eedccda22d8df8c4971d509cfb. Their crisis layers contain 228 initial polygons dated 7 December and 495 monitoring records dated 7 or 10 December. Each source was independently normalized, rasterized and scored; the three comparisons remain separate.

The Environment Agency delivery for request EIR2026/42104 was received on 18 September 2026 and registered outside Git as one 46,688,104,848-byte archive with SHA-256 40ae1d7df10b105b8abb195c43b53e9576683a654e781686406f286cee5da145. Its four top-level archives contain historical Carlisle survey material, the 2007 Eden ABD extract, the 2010 Carlisle FAS package and the 2012 Cumbria Tidal package. A central-directory-only inventory found native TUFLOW/ISIS material, reports, outputs, cross-sections, terrain and survey files, with no encrypted files, unsafe paths or duplicate normalized paths. This establishes receipt, not scientific usability.

The shared intake command accepts --kind cumbria-model, resolves artifacts only beneath an explicit external data root, recomputes byte length and SHA-256, leaves originals in place, writes a new non-overwriting receipt and performs no archive extraction. The registered package declares all ten requested components as incomplete or metadata-only. Model-group mapping, pre-event lineage, licence coverage for embedded third-party material, CRS, units, datum and component completeness still require document-level review. Product 4, post-event groups 2039/9458, observed flood geometry and automatic replay promotion remain rejected. Even a fully reviewed package can become only a candidate for a separate physical-gate assessment—not replay-eligible by intake alone.

This removes the static-grid, forcing, isolated-kernel, event-binding, execution, prediction-freeze, reference-identity and blind-comparison blockers without hiding the ten effective terrain gaps or the NoData pixels inside partially covered windows. The result demonstrates that the current public-only physics is not spatially adequate as a Carlisle flood reconstruction. Historical EA material is now present but remains unqualified: it cannot enter the official-model reconstruction until component review establishes event-valid channel geometry, boundary conditions, defence state and software/spatial semantics. That review is an optional comparison or separately versioned physics-upgrade track and cannot rewrite the retained negative baseline.

The deterministic manifest is tests/ground-truth/cumbria-2015/manifest.json. Re-run the open-service checks with:

npm run audit:cumbria-access
npm run audit:cumbria-lidar-catalog
npm run audit:cumbria-hydrography
npm run audit:cumbria-hydraulic-context
npm run audit:cumbria-boundary-protocol
npm run audit:cumbria-hydraulic-domain
npm run audit:cumbria-public-baseline
npm run materialize:cumbria-dtm -- --data-root <external-directory>
npm run materialize:cumbria-dtm-masks -- --data-root <external-directory>
npm run materialize:cumbria-clc2012 -- --data-root <external-directory> --execute
npm run materialize:cumbria-imerg -- --data-root <external-directory> --execute
npm run materialize:cumbria-spatial-evidence -- --data-root <external-directory> --check
npm run materialize:cumbria-solver-grids -- --data-root <external-directory> --check
npm run materialize:cumbria-forcing -- --data-root <external-directory> --check
npm run materialize:cumbria-event-input-binding -- --data-root <external-directory> --check
npm run preflight:cumbria-event-run -- --data-root <external-directory>
npm run prepare:cumbria-model-request
npm run plan:cumbria-dtm-materialization
npm run plan:cumbria-spatial-grid
npm run verify:cumbria-spatial-composition-fixture
npm run verify:cumbria-replacement-solver-protocol
npm run verify:cumbria-local-inertial-kernel
npm run verify:cumbria-blind-evaluation-protocol
npm run evaluate:cumbria-blind-prediction -- --data-root <external-directory> --check
npm run materialize:cumbria-evaluation-diagnostics -- --data-root <external-directory> --publication-output docs/screenshots/cumbria-blind-evaluation.svg --check
npm run materialize:cumbria-failure-inventory -- --data-root <external-directory> --check

The CLC, IMERG, spatial-evidence, solver-grid and forcing commands are dry runs unless --execute is supplied. All reject a data root inside the repository or OneDrive. IMERG reuses the canonical NASA cache, and forcing acquisition additionally checkpoints transformed per-granule amounts outside Git so an interrupted run can resume without discarding completed granules. For the composed cell, static solver grids and forcing, --check rebuilds or verifies the pinned result with zero network requests and zero writes.

When a delivery arrives, prepare a draft package beside the originals and run the intake with explicit paths outside Git:

npm run intake:external-evidence -- --kind cumbria-model --draft <draft.json> --data-root <delivery-directory> --output <new-receipt.json>

For Amsterdam data owners and collaborators

GeoLens already uses public Waternet infrastructure, AHN4 terrain, BGT physical surfaces, PDOK/GWSW context, NASA IMERG rainfall, CORINE Land Cover and Copernicus GLO-30.

The specific missing relationship is not another rainfall or terrain layer. It is authoritative attachment evidence connecting a contributing surface to an observed stormwater destination.

Useful material could include:

  • an Amsterdam owner-published BGT Inlooptabel;
  • a hydraulic-model surface-to-inlet or surface-to-outfall relation;
  • an exact crosswalk between BGT surface identifiers and Waternet assets;
  • documentation of the relevant relationship semantics and identifiers;
  • guidance to the municipal or Waternet team responsible for these data.

GeoLens will not infer an observed sewer attachment from proximity, polygon containment or a conditioned terrain outlet. Those may remain experimental proxies, but they cannot be represented as authoritative infrastructure evidence.

The non-negotiable rules

Missing is not zero

Zero is valid only when it is observed or legitimately derived. Provider failure, missing coverage and incomplete time windows remain explicit states.

Real evidence comes before interpretation

Important values retain their provider, dataset, version, observation time, acquisition time, source resolution, transformation and quality state.

Synthetic data cannot masquerade as real evidence

Fixtures are allowed for deterministic tests and explicit demos. They are structurally labelled as synthetic and cannot enter the real-data runtime as observations.

Physical quantities come before generic scores

GeoLens prefers rainfall in millimetres, elevation in metres, slope in degrees, runoff depth, runoff volume and downstream accumulation. A normalised score is used only when its meaning is explicit.

Uncertainty can stop the chain

Unknown pipe direction, missing attachment evidence or incomplete provider coverage can block propagation. A blocked result is more useful than a plausible-looking result built on invented assumptions.

What GeoLens is

GeoLens is:

  • a spatial evidence engine;
  • a common evidence and missing-data contract;
  • a set of real environmental-data providers;
  • an inspectable deterministic runoff derivation;
  • typed catchment and stormwater-network models;
  • a bounded API and visual evidence inspector;
  • a foundation for later environmental and infrastructure applications.

What GeoLens is not

GeoLens does not currently claim to provide:

  • flood probability or flood depth;
  • pipe capacity, surcharge or sewer overflow probability;
  • a calibrated hydraulic simulation;
  • groundwater recharge;
  • damage or financial-loss estimates;
  • a generic 0–1 risk score;
  • continent-scale real-time operation;
  • AI-generated assessment, recommendations or confidence.

AI, Gemini, RAG, mineral prospectivity and the previous generic multi-hazard framing are outside the active runtime.


Technical guide

The remainder of this document is for developers, data providers and reviewers who want to reproduce or inspect the implementation.

Evidence contract

Every important evidence value retains:

  • provider and dataset;
  • dataset version when available;
  • observation time or requested window;
  • acquisition time;
  • coordinate or H3 representation;
  • original spatial resolution;
  • sampling and transformation method;
  • transformation version;
  • quality status and missing reason;
  • provider-specific source metadata.

Canonical evidence states are:

available
missing
stale
out_of_coverage
auth_required
rate_limited
upstream_error
invalid_response
incomplete_window
synthetic_fixture

Observed zero is valid evidence. Missing is not zero.

Active architecture

apps/
  api/                 Fastify spatial-evidence API
  web/                 Next.js institutional site and evidence inspector

packages/
  evidence/            canonical Evidence<T> model and invariants
  providers/           IMERG, GLO-30, CLC, AHN and BGT providers
  stormwater/          runoff, catchments, topology and Waternet/GWSW models
  proof-zero/          end-to-end environmental composition

nasa-precip-engine/    canonical Python earthaccess + xarray IMERG service
copernicus-engine/     local acquisition tooling outside the active npm runtime

Active npm workspaces:

apps/api
apps/web
packages/evidence
packages/providers
packages/stormwater
packages/proof-zero

The Python IMERG service is the only production precipitation path. GeoLens does not maintain a second TypeScript precipitation implementation with a synthetic zero fallback.

Quick start

Requirements

  • Node.js 20 or newer;
  • npm 10 or newer;
  • Python 3.11 or newer;
  • NASA Earthdata credentials for live IMERG;
  • an official local CLC 2018 V2020_20u1 100 m GeoTIFF for live land-cover evidence.

Copernicus GLO-30, PDOK AHN4, PDOK BGT, Waternet public infrastructure and public PDOK/GWSW context do not require credentials.

Install

npm install

cd nasa-precip-engine
python -m pip install -r requirements.txt
cd ..

Create local configuration files:

Copy-Item nasa-precip-engine/.env.example nasa-precip-engine/.env
Copy-Item apps/api/.env.example apps/api/.env

Never commit environment files, credentials, service-key JSON or PEM files.

Configure NASA IMERG

Set nasa-precip-engine/.env:

EARTHDATA_USERNAME=...
EARTHDATA_PASSWORD=...
API_HOST=127.0.0.1
API_PORT=8001
LOG_LEVEL=INFO

IMERG_CACHE_DIR=D:/GeoLens/cache/imerg
IMERG_DISK_CACHE_TTL_SECONDS=2592000
GEOLENS_PYTHON=D:/GeoLens/venvs/nasa-precip/Scripts/python.exe
GEOLENS_TEMP_DIR=D:/GeoLens/tmp

IMERG source resolution is approximately 0.1 degree. H3 does not change that precision. Acquisition is bounded to the requested H3 scope plus one source-cell sampling margin. Cache identities include the area of interest so two places cannot share an accumulation.

Only complete, available windows may be persisted in the optional disk cache. An unavailable or incomplete cached result never becomes an observation or a zero.

The default service binds to the local machine. Use API_HOST=0.0.0.0 only when network access is intentional and protected.

Configure CORINE Land Cover

Keep the official European raster outside the repository. A suitable Windows layout is:

D:/GeoLens/data/clc/
  u2018_clc2018_v2020_20u1_raster100m/
    DATA/
      U2018_CLC2018_V2020_20u1.tif

Set apps/api/.env:

PORT=3003
NASA_PRECIP_SERVICE_URL=http://127.0.0.1:8001
CLC_RASTER_PATH=D:/GeoLens/data/clc/u2018_clc2018_v2020_20u1_raster100m/DATA/U2018_CLC2018_V2020_20u1.tif

If the raster is absent, unreadable or outside coverage, CLC remains explicitly unavailable.

Run

Start IMERG, the API and the web application together:

npm run dev

Default local endpoints:

The root launcher starts services in dependency order and waits for their health gates. The first uncached IMERG acquisition can take several minutes.

APIs

Observed Amsterdam infrastructure

GET /api/infrastructure/amsterdam-waternet acquires a small bounded response from the official Amsterdam Waternet infrastructure API.

Default bounding box:

latitude  52.3375 – 52.3395
longitude 4.8978 – 4.8995

The response keeps these layers separate:

  • observed Waternet topology and source receipts;
  • pipe-invert direction and outfall connectivity;
  • PDOK/GWSW management-area context;
  • raw AHN4 terrain evidence;
  • BGT physical-surface classification;
  • the experimental conditioned BGT/AHN surface proxy;
  • real IMERG, CLC and GLO-30 evidence;
  • derived runoff and catchment contribution;
  • the authoritative BGT Inlooptabel attachment boundary;
  • the reason network propagation was or was not attempted.

The GWSW polygon containing the selected outfall is context only. Point containment does not prove that a surface drains to an outfall.

The companion GET /api/infrastructure/amsterdam-waternet/attachment-intake exposes the delivery gate for the owner-published BGT Inlooptabel requested from Amsterdam/Waternet. Its current state is missing; both attachment assessment and propagation remain blocked. A future package must retain the STOWA 2025 relation semantics, publisher authority, bounded selection, source records and content-addressed original artifacts. Receipt, integrity review and exact topology matching are separate operations. A reviewed package becomes only ready_for_exact_observed_topology_match: propagation remains blocked until its published asset code matches one observed Waternet pipe uniquely. Proximity, polygon containment and the conditioned BGT/AHN outlet cannot be promoted into observed attachment evidence.

Cumbria model-delivery intake

GET /api/benchmarks/cumbria-2015/model-evidence-intake exposes the current Environment Agency Products 5/6/7 delivery state. It reports received, ten explicit incomplete or metadata-only component records, blocked hydraulic-context assessment and blocked replay eligibility. The endpoint exposes neither the external archive path, its checksum nor evaluation geometry. Component review and a later physical-gate decision remain separate operations; receipt cannot promote the delivery automatically.

GET /api/benchmarks/cumbria-2015 exposes the publication-safe retained result: event and grid identity, forcing counts, frozen prediction, three separate comparison records, failure diagnosis and the post-evaluation physics gate. It contains no source arrays, local paths or evaluation geometry and cannot execute or reopen the evaluation.

Emilia-Romagna historical benchmark

GET /api/benchmarks/emilia-romagna-2023 returns a compact, versioned projection of the verified external checkpoint. It does not load or redistribute the 746 MB source archive.

The response exposes:

  • the manifest version, artifact count and integrity method;
  • event window, bounded area, metric grid and H3 representation choice;
  • provider, dataset version, native resolution, role and state for each major source;
  • complete IMERG granule coverage and deterministic runoff quantities;
  • mass balance, model version and the incomplete DBTR window;
  • the post-freeze blind evaluation and retained near-random negative result;
  • the independent ARPAE station comparison;
  • every conditioned-replay evidence gate, including missing and metadata-only states;
  • permitted and forbidden scientific claims.

The institutional Case 02 page reads this endpoint through its evidence inspector. API unavailability is shown as an error; it cannot silently become a valid-looking benchmark result.

The companion GET /api/benchmarks/emilia-romagna-2023/hydraulic-evidence-intake exposes the external-delivery gate for the evidence requested from ARPAE. The current state is missing and replay eligibility is blocked. A future delivery must identify content-addressed artifacts and explicitly cover antecedent state, Montone and Rabbi inflows, downstream boundary, breach behaviour, embankment crests, bare-earth terrain, and channel geometry/roughness. Receipt, structural verification, scientific review and replay eligibility are separate states: receiving a file cannot promote it automatically. Missing components remain missing, chart digitisation and observed-extent leakage are rejected, and synthetic fixtures can test the contract but can never become real replay evidence. Original delivered files remain outside Git.

The companion GET /api/benchmarks/emilia-romagna-2023/map-manifest serves a deterministic publication-safe spatial projection. Five inspectable layers are aggregated from the pinned 30 m grid onto a nominal 300 m display grid: terrain-only D8 contributing area, mean GLO-30 elevation, dominant CORINE land-cover group, known DBTR permanent-water presence and event runoff concentration. Every layer retains native resolution, aggregation, evidence state, transformation version and attribution.

The event-runoff projection became renderable only after the official NASA Earthdata and GPM policies confirmed free use with source acknowledgement; it cites GPM_3IMERGHH V07 by DOI and does not redistribute source granules. The endpoint remains fail-closed for the observed V7 flood extent and ARPAE station geometry, which are registered but carry no map data while redistribution is restricted or still under review. The browser therefore cannot turn an unavailable layer into a visual zero, and neither concentration layer can be mistaken for inundation extent.

The checked-in display payload is reproducible from verified external artifacts:

npm run materialize:emilia-map -- --data-root C:\Users\dacan\GeoLens\data\emilia-romagna-2023
npm run verify:emilia-map -- --data-root C:\Users\dacan\GeoLens\data\emilia-romagna-2023

External evidence intake

GeoLens includes one local intake command for future ARPAE, Amsterdam/Waternet and Cumbria model deliveries. It reads originals from a data directory outside Git, computes their byte counts and SHA-256 identities, validates the appropriate scientific contract and writes a new package receipt. It does not copy the originals and refuses to overwrite an existing receipt.

For future Waternet and Environment Agency packages, the scientific reference is pre-external-evidence-baseline-v1 at commit 938b18fb66925e36236ea04a49eefdb2ca9826cb. Review records and any resulting prediction provenance must retain that identity. Intake may normalize representation; it cannot change model semantics, establish an undocumented attachment or grant calibration access to withheld evaluation evidence.

npm run intake:external-evidence -- --kind arpae --draft C:\path\to\arpae-draft.json --data-root C:\path\to\arpae-delivery --output C:\path\to\receipts\arpae-package.json

npm run intake:external-evidence -- --kind amsterdam --draft C:\path\to\amsterdam-draft.json --data-root C:\path\to\amsterdam-delivery --output C:\path\to\receipts\amsterdam-package.json

npm run intake:external-evidence -- --kind cumbria-model --draft C:\path\to\cumbria-model-draft.json --data-root C:\path\to\cumbria-model-delivery --output C:\path\to\receipts\cumbria-model-package.json

The draft is the corresponding package JSON with artifact entries containing id, role, portable relativePath and mediaType. Any draft bytes, sha256 or local source-path fields are discarded and recomputed from data-root. Artifact paths and resolved symlinks must remain inside that root.

A successful command means only structurally_valid. It does not complete scientific review, prove an Amsterdam topology match or make the Emilia-Romagna replay eligible. Those remain separate fail-closed decisions.

Generic Proof 0

POST /api/proof-zero/run accepts a bounded typed GeoJSON network and an explicit reference time. A complete example is available in apps/web/app/lib/fixture.ts.

Supported entities:

  • Point node with type inlet, manhole or outfall;
  • LineString pipe with type pipe;
  • Polygon catchment with type catchment and an explicit outlet_node_id.

Proof 0 guardrails:

Limit Value
Geographic span 0.25° × 0.25°
GeoJSON features 500
Coordinates 10,000
Nodes 100
Pipes 200
Catchments 50
Catchment H3 cells 500
Request body 1 MiB

These are bounded research-system limits, not continent-scale performance claims.

Technical benchmark notes

Amsterdam direction and attachment semantics

Waternet endpoint invert levels are retained without rounding. The orientation model uses an inclusive 0.05 m resolvable-drop boundary and a separately visible 0.000001 m numeric comparison tolerance.

Direction can be known, unknown or ambiguous. The numeric tolerance handles serialisation noise; it is not a claim about survey accuracy.

The authoritative attachment boundary is modelled as STOWA 2025 BGT Inlooptabel or equivalent owner-published evidence. No such bounded Amsterdam relation has yet been located in the public catalogues.

The executable intake contract records five package states: missing, received, under_review, verified and rejected. Even verified means only that an external delivery may enter the existing exact-identifier assessment; it does not mean that a network attachment or propagated flow has been established. Synthetic fixtures can validate the contract but can never become observed infrastructure evidence.

Emilia-Romagna reproducibility boundary

Case 02 is a retrospective reconstruction, not an as-known-at-the-time forecast. IMERG V07 was released after the event and is therefore labelled as retrospective model input.

The official regional flood extent is evaluation-only. It remained unread until protocol commit 110a217 froze the prediction, score and metrics. It cannot enter model input or calibration.

The common evaluation grid is EPSG:32632 at 30 m, 335 × 420 cells, over bounds [737790, 4895070, 747840, 4907670]. A cell-centre mask retains 130,307 eligible cells.

Important data-quality boundaries remain explicit:

  • missing CLC classes use -1 and never class 0;
  • NaN marks values outside the analysis area or unavailable numeric evidence;
  • 119 DBTR features added after 16 May 2023 are excluded and counted;
  • the current DBTR extract cannot reconstruct deleted or overwritten historical geometry;
  • the regional PST terrain audit contains 560,965 missing values out of 5,069,731 pixels;
  • no PST gap is silently filled from GLO-30;
  • published charts are not digitised into unavailable numerical hydrographs;
  • missing discharge, breach and boundary evidence keeps a conditioned replay blocked.

The ARPAE delivery contract makes that boundary executable: a package may be received, under_review, verified or rejected, while replay eligibility remains independently blocked until every required component is accepted as external evidence. Invalid or incomplete deliveries cannot become zero-filled model input.

External benchmark inputs remain outside Git. While the D volume is unavailable, the verified working copy is under C:/Users/dacan/GeoLens/data/emilia-romagna-2023; the manifest uses portable relative paths so it can move back to D without changing dataset identity.

Manifest v1.16.0 pins 55 benchmark artifacts totaling 746,444,721 bytes and records the completed NASA/GPM use-policy review. This includes the canonical IMERG cache and portable source grid, DBTR source and metadata, derived masks, terrain routing, runoff arrays, evaluation receipts and audited official reports. Restricted observed geometry and source archives are not redistributed through Git.

For the complete audit trail, phase state and source-by-source limitations, read REFOUNDATION_PLAN.md and the Emilia-Romagna manifest.

The paths below point to the temporary verified copy on C. Change the four variables when the data move to another volume; the manifest identity and verification rules do not change.

$benchmarkRoot = 'C:\Users\dacan\GeoLens\data\emilia-romagna-2023'
$clcRaster = 'C:\Users\dacan\GeoLens\data\clc\clc2018-forli-feature-service\U2018_CLC2018_V2020_20u1_Forli_feature_service_100m.tif'
$dbtrSource = 'C:\Users\dacan\GeoLens\downloads\dbtr-source\estraz_procons.gpkg'
$imergCacheRoot = 'C:\Users\dacan\GeoLens\cache\imerg'

npm run materialize:emilia-inputs -- $benchmarkRoot $clcRaster
npm run materialize:emilia-xdbtr -- --data-root $benchmarkRoot --source $dbtrSource
npm run materialize:emilia-imerg-cache -- --data-root $benchmarkRoot --metadata "$imergCacheRoot\v07_20230518T000000Z_48h_669b94d37ff0.json" --netcdf "$imergCacheRoot\v07_20230518T000000Z_48h_669b94d37ff0.nc"
npm run verify:emilia-inputs -- $benchmarkRoot
npm run materialize:emilia-terrain-routing -- $benchmarkRoot
npm run verify:emilia-terrain-routing -- $benchmarkRoot
npm run materialize:emilia-event-runoff -- $benchmarkRoot
npm run verify:emilia-event-runoff -- $benchmarkRoot
npm run evaluate:emilia-concentration -- --data-root $benchmarkRoot
npm run verify:ground-truth -- $benchmarkRoot

Verification

Run deterministic verification:

npm run typecheck
npm test
npm run build

Live providers are opt-in and separate from deterministic fixtures:

$env:GEOLENS_RUN_LIVE_PROVIDER_TESTS = '1'
$env:CLC_RASTER_PATH = 'D:/GeoLens/data/clc/u2018_clc2018_v2020_20u1_raster100m/DATA/U2018_CLC2018_V2020_20u1.tif'
$env:GEOLENS_IMERG_REFERENCE_TIME = '2026-08-20T00:00:00Z'
npm run test:live

Waternet, BGT and GWSW live checks:

$env:GEOLENS_LIVE_WATERNET = '1'
$env:GEO_LENS_LIVE_BGT = '1'
npm run build --workspace=@geo-lens/providers
npm run build --workspace=@geo-lens/stormwater
node --test packages/providers/test/live-bgt.test.cjs
node --test packages/stormwater/test/live-amsterdam-wfs.test.cjs packages/stormwater/test/live-amsterdam-gwsw.test.cjs

A live-provider failure may reflect credentials, network, rate limits or incomplete upstream coverage. It must never produce a valid-looking zero.

Scientific limits

  • Runoff v0 is deterministic and inspectable, but not a calibrated flood model.
  • Land-cover-derived imperviousness and runoff parameters are model inputs or proxies.
  • GLO-30 slope and AHN/BGT terrain conditioning have different roles and remain separate.
  • H3 resolution never replaces provider resolution.
  • Waternet direction uses endpoint invert evidence; insufficient evidence remains unknown or ambiguous.
  • The conditioned Amsterdam outlet is a model boundary, not an observed sewer attachment.
  • Propagation does not model pipe capacity, storage, travel time, surcharge or overflow.
  • No percentage confidence or production-readiness claim is generated without a separate validation procedure.

Refoundation and project history

GeoLens originally contained a broad multi-hazard product, AI analysis, mineral exploration and contradictory data paths. The refoundation deliberately reduced the active system to one physically meaningful, provenance-complete chain.

The pre-overhaul repository is preserved on branch codex/pre-overhaul-snapshot-20260822. Do not rewrite that branch.

The public release line remains v0.1.0-alpha.4. The separate annotated tag pre-external-evidence-baseline-v1 freezes commit 938b18fb66925e36236ea04a49eefdb2ca9826cb for the Waternet and Environment Agency external-evidence tests; it is a scientific baseline, not a release claim. Current main may contain later maintenance and protocol-preserving work.

When repository materials disagree, use this order:

  1. AGENTS.md
  2. REFOUNDATION_PLAN.md
  3. verified runtime behaviour
  4. tests expressing intended behaviour
  5. implementation
  6. historical documentation

The durable execution state belongs in the plan, code, tests and commits, not in generated completion reports.

About

Experimental spatial evidence engine: real rainfall, terrain, land cover and stormwater infrastructure into traceable runoff and downstream physical state.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages