Skip to content

New YAML schema - #119

Open
AgentOxygen wants to merge 9 commits into
mainfrom
yaml-schema
Open

AgentOxygen wants to merge 9 commits into
mainfrom
yaml-schema

Conversation

@AgentOxygen

Copy link
Copy Markdown
Collaborator

This is a large change that needs careful revision and feedback on before merging.

A centrally defined schema for all of our YAML files. This makes many of the assumptions made throughout the source code explicit and allows us to enforce them when we inevitably change the YAMLs to accommodate updates to our upstream spreadsheets.

The new schema is defined in one file using JSON-Schema, which is added as a dependency:
data/schemas/mapping.schema.yaml

A lot of the entries in the YAML files were unused or otherwise redundant and could be replaced with reasonable defaults. This greatly reduced the size and complexity of the YAML files. For many variables, a one-line rename mapping is all that is needed, so entries take on either a "short form" or "long form" style.

Short form example (simple rename):

variables:
  tas_tavg-h2m-hxy-u: TREFHT

Long form example:

variables:
  clt_tavg-u-hxy-u:
    formula:
      mon: CLDTOT
      day: CLDTOT_d
    units: "1" 
    positive: down
    grids: [gn, gr]

Changes from the old format:

  • formula can be delineated between daily and monthly in the same block
  • sources now derives from the formula directly rather than restating
  • table table gets set by the specified realm parameter
  • units sets formula result units if they differ from CMIP7 table
  • levels comes from the CMIP7 table entry dimensions
  • regrid_method assumes conservative, but overrides come from data/intensive_vars.yaml
  • cell_methods derived from CMIP7 table
  • long_name derived from CMIP7 table
  • standard_name derived from CMIP7 table

This is a minimal implementation in the interest of avoiding large PRs. As a result, I suppressed the linter rule for too many lines. This should be removed in a later clean up.

@AgentOxygen
AgentOxygen marked this pull request as draft September 28, 2026 19:38

@maritsandstad maritsandstad left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, I think this looks like a really good way to go, and really amazing refactoring.

However, I am getting a slight bit of whiplash from it, and think we might need a bit more time and validation to check that this rewrite isn't loosing us something important that we have tested before, but we don't have test coverage for (which I think might be a significant amount). Maybe you two @AgentOxygen and @mvertens are completely on top of this, and it might be fine, at least for my part I don't feel like I have a grasp of it that is through enough that I want to merge it in directly, especially given that I think we will start doing production runs fairly soon (hopefully).

Comment on lines +42 to +49
grids:
description: >-
Default output grids for every entry in this file: gn keeps the native
grid, gr regrids to lat/lon. Default [gr].
type: array
minItems: 1
uniqueItems: true
items: {enum: [gn, gr]}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Flagging; we will want to change this to the updated CMOR-table defined grids for CESM and NorESM as defined in the EMD

@AgentOxygen

Copy link
Copy Markdown
Collaborator Author

However, I am getting a slight bit of whiplash from it, and think we might need a bit more time and validation to check that this rewrite isn't loosing us something important that we have tested before, but we don't have test coverage for (which I think might be a significant amount). Maybe you two @AgentOxygen and @mvertens are completely on top of this, and it might be fine, at least for my part I don't feel like I have a grasp of it that is through enough that I want to merge it in directly, especially given that I think we will start doing production runs fairly soon (hopefully).

That's totally fair. Especially since we are expecting to start production runs soon, it would be safer to leave this in draft form until we have implemented more test coverage separately. The main reason I ended up with a large refactor was because writing a schema for the existing structure was terribly verbose and added more complexity than it was worth.

For now I can focus on just extending the existing test suite as I develop our new functions. We can hopefully revisit this after testing coverage is improved.

@AgentOxygen AgentOxygen added the enhancement New feature or request label Sep 29, 2026
@mvertens

Copy link
Copy Markdown
Contributor

@AgentOxygen - this is great! I had claude review the yaml files - here is a summary of the issues that were found.

@marit @AgentOxygen - I think this refactor has actually exposed some real problems that need to be addressed and I am wondering if we should migrate it out of draft form once we resolve these issues. A meeting this Friday would be helpful.

Mapping entries that don't match the CMOR tables

54 entries across the data/*_to_cmip7_*.yaml files use a variable name that
does not appear in the CMOR table for the realm they are listed under.

Checked against the cmip7-cmor-tables submodule as it is pinned in this
checkout (cf46e3c, branch noresm-dev).

Background: how a name is built

A CMIP7 variable name has a short root name, then a suffix of codes:

siconc_tavg-u-hxy-u
│      │    │ │   │
│      │    │ │   └─ area:       what area the value is averaged over
│      │    │ └───── horizontal: gridded, or a single global number
│      │    └─────── vertical:   which level, or none
│      └──────────── time:       averaged, instantaneous, minimum, maximum
└─────────────────── root:       what the quantity is (sea-ice concentration)

The root says what the quantity is. The suffix says how it was sampled. The
same letter means different things in different slots -- the u in the
vertical slot and the u in the area slot both mean "not applicable", and
neither has anything to do with the wind variables ua / uas, which are
root names.

The area code is the one that matters below. si means the average was taken
over only the ice-covered part of the cell. u means it was taken over the
whole cell. For sea-ice thickness you want si, because you average thickness
over the ice. For sea-ice concentration you want u, because concentration is
the fraction of the whole cell that is ice.

Why a mismatch matters

The mapping files now hold one line per variable and look up everything else
in the CMOR table. _add_table_fields in
src/cmip7_prep/mapping_compat.py:688 searches
cmip7-cmor-tables/tables/CMIP7_<realm>.json for the exact name and copies
over the units and the level information.

The search is on the full name, as one string. If the name isn't found, nothing
is raised -- the code fills in the realm it was given and returns. The variable
ends up with no units and no level information, and nothing says so.

Regridding is not affected by any of this. The conservative-or-bilinear choice
is made earlier, in expand_entry (src/cmip7_prep/mapping_compat.py:437),
from data/intensive_vars.yaml, using only the root name. No table is
consulted.

The four kinds of mismatch

count what it is what to do
A 2 the name doesn't exist in CMIP7 in any form decide what was meant
B 19 the name is correct but sits in a different realm's table code change
C 19 the root name is nowhere in CMIP7 at all check the data request
D 14 two entries for the same quantity, only one of which matches find out why

Only B is fixed in code. A, C and D need someone to decide what was intended.

A. No matching name (2)

Both are CESM relative humidity. The atmos table has hurs only at 2 m
(h2m in the vertical slot). Neither of these two entries says h2m, and
there is no table entry that would accept them.

model realm YAML entry closest names in the table
cesm atmos hurs_tavg-al-hxy-u hurs_tavg-h2m-hxy-u, hurs_tpt-h2m-hxy-u, and 4 more, all h2m
cesm atmos hurs_tavg-u-hxy-u same

Someone needs to say which height these were meant to be at.

B. Right name, wrong table (19)

The name is spelled correctly and exists in CMIP7 -- just in a different
realm's file than the one being searched. The code only ever looks in the
realm it was given, so it never finds them.

The groupings are sensible ones. Irrigation fields belong in the CESM atmos
mapping from a modelling point of view, even though CMIP7 files them under
land. The problem is the lookup, not the mapping.

model realm searched YAML entry actually filed under
cesm land evspsbl_tavg-u-hxy-lnd atmos
cesm atmos evspsblsoi_tavg-u-hxy-u land
cesm atmos evspsblveg_tavg-u-hxy-u land
cesm atmos irrDem_tavg-u-hxy-u land
cesm atmos irrGw_tavg-u-hxy-u land
cesm atmos irrLut_tavg-u-hxy-u land
cesm atmos irrSurf_tavg-u-hxy-u land
cesm ocean siareaacrossline_tavg-u-ht-u seaIce
cesm ocean sihc_tavg-u-hxy-sea seaIce
cesm seaIce sialgc_tavg-u-hxy-si ocnBgchem
noresm seaIce sialgc_tavg-u-hxy-si ocnBgchem
cesm seaIce sichl_tavg-u-hxy-si ocnBgchem
noresm seaIce sichl_tavg-u-hxy-si ocnBgchem
cesm seaIce sigpp_tavg-u-hxy-si ocnBgchem
noresm seaIce sigpp_tavg-u-hxy-si ocnBgchem
cesm seaIce sino3_tavg-u-hxy-si ocnBgchem
noresm seaIce sino3_tavg-u-hxy-si ocnBgchem
cesm seaIce sisi_tavg-u-hxy-si ocnBgchem
noresm seaIce sisi_tavg-u-hxy-si ocnBgchem

The fix is for _add_table_fields to search the other CMIP7_*.json files
when the variable isn't in the one for the realm it was given, and to use
whichever file does have it.

That file's name is also the answer to which realm the variable belongs to, so
table should be set from it rather than left as the realm the run was started
with. table is what CMOR loads to write the file, so getting it wrong matters
as much as the missing units.

One rule still has to be decided: what to do if a name turns up in more than
one CMIP7_*.json file. It doesn't come up for these 19 -- each is in exactly
one -- but the search shouldn't just take whichever file it happens to read
first.

C. Not in CMIP7 at all (19)

For these, every table was searched for the root name on its own, ignoring the
suffix. Nothing came back. So this is not a suffix problem -- the quantity
itself is absent from this build of the tables.

Four of them (sitemptop, sisnthick, sipr, sisndmasssnf) were standard
CMIP6 sea-ice variables. The likely explanation is that they were dropped or
renamed between CMIP6 and CMIP7. That is worth confirming against the data
request before anything is removed.

model realm YAML entry
cesm seaIce siflswdtop_tavg-u-hxy-si
noresm seaIce siflswdtop_tavg-u-hxy-si
cesm seaIce siflswutop_tavg-u-hxy-si
noresm seaIce siflswutop_tavg-u-hxy-si
cesm ocean sig1_tavg-u-hxy-sea
cesm seaIce sipr_tavg-u-hxy-si
noresm seaIce sipr_tavg-u-hxy-si
cesm seaIce sirdgthick_tavg-u-hxy-si
noresm seaIce sirdgthick_tavg-u-hxy-si
cesm seaIce sisndmassmelt_tavg-u-hxy-si
noresm seaIce sisndmassmelt_tavg-u-hxy-si
cesm seaIce sisndmasssnf_tavg-u-hxy-si
noresm seaIce sisndmasssnf_tavg-u-hxy-si
cesm seaIce sisndmasssubl_tavg-u-hxy-si
noresm seaIce sisndmasssubl_tavg-u-hxy-si
cesm seaIce sisnthick_tavg-u-hxy-si
noresm seaIce sisnthick_tavg-u-hxy-si
cesm seaIce sitemptop_tavg-u-hxy-si
noresm seaIce sitemptop_tavg-u-hxy-si

D. Two entries, one of which matches (14)

These are seven quantities, each appearing twice in both sea-ice files. One of
the two names is in the CMOR table; the other is not. The two entries read
from the same model field.

siconc in data/noresm_to_cmip7_seaIce.yaml is the pattern:

siconc_tavg-u-hxy-si:      # not in the table
  formula: {day: siconc_d, mon: siconc}
  units: '%'               # units written out by hand

siconc_tavg-u-hxy-u:       # in the table
  formula: {day: siconc_d, mon: siconc}
                           # no units; taken from the table

The entry that doesn't match the table has its units written out in the file.
That looks deliberate rather than accidental -- whoever added it knew the
lookup would come back empty and supplied the units directly.

So this is not a set of typos, and renaming the unmatched entry would collide
with the matched one that is already there. What is needed is an explanation of
why both are wanted. Both files gained these in 85ac655, "Updated data files
to new schema".

quantity in the table also present, not in the table
sea-ice concentration siconc_tavg-u-hxy-u siconc_tavg-u-hxy-si
ice concentration by thickness sidconcth_tavg-u-hxy-sea sidconcth_tavg-u-hxy-si
ice mass transport, x sidmasstranx_tavg-u-hxy-u sidmasstranx_tavg-u-hxy-si
ice mass transport, y sidmasstrany_tavg-u-hxy-u sidmasstrany_tavg-u-hxy-si
melt-pond refrozen ice simprefrozen_tavg-u-hxy-simp simprefrozen_tavg-u-hxy-si
ice velocity siv_tavg-u-hxy-si siv_tavg-u-hxy-ifs

Each row appears in both cesm_to_cmip7_seaIce.yaml and
noresm_to_cmip7_seaIce.yaml, which is why the count is 12 for these six and
not 6.

The remaining two of the 14 are sivol_tavg-u-hxy-si, one in each file. They
are listed here only because the check groups entries by root name, and both
files also carry sivol_tavg-u-hm-u. Those two are not a pair:
sivol_tavg-u-hm-u is a single global total in 1e3 km^3, built by summing
over cells, while sivol_tavg-u-hxy-si is the gridded field in m3 m-3. Both
are legitimate and neither needs fixing.

How this was checked

For each mapping file, take the variable names and look each one up in
cmip7-cmor-tables/tables/CMIP7_<realm>.json for that file's realm. For the
ones not found, sort them by asking, in order:

  1. Does the same file also contain a name with the same root that the table
    does have? If so it is category D.
  2. Does the exact name appear in some other realm's table? Category B.
  3. Does the root name appear anywhere in any table? If yes, category A. If no,
    category C.

Step 1 is the one that matters. Without it, the 14 entries in category D look
like misspellings of the entries sitting next to them.

@maritsandstad

Copy link
Copy Markdown
Collaborator

So this is all super useful, and I don't think it necessarily needs to stay in draft, just we want to do some validation and testing, also possibly discuss how we think about the direction and implication for our csv to yaml setup (I think a lot might be superfluous in the current code that we added to conform to the old yaml scheme).

One question I have for instance is what is expected for the unit conversion. To me it looks like the intention to work towards doing it automatically from comparing inputs to requests? If so, I'm just concerned whether that this doesn't cause some sort of double conversion (for NorESM we've generally gotten scientists to make those types of conversions explicitly by writing in formulas).

Comment on lines -1 to -4
dataset_overrides:
institution_id: NCC
nominal_resolution: 200 km
source_id: NorESM3

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So, one more flag, in this rewrite, where does this information live?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, actually really good to get this out from here, it is not used anyways. Looking currently at how to change so we can use it with updated tables version.

@mvertens

Copy link
Copy Markdown
Contributor

@AgentOxygen - I have merged this to main and resolved conflicts - and I also needed to add a change to cmor_utils.py to added import of lru_cache. Can I push back to your branch at this point?

@maritsandstad - I ran run_reference_case in my sandbox (~/test_cmor is /scratch/mvertens/noresm3/test_cmor/):

 python run_reference_case.py --model noresm --model-res NorESM3-LM --experiment piControl --case-dir /nird/datalake/NS9560K/noresm3/cases/n1850GaxgGHG.LM.nor30b25.528.20260928 --outdir ~/test_cmor/n1850GaxgGHG.LM.nor30b25.528.20260928 --years 1526:1527:2 --plots --html

and the validation summary is here:
https://ns9560k.web.sigma2.no/datalake/diagnostics/noresm/mvertens/n1850GaxgGHG.LM.nor30b25.528.20260928/validation_reports/index.html

If the validation looks reasonable I propose that I push back to the branch origin/yaml-schema and merge it. @AgentOxygen @maritsandstad - does that sound reasonable?

@maritsandstad

Copy link
Copy Markdown
Collaborator

@mvertens , yes this sound reasonable to me. However, do you have a link to a summary without this? I am mainly concerned with whether this produces changes to what we are able to produce.

@mvertens

Copy link
Copy Markdown
Contributor

@maritsandstad - the original validation reports are in /projects/NS9560K/www/diagnostics/noresm/n1850GaxgGHG.LM.nor30b25.528.20260928.main
and the new ones are in
/projects/NS9560K/www/diagnostics/noresm/n1850GaxgGHG.LM.nor30b25.528.20260928

@mvertens

mvertens commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

@mvertens

Copy link
Copy Markdown
Contributor

I did not have plots for the original - so I think I need to regenerate them to have a valid comparison.

@maritsandstad

Copy link
Copy Markdown
Collaborator

No worries, the original definitely does not do any better. I'm fine with the validation here, we need to get people to check things anyway, this doesn't look to be breaking anything to me.

@mvertens

Copy link
Copy Markdown
Contributor

@maritsandstad @AgentOxygen - in looking at the validation reports and comparing the plots it looks like this PR is showing up more plots than what we get from main:
The output from main:
https://ns9560k.web.sigma2.no/datalake/diagnostics/noresm/mvertens/n1850GaxgGHG.LM.nor30b25.528.20260928.main/validation_reports/index.html
The output from this PR:
https://ns9560k.web.sigma2.no/datalake/diagnostics/noresm/mvertens/n1850GaxgGHG.LM.nor30b25.528.20260928/validation_reports/index.html

@maritsandstad

Copy link
Copy Markdown
Collaborator

I am happy for conflicts to be resolved and this to be merged in, and then I think we need to have a coordinated validation round with the scientists on our end again soon anyways @mvertens

@mvertens
mvertens marked this pull request as ready for review October 11, 2026 16:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants