Skip to content

Commit 292e1ad

Browse files
chross22claude
andcommitted
Make the README orientation and move the depth into vignettes
It had reached 11,600 words - a forty-five minute read - because every capability added a section rather than a line. The structure had also drifted: it opened by promising five sources and then spent eight hundred lines on Copernicus before the others appeared. Nothing is deleted. Two new articles take what left: Working with what comes back - resampling between grids and time steps, filling satellite gaps, the four plots, and what n_workers actually buys. None of it is source-specific, which is the point: a grid from any of the five resamples and plots the same way, and that was hard to see while the material sat inside the Copernicus walkthrough. Static and basin-scale covariates - seafloor terrain and the climate indices, including what TPI measures, why the indices are not interchangeable, and the LCR and AMOC caveats. The deeper halves of "Satellite or model?", "Bottom salinity" and "Downloads run in parallel" moved into the existing articles, each leaving a summary and a link. What stays in the README is what the ask was: the basic functions, and the need-to-know facts about each data source - what it gives, over what record, at what resolution, and what it will not do. A function reference table replaces the old lookup section. Every export, grouped by purpose, one line each. That is what moving the depth out required - the README now says what each function is and where to read more, rather than explaining each at length - and it keeps the check that every export is named in user-facing documentation honest. Sections are reordered so the five sources sit together and the shared tooling follows them, the contents list is rebuilt to match, and every internal link was checked: two pointed at sections that had moved. One thing worth recording because it only shows up at build time: knitr rejects duplicate chunk labels, and moving a section between documents can collide with a label already there. The bottom-salinity chunk did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent b971905 commit 292e1ad

7 files changed

Lines changed: 933 additions & 1163 deletions

File tree

‎NEWS.md‎

Lines changed: 26 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -187,6 +187,32 @@
187187
* `variable_dictionary("wind")` filters the catalog to the wind variables, and
188188
`variable_dataset(frequency = "hourly")` reports which have an hourly dataset.
189189

190+
## Documentation
191+
192+
* **The README is orientation now, and the depth is in vignettes.** It had grown
193+
to 11,600 words — a forty-five minute read — because every capability added a
194+
section to it. It is now 9,000, and what left went into two new articles
195+
rather than being deleted:
196+
197+
- **Working with what comes back** — resampling between grids and time steps,
198+
filling satellite gaps, the four plots, and what `n_workers` actually buys.
199+
- **Static and basin-scale covariates** — seafloor terrain and the climate
200+
indices, including what TPI measures and why the indices are not
201+
interchangeable.
202+
203+
The deeper parts of "Satellite or model?", "Bottom salinity" and "Downloads
204+
run in parallel" moved into the existing articles, each leaving a short
205+
summary and a link behind. Four articles now: getting started, choosing a
206+
source, working with the result, and the other covariates.
207+
208+
* **A function reference table.** Every export, grouped by what it is for, in
209+
one place in the README. That is what the depth moving out required: the
210+
README says what each function is and where to read more, rather than
211+
explaining every one at length.
212+
213+
* The contents list was rebuilt to match, and every internal link checked — two
214+
pointed at sections that had moved to a vignette.
215+
190216
## Bug fixes
191217

192218
* **`upscale_time()` and `downscale_time()` did not record the step they

‎README.Rmd‎

Lines changed: 154 additions & 555 deletions
Large diffs are not rendered by default.

‎README.md‎

Lines changed: 172 additions & 608 deletions
Large diffs are not rendered by default.

‎_pkgdown.yml‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -93,3 +93,12 @@ reference:
9393
- erddap_dictionary
9494
- product_url
9595
- as_markdown
96+
97+
articles:
98+
- title: Articles
99+
navbar: ~
100+
contents:
101+
- datamatch
102+
- sources
103+
- working-with-data
104+
- covariates

‎vignettes/covariates.Rmd‎

Lines changed: 230 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,230 @@
1+
---
2+
title: "Static and basin-scale covariates"
3+
output: rmarkdown::html_vignette
4+
vignette: >
5+
%\VignetteIndexEntry{Static and basin-scale covariates}
6+
%\VignetteEngine{knitr::rmarkdown}
7+
%\VignetteEncoding{UTF-8}
8+
---
9+
10+
```{r setup, include = FALSE}
11+
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
12+
library(datamatch)
13+
```
14+
15+
Not every useful covariate is a gridded field on a monthly step. Two other kinds
16+
are available, and they behave differently from the gridded data and from each
17+
other: **seafloor terrain**, which does not vary in time at all, and
18+
**basin-scale climate indices**, which do not vary in space.
19+
20+
Neither is fetched with `matchData()`. They have their own attach functions,
21+
because neither is a spatiotemporal nearest-feature join.
22+
23+
24+
Not everything useful is a gridded Copernicus variable on a monthly time step.
25+
Two other kinds are available, and they behave differently from the Copernicus
26+
data and from each other.
27+
28+
### Seafloor terrain
29+
30+
This does not vary by month or year. So it is fetched once for a study area and
31+
attached to every time step. It comes from NOAA ETOPO via `marmap`, a suggested
32+
package rather than a hard dependency:
33+
34+
```{r bathy, eval = FALSE}
35+
bathy <- fetch_bathymetry(
36+
bounding_box = list(xmin = -70, xmax = -66, ymin = 41, ymax = 44)
37+
)
38+
39+
observations <- attach_bathymetry(observations, bathy, c("DEPTH", "SLOPE", "TPI"))
40+
```
41+
42+
`bathymetry_variables()` lists them: `DEPTH`, `SLOPE`, `ASPECT`, and `TPI`.
43+
44+
`TPI`, the topographic position index, is a cell's depth relative to the mean of
45+
the eight around it. It answers a question depth cannot. A 100 m bank top and a
46+
100 m basin floor are the same depth and very different places. TPI is positive
47+
on the first, negative on the second, and near zero on both flat bottom and
48+
uniform slopes.
49+
50+
It is scale-dependent by construction. It describes position within the
51+
immediate neighbourhood, so its meaning follows the
52+
resolution of the grid it was computed on.
53+
54+
`attach_bathymetry()` takes either a plain data frame with coordinate columns or
55+
an `sf` object, so it works on observations and on `accessEnvDat()` output alike.
56+
57+
### Climate indices
58+
59+
These are monthly, and have **no spatial dimension**. One value describes the
60+
whole basin, so every observation in a month receives the same number:
61+
62+
```{r indices, eval = FALSE}
63+
observations <- attach_climate_index(observations, c("NAO", "AMO"))
64+
65+
index_dictionary() # NAO, AO, AMO, PDO, LCR, AMOC
66+
```
67+
68+
That makes them a different kind of covariate from local temperature. They tell
69+
you about *when* — what year and season it was. They tell you nothing about
70+
*where* within a region conditions are better.
71+
72+
So they are useful for interannual questions ("was this a warm-regime year?") and
73+
useless for spatial ones. A model given only indices cannot produce a map.
74+
75+
Six are available, and they are not interchangeable.
76+
77+
| Name | What it measures | Units | Timescale | Record |
78+
|---|---|---|---|---|
79+
| `NAO` | Pressure difference, Icelandic Low to Azores High | standardized anomaly | Year to year | 1950– |
80+
| `AO` | Strength of the polar vortex | standardized anomaly | Year to year | 1950– |
81+
| `AMO` | Detrended North Atlantic SST anomaly | degrees C | Multidecadal | 1948– |
82+
| `PDO` | Leading mode of North Pacific SST | standardized anomaly | Decadal | 1948– |
83+
| `LCR` | Retroflection of the Labrador Current | fraction | Year to year | 1993–2014 |
84+
| `AMOC` | Overturning transport at 26.5°N | Sv | Year to year | 2004– |
85+
86+
Most are **standardized anomalies** — in standard deviations, not anything
87+
physical. Only `AMO` (degrees C) and `AMOC` (Sverdrups) carry real units, so a
88+
coefficient fitted to one of those is not comparable with one fitted to `NAO`.
89+
`index_dictionary()` carries the units at runtime.
90+
91+
**`NAO`** is the usual first choice in the Northwest Atlantic. It sets the
92+
strength and track of the westerlies, and with them heat flux and mixing over the
93+
shelves. Its winter values carry most of the signal. Summer values are noisier
94+
and mean less.
95+
96+
**`AO`** is the hemispheric version of the same thing: polar vortex strength
97+
rather than an Atlantic-specific pressure dipole. In winter it correlates
98+
strongly with `NAO`. Including both usually buys little and costs collinearity,
99+
so pick one unless you have a reason.
100+
101+
**`AMO`** is a slow background state, not a year-to-year signal. It varies over
102+
decades. Within a study period of ten or twenty years it may act more like a
103+
trend than a covariate. One caution: it is *derived from* North Atlantic SST, so
104+
using it to predict local SST is partly circular.
105+
106+
**`PDO`** is North Pacific and included for completeness. It has limited relevance
107+
to Atlantic shelf systems, and a relationship found with it is worth treating
108+
sceptically.
109+
110+
The last two are different in kind, and get their own sections below.
111+
112+
#### The overturning circulation
113+
114+
`AMOC` is the strength of the Atlantic overturning in Sverdrups, measured
115+
directly by the RAPID mooring array at 26.5°N. Alone among these it is a
116+
measurement rather than a pattern derived from pressure or SST.
117+
118+
```{r amoc, eval = FALSE}
119+
amoc <- fetch_climate_index("AMOC") # monthly, April 2004 onward
120+
observations <- attach_climate_index(observations, "AMOC")
121+
```
122+
123+
Being real is also its limitation, in two ways. It **starts in April 2004**, so
124+
it cannot reach back over a longer survey record. And it is measured **far south
125+
of the shelf**: it describes the basin-scale circulation the Labrador and slope
126+
currents sit within, not conditions on the Gulf of Maine or Scotian Shelf. A
127+
weakening AMOC is associated with a warming Northwest Atlantic shelf, but that
128+
is a chain of several links, so check the sign of any relationship against `LCR`
129+
and `AMO` rather than assuming it.
130+
131+
Two practical notes. RAPID publishes this only as NetCDF, so it needs the
132+
`ncdf4` package — a `Suggests`, installed with `install.packages("ncdf4")`. And
133+
the published series is twelve-hourly, averaged to monthly here; the file is over
134+
a megabyte, so it is cached like the Copernicus downloads.
135+
136+
> Moat BI, Smeed DA, Rayner D, Johns WE, Smith R, Volkov D, Elipot S, Petit T,
137+
> Kajtar J, Baringer MO, Collins J (2026). Atlantic meridional overturning
138+
> circulation observed by the RAPID-MOCHA-WBTS array at 26°N from 2004 to 2024
139+
> (v2024.1a). British Oceanographic Data Centre, NERC, UK.
140+
> <https://doi.org/10.5285/48d0bf43-0598-ceb2-e063-7086abc062f1>
141+
142+
#### Staying current
143+
144+
Four of the six indices are still growing. Downloads are cached, and the cache
145+
expires on an interval matched to how often each provider actually publishes, so
146+
a living index re-downloads on its own without being asked:
147+
148+
```{r indices-status, eval = FALSE}
149+
climate_index_status() # what is cached, how old, what is due
150+
refresh_climate_index() # force a re-fetch of everything still growing
151+
refresh_climate_index("AMOC")
152+
```
153+
154+
| Updates | Cache reused for | Indices |
155+
|---|---|---|
156+
| Monthly | 7 days | `NAO`, `AO`, `AMO`, `PDO` |
157+
| Roughly yearly | 30 days | `AMOC` |
158+
| Never | forever | `LCR` |
159+
160+
`LCR` finished at 2014, so re-downloading it cannot produce anything new and it
161+
is skipped rather than fetched again.
162+
163+
Two behaviours are worth knowing, because both are deliberate.
164+
165+
**Staleness is judged from the data, not the cache.** A fresh download of a file
166+
the provider stopped updating is still stale. If a series ends further behind the
167+
present than its source's usual publishing lag, you are told, along with the
168+
command to force a refresh and the note that if refreshing changes nothing then
169+
the provider has not published either.
170+
171+
**A failed download falls back to the cached copy, with a warning.** A provider
172+
being briefly unreachable should not become an error here when usable data is
173+
already on disk. The warning is what keeps the old copy from being mistaken for
174+
a current one.
175+
176+
#### The Labrador Current retroflection index
177+
178+
`LCR` is the odd one out among the five, and often the most directly useful. The
179+
other four are atmospheric or SST patterns. This one describes a *current*: how
180+
much of the Labrador Current turns eastward at the Grand Banks instead of
181+
continuing southwest along the shelf.
182+
183+
Positive values mean stronger retroflection. That means less cold, fresh,
184+
oxygen-rich Labrador water reaching the Scotian Shelf and Gulf of Maine. For
185+
shelf water properties that is a shorter causal chain than the NAO, which acts on
186+
them only indirectly through winds.
187+
188+
```{r lcr, eval = FALSE}
189+
lcr <- fetch_climate_index("LCR") # monthly, 1993-2014
190+
observations <- attach_climate_index(observations, "LCR")
191+
```
192+
193+
It is the published output of one study rather than an operational product, so
194+
**cite it when you use it**:
195+
196+
> Jutras M, Dufour CO, Mucci A, Talbot LC (2023) Large-scale control of the
197+
> retroflection of the Labrador Current. *Nature Communications* **14**:2623.
198+
> <https://doi.org/10.1038/s41467-023-38321-y>
199+
200+
`index_dictionary()` prints that citation, and
201+
`as.data.frame(index_dictionary())$reference` carries it at runtime.
202+
203+
Three limits worth knowing.
204+
205+
It **covers 1993–2014 only**, so it cannot be attached to recent observations.
206+
207+
The series is fetched from the paper's published source data, not recomputed. The
208+
authors derived it by seeding 966 virtual particles per week across a line at
209+
(53°N, 56.7°W)–(54.3°N, 52.0°W). Each was tracked for three years through
210+
GLORYS12V1 velocities with [OceanParcels](https://parcels-code.org/). The index
211+
is the difference between the counts crossing hydrographic sections on the
212+
Labrador and Scotian Shelves.
213+
214+
Extending the record past 2014 means redoing that computation, not calling a
215+
different function.
216+
217+
**The values are the raw index, not the normalized one plotted in the paper.**
218+
The published source data runs roughly −0.09 to 0.18, consistent with a fraction
219+
of the seeded particles. Figure 3a of the paper shows a detrended, smoothed
220+
series normalized to [−1, 1], spanning about −0.6 to +0.5.
221+
222+
Both describe the same quantity, but the numbers are not comparable. Don't read a
223+
value here against that figure. To get the paper's variant, apply its chain
224+
yourself: detrend, 12-month rolling mean, rescale to [−1, 1], subtract the
225+
1993–2015 mean.
226+
227+
One interaction to know. `attach_climate_index()` joins on year and month, and
228+
`upscale_time(to = "year")` stamps its output `MONTH = 1`. So attaching an index
229+
to annual data would give every year January's value. Aggregate the index to a
230+
year first instead.

‎vignettes/sources.Rmd‎

Lines changed: 90 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -206,3 +206,93 @@ index_dictionary() # carries the climate index references
206206

207207
The README's References section lists the Copernicus product DOIs. Cite whichever
208208
you actually fetched from.
209+
210+
## Satellite or model?
211+
212+
Copernicus serves chlorophyll and primary production from two very different
213+
sources, and both are available:
214+
215+
| Name | Source | Resolution | Trade-off |
216+
|---|---|---|---|
217+
| `CHL`, `PP` | Copernicus-GlobColour, satellite | 4 km | Observed, finer — but surface-only and gappy under persistent cloud |
218+
| `CHL_MODEL`, `NPP_MODEL` | Biogeochemistry reanalysis | 0.25° | Gap-free and depth-resolved — but simulated, and coarser |
219+
220+
The plain names default to satellite, since the values are observed rather than
221+
simulated. Switch to the model versions where cloud gaps would matter more than
222+
resolution.
223+
224+
One caution: **satellite `PP` and model `NPP_MODEL` are not the same quantity.**
225+
`PP` is depth-integrated (mg/m2/day), `NPP_MODEL` volumetric (mg/m3/day).
226+
Substituting one for the other is a units error, not a resolution difference.
227+
228+
Phytoplankton functional types come from the same satellite plankton dataset as
229+
`CHL`, so they can be fetched together:
230+
231+
```{r pft, eval = FALSE}
232+
env <- accessEnvDat(
233+
vars = c("CHL", "DIATO", "DINO"), # one request, one dataset
234+
years = 2003:2017, months = 1:12,
235+
bounding_box = list(xmin = -76, xmax = -65, ymin = 35, ymax = 45)
236+
)
237+
```
238+
239+
`DIATO` and `DINO` are diatom and dinophyte chlorophyll — the spring-bloom
240+
species large copepods prefer, and the later stratified-water group
241+
respectively.
242+
243+
## Bottom salinity
244+
245+
`BOTS` pairs with `BOTT`, but it is not fetched the same way, because
246+
**GLORYS12V1 does not publish it.** The reanalysis has temperature at the sea
247+
floor and no salinity counterpart. So it is derived: the full salinity column is
248+
fetched and the deepest wet level in each cell kept.
249+
250+
```{r bots-readme, eval = FALSE}
251+
bots <- accessEnvDat(vars = "BOTS", years = 2010:2014, months = 1:12,
252+
bounding_box = bb)
253+
#> BOTS is not published by this product. Deriving it from the full 'so' column
254+
#> and keeping the deepest wet level in each cell; the depth used comes back as
255+
#> BOTS_depth.
256+
```
257+
258+
The depth each value came from is returned as `BOTS_depth` rather than left to
259+
be assumed. That matters because the deepest wet *model level* is not the sea
260+
floor: level spacing coarsens with depth, so in deep water the value can sit a
261+
long way above the bottom. In shelf water it is within a few metres.
262+
263+
Two consequences, both deliberate:
264+
265+
- **It must be fetched on its own.** The whole depth column is a different
266+
request from the single level `SST` wants, so mixing them is an error rather
267+
than a quiet second download. Call twice and chain `matchData()`.
268+
- **It costs far more.** Roughly fifty levels are downloaded over the same box to
269+
keep one, so a large box is much slower than the same box of `SST`.
270+
271+
In forecast mode none of this applies. The analysis-and-forecast product
272+
publishes sea-floor salinity outright, so `BOTS` is an ordinary variable there,
273+
fetches alongside `BOTT`, and returns no `BOTS_depth`:
274+
275+
```{r bots-forecast, eval = FALSE}
276+
accessEnvDat(vars = c("BOTT", "BOTS"), years = 2026, months = 8,
277+
bounding_box = bb, mode = "forecast")
278+
```
279+
280+
The two are the same quantity by construction but not the same number — one is
281+
the deepest level of a 50-level grid, the other Copernicus's own diagnostic — so
282+
a record spanning both modes has a seam in it.
283+
284+
```{r lookup, eval = FALSE}
285+
variable_dictionary("biogeochemical") # filter by product
286+
variable_dictionary("wind") # or just the winds
287+
fvcom_dictionary() # the FVCOM catalog, for accessFVCOM()
288+
fvcom_archives() # which FVCOM archives are built in
289+
hycom_dictionary() # the HYCOM catalog, for accessHYCOM()
290+
hycom_archives() # which HYCOM archives can be read
291+
hycom_covering("2019-06-15") # which of them span a given date
292+
ccmp_dictionary() # the CCMP catalog, for accessCCMP()
293+
ccmp_versions() # which CCMP versions can be read
294+
erddap_dictionary() # MUR and VIIRS, for accessERDDAP()
295+
erddap_datasets() # which ERDDAP datasets ship
296+
variable_dataset(c("SST", "CHL")) # which dataset each comes from
297+
as.data.frame(variable_dictionary())$description # full descriptions
298+
```

0 commit comments

Comments
 (0)