|
1 | 1 | # Salmon Data Standards Workshop |
2 | 2 |
|
3 | | -This workshop helps salmon biologists, data stewards, and data scientists turn the included **NuSEDS Fraser Coho sample** into a reviewable **Salmon Data Package**, following the same dataset from first inspection through metadata, semantic review, EML export, and a catalog dry run. No personal dataset is required; an optional activity near the end applies the workflow to your own data. The common biologist pathway is practical FAIR publication: preserve the context behind local data, link selected meanings to shared definitions, export reviewed metadata as a validated **EML 2.2 file**, and prepare or perform an authorized upload to an EML-aware catalog such as KNB. Depending on the learner's goals, the pathway can extend to proposing missing shared terms or planning how an organizational vocabulary or ontology will be governed and mapped to the Salmon Domain Ontology. Code-driven activities include R and Python examples, with separate spreadsheet instructions where a no-code path is useful. |
| 3 | +This six-hour workshop follows one **NuSEDS Fraser Coho 2023–2024 dataset** from human interpretation to a Salmon Data Package and a KNB Test Node publication exercise. Everyone uses the included `nuseds-fraser-coho-2023-2024.csv`: **173 rows and 14 columns**. No personal dataset is needed. |
4 | 4 |
|
5 | | -The material is being refactored for the Salmon Ontology Development Working Group from an ontology-development-first course into an SDP-first learning path. An ontology is a maintained set of concepts and definitions that also records how the concepts relate. Biologists can complete the common pathway by reusing shared definitions where they help; they do not need to build an ontology or give every field an ontology term. The advanced stewardship pathway addresses organizational vocabularies, ontologies, and bridge mappings for learners whose roles require them. |
| 5 | +Participants first draw how observations, results, populations, methods, and units relate. They write and peer-review a data dictionary before asking software to suggest mappings. They then compare human decisions with recorded AI outputs, document code meanings and gaps, and inspect validation, EML, and a test catalog record. The reference package is a technical draft pending Bruno and Tom's domain review; a passing software check does not supply that review. |
6 | 6 |
|
7 | | -## Learning Path |
| 7 | +## The six-hour sequence |
8 | 8 |
|
9 | | -1. **Structure first**: create a draft Salmon Data Package from the bundled `nuseds-fraser-coho-sample.csv` with R/`metasalmon` or Python/`metasalmonpy`; spreadsheet learners open a facilitator-generated package from that same sample. |
10 | | -2. **Context next**: write dataset, table, column, code, caveat, and method notes that travel with the data. |
11 | | -3. **Meaning where it matters**: review suggested term mappings, focusing first on measurement columns and important code lists. |
12 | | -4. **Contribution and stewardship paths**: route unresolved terms to the shared Salmon Domain Ontology, GC DFO Salmon Ontology, or a local/profile vocabulary or ontology; where needed, plan how organizational terms will be governed and mapped to shared anchors. |
13 | | -5. **Publication**: map the reviewed sample SDP into EML, validate the export, and preview the exact catalog deposit in a credential-free dry run. A live upload requires separate publication authority. |
14 | | -6. **Optional transfer**: near the end, start a separate draft package with your own dataset or make a transfer plan using the sample. |
| 9 | +| Chapter | Activity | Minutes | |
| 10 | +| --- | --- | ---: | |
| 11 | +| 1 | See the complete workflow and the shared data | 55 | |
| 12 | +| 2 | Draw a human concept graph | 55 | |
| 13 | +| 3 | Write a dictionary, decompose a measurement, and peer-review | 65 | |
| 14 | +| 4 | Build the Salmon Data Package | 40 | |
| 15 | +| 5 | Review mappings and compare human and AI reasoning | 70 | |
| 16 | +| 6 | Describe codes and record term gaps | 30 | |
| 17 | +| 7 | Validate, export EML, and inspect test publication | 45 | |
| 18 | +| | **Teaching and activities; add breaks and lunch** | **360** | |
15 | 19 |
|
16 | | -The shared example is the 30-row, 17-column teaching sample, covering selected years from 1996–2024. It stays under `raw_data/`, with the reproducible build in `scripts/build_sdp.R` or `scripts/build_sdp.py` and its generated package at `output/fraser-coho-example-sdp`. The separate 173-row, 2023–2024 example is outside this workshop pathway. |
| 20 | +## Start here |
17 | 21 |
|
18 | | -## Audience |
| 22 | +- [Setup](learners/setup.md): download the workshop kit and choose spreadsheet, R, or Python tools. |
| 23 | +- [Glossary](learners/glossary.md): plain-language terms linked from their first use in the lesson. |
| 24 | +- [Field reference](learners/field-reference.md): package files, required fields, and links to the canonical SDP specification. |
| 25 | +- [Reference and teaching record](learners/reference.md#teaching-record): checkpoints, software boundaries, and test catalog status. |
| 26 | +- [Instructor notes](instructors/instructor-notes.md): timing, preparation, and review criteria. |
| 27 | +- [Extended practice](learners/extended-practice.md): five optional labs using the same source, after the human graph and dictionary checkpoints. |
| 28 | +- [Advanced extension](learners/advanced.md): optional ontology formalization after the six-hour workshop, using the same human graph. |
19 | 29 |
|
20 | | -This workshop is designed for mixed groups: |
| 30 | +Extended practice is outside the 360-minute schedule. The labs investigate repeated population–year records, missingness and method context, code sources, a reviewed metadata edit across R and Python, and validation/EML/manifest evidence. They use the included data and local artifacts without requiring live AI or a deposit. |
21 | 31 |
|
22 | | -- operational salmon biologists who mostly work in Excel; |
23 | | -- data stewards standardizing datasets for sharing; |
24 | | -- R or Python users who want a reproducible SDP workflow; |
25 | | -- ontology maintainers who need better evidence from contributors. |
| 32 | +The project is `fraser-coho-workshop/`. Keep the source at `raw_data/nuseds-fraser-coho-2023-2024.csv`, the build at `scripts/build_sdp.R` or `scripts/build_sdp.py`, and the working package at `output/fraser-coho-workshop-sdp`. Dataset ID `fraser-coho-workshop` and table ID `escapement` remain consistent throughout. The kit contains `draft-sdp`, `seeded-sdp`, and `reference-sdp` checkpoints from the same data. |
26 | 33 |
|
27 | | -No terminology-standards background is assumed. Session 1 defines semantic links, vocabularies, code lists, and ontologies in plain language; later standards are introduced only when they help with a concrete review decision. |
| 34 | +## Tools and access |
28 | 35 |
|
29 | | -## R and Python implementations |
| 36 | +The lesson pins R `metasalmon` to **v0.5.0** and Python `metasalmonpy` to **v0.4.0**. The Python release does not yet provide the R-native review and metadata setters; its lane uses the documented alternatives and supplied evidence. Spreadsheet participants review the same files and decisions. |
30 | 37 |
|
31 | | -The R package `metasalmon` and Python package `metasalmonpy` are intended to remain behaviorally aligned, but the current workshop records an open catch-up window: R is pinned to `metasalmon` `v0.5.0`, while Python is pinned to `metasalmonpy` `v0.4.0`. Session 4's native semantic-review workflow is therefore taught only in R until the Python port lands. Examples use idiomatic syntax for each language rather than forcing literal API mimicry; deliberate differences are recorded in the [metasalmonpy parity guide](https://salmon-data-mobilization.github.io/metasalmonpy/guides/parity.html). |
| 38 | +The human–AI comparison is required, using supplied recorded outputs. Live AI calls are optional. Participants need no AI account, API key, purchased credits, or catalog account to complete the workshop. Optional OpenRouter practice uses a free model with no paid fallback; optional Ollama practice can use a local model. |
32 | 39 |
|
33 | | -Python semantic seeding on this sample fails with the tested `metasalmonpy 0.4.0` / pandas `3.0.5` combination. Session 3 uses facilitator-supplied R candidate evidence for the same sample; Python creation and metadata editing remain local. Replace that handoff only after a released combination passes the seeded sample rebuild and the workshop pins are updated. |
| 40 | +Catalog work uses a separately authorized **KNB Test Node** teaching record under Brett's identity. It does not authorize a production deposit or imply Bruno and Tom have reviewed the scientific meanings. See the [teaching-record status](learners/reference.md#teaching-record) before presenting a live result. |
34 | 41 |
|
35 | | -## Formats |
| 42 | +## Repository maintenance |
36 | 43 |
|
37 | | -The same materials support two delivery modes: |
| 44 | +Edit lesson sources under `episodes/` and `learners/`; do not edit generated `site/` output. Kit source files are under `episodes/files/fraser-coho-workshop/`, and the downloadable ZIP is `episodes/files/fraser-coho-workshop.zip`. [Entrypoints](docs/entrypoints.md) records the source map, build commands, and checks. |
38 | 45 |
|
39 | | -- **One-hour introduction**: end-goal framing, SDP anatomy and example CSVs, a short package demo, one measurement mapping review, and an SDP-to-EML/catalog preview. |
40 | | -- **Full-day workshop**: hands-on package creation, context capture, mapping review, measurement decomposition, code-list review, term-request planning, EML export, and a credential-free KNB publication dry run, followed by an optional bring-your-own-dataset activity. |
41 | | - |
42 | | -## Repository Contents |
43 | | - |
44 | | -- `episodes/`: learner-facing workshop sessions. |
45 | | -- `learners/setup.md`: setup guidance for R/metasalmon, Python/metasalmonpy, and spreadsheet participants. |
46 | | -- `learners/reference.md`: glossary, decision aids, and core workflow checks. |
47 | | -- `instructors/instructor-notes.md`: facilitation plans for one-hour and full-day delivery. |
48 | | -- `profiles/learner-profiles.md`: persona notes for designing and testing the workshop. |
49 | | -- `docs/entrypoints.md`: short map of the lesson entry points and local checks. |
50 | | - |
51 | | -## Related Components |
52 | | - |
53 | | -- [Salmon Data Package specification](https://github.com/salmon-data-mobilization/smn-data-pkg) |
54 | | -- [metasalmon R package](https://github.com/salmon-data-mobilization/metasalmon) |
55 | | -- [metasalmonpy Python package](https://github.com/salmon-data-mobilization/metasalmonpy) |
56 | | -- [Salmon Domain Ontology](https://github.com/salmon-data-mobilization/salmon-domain-ontology) |
57 | | -- [GC DFO Salmon Ontology](https://github.com/dfo-pacific-science/dfo-salmon-ontology) |
58 | | - |
59 | | -## Development Status |
60 | | - |
61 | | -This lesson is under active development. |
62 | | - |
63 | | -- Explore the **Issues** tab to find ways to contribute. |
64 | | -- Join our discussions on the **SDM Discord Server**. |
65 | | -- Share your feedback to improve and expand the workshop's impact. |
66 | | - |
67 | | ---- |
68 | | - |
69 | | -For more information, visit the [Salmon Data Mobilization GitHub Organization](https://github.com/salmon-data-mobilization) or contact us directly. |
| 46 | +The [Salmon Data Package specification](https://github.com/salmon-data-mobilization/smn-data-pkg/blob/main/SPECIFICATION.md), [metasalmon](https://github.com/salmon-data-mobilization/metasalmon), and [metasalmonpy](https://github.com/salmon-data-mobilization/metasalmonpy) own the underlying formats and software contracts. |
0 commit comments