Skip to content

Commit 190df30

Browse files
committed
Rebuild workshop around Fraser Coho human interpretation workflow
1 parent a0d4628 commit 190df30

137 files changed

Lines changed: 11627 additions & 3005 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.gitattributes‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,9 @@
1+
# The teaching kit has a byte-level source and checksum inventory contract.
2+
# Preserve its recorded encodings and line endings on every checkout platform.
3+
/episodes/files/fraser-coho-workshop/** -text
4+
/episodes/files/fraser-coho-workshop.zip -text
5+
# The official CSV has CRLF record endings (and embedded LF inside fields).
6+
# Field W also ends in a source-provided space. Preserve that source value;
7+
# retire these whitespace settings when a deliberately refreshed source no
8+
# longer contains those endings. Never normalize the archived bytes.
9+
/episodes/files/fraser-coho-workshop/raw_data/official-nuseds-dictionary.csv whitespace=cr-at-eol,-blank-at-eol

‎.github/workflows/sandpaper-main.yaml‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -37,6 +37,11 @@ jobs:
3737
- name: "Checkout Lesson"
3838
uses: actions/checkout@v4
3939

40+
- name: "Check fixed dataset and teaching kit"
41+
# This checks distributed bytes and lesson artifacts. Scientific review,
42+
# the independent SDP validator and a live catalog need separate evidence.
43+
run: python3 scripts/check-workshop.py
44+
4045
- name: "Set up R"
4146
uses: r-lib/actions/setup-r@v2
4247
with:

‎README.md‎

Lines changed: 30 additions & 53 deletions
Original file line numberDiff line numberDiff line change
@@ -1,69 +1,46 @@
11
# Salmon Data Standards Workshop
22

3-
This workshop helps salmon biologists, data stewards, and data scientists turn the included **NuSEDS Fraser Coho sample** into a reviewable **Salmon Data Package**, following the same dataset from first inspection through metadata, semantic review, EML export, and a catalog dry run. No personal dataset is required; an optional activity near the end applies the workflow to your own data. The common biologist pathway is practical FAIR publication: preserve the context behind local data, link selected meanings to shared definitions, export reviewed metadata as a validated **EML 2.2 file**, and prepare or perform an authorized upload to an EML-aware catalog such as KNB. Depending on the learner's goals, the pathway can extend to proposing missing shared terms or planning how an organizational vocabulary or ontology will be governed and mapped to the Salmon Domain Ontology. Code-driven activities include R and Python examples, with separate spreadsheet instructions where a no-code path is useful.
3+
This six-hour workshop follows one **NuSEDS Fraser Coho 2023–2024 dataset** from human interpretation to a Salmon Data Package and a KNB Test Node publication exercise. Everyone uses the included `nuseds-fraser-coho-2023-2024.csv`: **173 rows and 14 columns**. No personal dataset is needed.
44

5-
The material is being refactored for the Salmon Ontology Development Working Group from an ontology-development-first course into an SDP-first learning path. An ontology is a maintained set of concepts and definitions that also records how the concepts relate. Biologists can complete the common pathway by reusing shared definitions where they help; they do not need to build an ontology or give every field an ontology term. The advanced stewardship pathway addresses organizational vocabularies, ontologies, and bridge mappings for learners whose roles require them.
5+
Participants first draw how observations, results, populations, methods, and units relate. They write and peer-review a data dictionary before asking software to suggest mappings. They then compare human decisions with recorded AI outputs, document code meanings and gaps, and inspect validation, EML, and a test catalog record. The reference package is a technical draft pending Bruno and Tom's domain review; a passing software check does not supply that review.
66

7-
## Learning Path
7+
## The six-hour sequence
88

9-
1. **Structure first**: create a draft Salmon Data Package from the bundled `nuseds-fraser-coho-sample.csv` with R/`metasalmon` or Python/`metasalmonpy`; spreadsheet learners open a facilitator-generated package from that same sample.
10-
2. **Context next**: write dataset, table, column, code, caveat, and method notes that travel with the data.
11-
3. **Meaning where it matters**: review suggested term mappings, focusing first on measurement columns and important code lists.
12-
4. **Contribution and stewardship paths**: route unresolved terms to the shared Salmon Domain Ontology, GC DFO Salmon Ontology, or a local/profile vocabulary or ontology; where needed, plan how organizational terms will be governed and mapped to shared anchors.
13-
5. **Publication**: map the reviewed sample SDP into EML, validate the export, and preview the exact catalog deposit in a credential-free dry run. A live upload requires separate publication authority.
14-
6. **Optional transfer**: near the end, start a separate draft package with your own dataset or make a transfer plan using the sample.
9+
| Chapter | Activity | Minutes |
10+
| --- | --- | ---: |
11+
| 1 | See the complete workflow and the shared data | 55 |
12+
| 2 | Draw a human concept graph | 55 |
13+
| 3 | Write a dictionary, decompose a measurement, and peer-review | 65 |
14+
| 4 | Build the Salmon Data Package | 40 |
15+
| 5 | Review mappings and compare human and AI reasoning | 70 |
16+
| 6 | Describe codes and record term gaps | 30 |
17+
| 7 | Validate, export EML, and inspect test publication | 45 |
18+
| | **Teaching and activities; add breaks and lunch** | **360** |
1519

16-
The shared example is the 30-row, 17-column teaching sample, covering selected years from 1996–2024. It stays under `raw_data/`, with the reproducible build in `scripts/build_sdp.R` or `scripts/build_sdp.py` and its generated package at `output/fraser-coho-example-sdp`. The separate 173-row, 2023–2024 example is outside this workshop pathway.
20+
## Start here
1721

18-
## Audience
22+
- [Setup](learners/setup.md): download the workshop kit and choose spreadsheet, R, or Python tools.
23+
- [Glossary](learners/glossary.md): plain-language terms linked from their first use in the lesson.
24+
- [Field reference](learners/field-reference.md): package files, required fields, and links to the canonical SDP specification.
25+
- [Reference and teaching record](learners/reference.md#teaching-record): checkpoints, software boundaries, and test catalog status.
26+
- [Instructor notes](instructors/instructor-notes.md): timing, preparation, and review criteria.
27+
- [Extended practice](learners/extended-practice.md): five optional labs using the same source, after the human graph and dictionary checkpoints.
28+
- [Advanced extension](learners/advanced.md): optional ontology formalization after the six-hour workshop, using the same human graph.
1929

20-
This workshop is designed for mixed groups:
30+
Extended practice is outside the 360-minute schedule. The labs investigate repeated population–year records, missingness and method context, code sources, a reviewed metadata edit across R and Python, and validation/EML/manifest evidence. They use the included data and local artifacts without requiring live AI or a deposit.
2131

22-
- operational salmon biologists who mostly work in Excel;
23-
- data stewards standardizing datasets for sharing;
24-
- R or Python users who want a reproducible SDP workflow;
25-
- ontology maintainers who need better evidence from contributors.
32+
The project is `fraser-coho-workshop/`. Keep the source at `raw_data/nuseds-fraser-coho-2023-2024.csv`, the build at `scripts/build_sdp.R` or `scripts/build_sdp.py`, and the working package at `output/fraser-coho-workshop-sdp`. Dataset ID `fraser-coho-workshop` and table ID `escapement` remain consistent throughout. The kit contains `draft-sdp`, `seeded-sdp`, and `reference-sdp` checkpoints from the same data.
2633

27-
No terminology-standards background is assumed. Session 1 defines semantic links, vocabularies, code lists, and ontologies in plain language; later standards are introduced only when they help with a concrete review decision.
34+
## Tools and access
2835

29-
## R and Python implementations
36+
The lesson pins R `metasalmon` to **v0.5.0** and Python `metasalmonpy` to **v0.4.0**. The Python release does not yet provide the R-native review and metadata setters; its lane uses the documented alternatives and supplied evidence. Spreadsheet participants review the same files and decisions.
3037

31-
The R package `metasalmon` and Python package `metasalmonpy` are intended to remain behaviorally aligned, but the current workshop records an open catch-up window: R is pinned to `metasalmon` `v0.5.0`, while Python is pinned to `metasalmonpy` `v0.4.0`. Session 4's native semantic-review workflow is therefore taught only in R until the Python port lands. Examples use idiomatic syntax for each language rather than forcing literal API mimicry; deliberate differences are recorded in the [metasalmonpy parity guide](https://salmon-data-mobilization.github.io/metasalmonpy/guides/parity.html).
38+
The human–AI comparison is required, using supplied recorded outputs. Live AI calls are optional. Participants need no AI account, API key, purchased credits, or catalog account to complete the workshop. Optional OpenRouter practice uses a free model with no paid fallback; optional Ollama practice can use a local model.
3239

33-
Python semantic seeding on this sample fails with the tested `metasalmonpy 0.4.0` / pandas `3.0.5` combination. Session 3 uses facilitator-supplied R candidate evidence for the same sample; Python creation and metadata editing remain local. Replace that handoff only after a released combination passes the seeded sample rebuild and the workshop pins are updated.
40+
Catalog work uses a separately authorized **KNB Test Node** teaching record under Brett's identity. It does not authorize a production deposit or imply Bruno and Tom have reviewed the scientific meanings. See the [teaching-record status](learners/reference.md#teaching-record) before presenting a live result.
3441

35-
## Formats
42+
## Repository maintenance
3643

37-
The same materials support two delivery modes:
44+
Edit lesson sources under `episodes/` and `learners/`; do not edit generated `site/` output. Kit source files are under `episodes/files/fraser-coho-workshop/`, and the downloadable ZIP is `episodes/files/fraser-coho-workshop.zip`. [Entrypoints](docs/entrypoints.md) records the source map, build commands, and checks.
3845

39-
- **One-hour introduction**: end-goal framing, SDP anatomy and example CSVs, a short package demo, one measurement mapping review, and an SDP-to-EML/catalog preview.
40-
- **Full-day workshop**: hands-on package creation, context capture, mapping review, measurement decomposition, code-list review, term-request planning, EML export, and a credential-free KNB publication dry run, followed by an optional bring-your-own-dataset activity.
41-
42-
## Repository Contents
43-
44-
- `episodes/`: learner-facing workshop sessions.
45-
- `learners/setup.md`: setup guidance for R/metasalmon, Python/metasalmonpy, and spreadsheet participants.
46-
- `learners/reference.md`: glossary, decision aids, and core workflow checks.
47-
- `instructors/instructor-notes.md`: facilitation plans for one-hour and full-day delivery.
48-
- `profiles/learner-profiles.md`: persona notes for designing and testing the workshop.
49-
- `docs/entrypoints.md`: short map of the lesson entry points and local checks.
50-
51-
## Related Components
52-
53-
- [Salmon Data Package specification](https://github.com/salmon-data-mobilization/smn-data-pkg)
54-
- [metasalmon R package](https://github.com/salmon-data-mobilization/metasalmon)
55-
- [metasalmonpy Python package](https://github.com/salmon-data-mobilization/metasalmonpy)
56-
- [Salmon Domain Ontology](https://github.com/salmon-data-mobilization/salmon-domain-ontology)
57-
- [GC DFO Salmon Ontology](https://github.com/dfo-pacific-science/dfo-salmon-ontology)
58-
59-
## Development Status
60-
61-
This lesson is under active development.
62-
63-
- Explore the **Issues** tab to find ways to contribute.
64-
- Join our discussions on the **SDM Discord Server**.
65-
- Share your feedback to improve and expand the workshop's impact.
66-
67-
---
68-
69-
For more information, visit the [Salmon Data Mobilization GitHub Organization](https://github.com/salmon-data-mobilization) or contact us directly.
46+
The [Salmon Data Package specification](https://github.com/salmon-data-mobilization/smn-data-pkg/blob/main/SPECIFICATION.md), [metasalmon](https://github.com/salmon-data-mobilization/metasalmon), and [metasalmonpy](https://github.com/salmon-data-mobilization/metasalmonpy) own the underlying formats and software contracts.

‎config.yaml‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ carpentry: 'incubator'
1717
# carpentry_description: "Custom Carpentry"
1818

1919
# Overall title for pages.
20-
title: 'Creating Salmon Data Packages With Shared Standards'
20+
title: 'Salmon Data Standards Workshop'
2121

2222
# Date the lesson was created (YYYY-MM-DD, this is empty by default)
2323
created: '2025-01-22'
@@ -72,12 +72,15 @@ episodes:
7272
- session-5.Rmd
7373
- session-6.Rmd
7474
- session-7.Rmd
75-
- bonus-session.Rmd
7675

7776
# Information for Learners
7877
learners:
7978
- setup.md
79+
- glossary.md
80+
- field-reference.md
8081
- reference.md
82+
- extended-practice.md
83+
- advanced.md
8184

8285
# Information for Instructors
8386
instructors:

0 commit comments

Comments
 (0)