Skip to content

Session status: data-availability corpus scraped and merged; methods reframed around reuse #40

Description

@larnsce

Status note capturing a working session on the data-availability corpus, so the state and decisions are not lost. Links the pieces together; not a work item itself.

Shipped (merged via #36)

Headline finding

DAS presence is near-0% before ~2019, rising to ~90-98% by 2021-2022, consistently across all three journals. This temporal shift is exactly the within-journal signal the whole-journal indexing was built to capture.

Methods discoveries (reframed the remaining plan) - three open issues

The larger outcome was realising much of the planned build already exists as open tools/datasets. Shift from "build our own" to "reuse what exists, scrape only the gaps, validate against the overlap."

Open / not done

One-line: data collection is finished and merged; in the process we found that much of the remaining build is already solved by existing open tools/datasets, so the next phase is reuse-and-validate rather than build-from-scratch.

Refs #32, #33, #34, #35, #37, #38, #39

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions