Skip to content

Integration #3

Description

@meehkal

This subcorpus should not grow separated from our "data backbone". Ideally, we maintain and develop a complete set of data, metadata and annotations for in a private folder and derive the gold subcorpus from there automatically. However, until the full implementation of this interaction, we build this gold subcorpus separated from the other data repositories.

Ideas:

  • For gold corpus annotations, we use separate ELAN tiers because they need to be added or corrected manually; the annotations of these tiers may undergo permanent changes (additions, corrections, etc.), but these changes are only done in the "data backbone working repository"; in the future, the present repository get its data automatically from the working repository
  • tiers: ref@ orth@ ft-eng@ word-gold@ lemma-gold@ pos-gold@ morph-gold@
  • For gold corpus metadata, we define a minimal subset of metadata categories; they are also derived automatically (see above) and included completely into the EAF file (either as attributes or in tier names)
  • Since our session names are unique we can include a link to the corresponding sessions in the CMDI catalogue in order to provide more complete metadata as well as multimedia or additional corpus annotations (if available)

Metadata

Metadata

Labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions