@TomazErjavec
affected partners: @maria-gav @E2GG @Dandelliony @dareqx
After today's discussion, two related issues have arisen, so I am merging it in one:
- What single component file contains?
- single newspaper issue (CZ)
- one article (UA)
- small random fragment (PL)
Here, we want to somehow encode what is present in the file. Shouldn't we introduce some taxonomy that says simply what is inside a file (/TEI/@ana) or part of a newspaper (//div/@ana)?
- What type of "text input" should be included?
- newspaper
- newspaper supplement like attachments
- magazines
We need to set stricter boundaries on what should be included in corpora, and if we want to also include the edge types like newspaper supplements, then we need to categorize them, to be filtered out if users do not want them.
@TomazErjavec
affected partners: @maria-gav @E2GG @Dandelliony @dareqx
After today's discussion, two related issues have arisen, so I am merging it in one:
Here, we want to somehow encode what is present in the file. Shouldn't we introduce some taxonomy that says simply what is inside a file (
/TEI/@ana) or part of a newspaper (//div/@ana)?We need to set stricter boundaries on what should be included in corpora, and if we want to also include the edge types like newspaper supplements, then we need to categorize them, to be filtered out if users do not want them.