diff --git a/Build/Sources-TEI/Makefile b/Build/Sources-TEI/Makefile new file mode 100644 index 000000000..16e5434f9 --- /dev/null +++ b/Build/Sources-TEI/Makefile @@ -0,0 +1,2 @@ +init-Sources-TEI: + rsync -av ../Distro/ParlaMint-*.TEI* . diff --git a/Build/Taxonomies/taxonomy-translation-include.tsv b/Build/Taxonomies/taxonomy-translation-include.tsv index d65c18ff3..df49f8f9c 100644 --- a/Build/Taxonomies/taxonomy-translation-include.tsv +++ b/Build/Taxonomies/taxonomy-translation-include.tsv @@ -5,6 +5,7 @@ ca ParlaMint-ES-CT cs ParlaMint-CZ da ParlaMint-DK de ParlaMint-AT +de ParlaMint-DE el ParlaMint-GR es ParlaMint-ES es ParlaMint-ES-CN diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index bb00cc58c..784eaff08 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -53,6 +53,7 @@ flowchart TB - [Create a GitHub account](https://github.com/signup) if you don't have one. - [Fork ParlaMint repository](https://github.com/clarin-eric/ParlaMint/fork) into your organization or private account. + - **uncheck "Copy the main branch only"** - Start the terminal on your computer and navigate to the folder where you want the ParlaMint local clone of the repository to be placed: ```bash @@ -60,8 +61,7 @@ flowchart TB git clone git@github.com:/ParlaMint.git ``` -- Set the data branch in your repository to be synchronized with the data branch in the ParlaMint repository: - +- (ONLY WHEN YOU FORGET TO UNCHECK "Copy the main branch only") Set the data branch in your repository to be synchronized with the data branch in the ParlaMint repository: ```bash cd ParlaMint git remote add upstream https://github.com/clarin-eric/ParlaMint.git diff --git a/Makefile b/Makefile index 240cbc056..6e382c7c1 100644 --- a/Makefile +++ b/Makefile @@ -5,8 +5,11 @@ #Special case #PARLIAMENTS = IL +#All Parliaments +PARLIAMENTS = AT BE BG CZ DE DK EE ES ES-CN ES-CT ES-GA ES-PV FI FR GB GR HR HU IL IS IT LV NL NO PL PT SE SI TR BA RS UA + #Parliaments for V5.0 -PARLIAMENTS = AT BE BG CZ DK EE ES ES-CT ES-GA ES-PV FI FR GB GR HR HU IS IT LV NL NO PL PT SE SI TR BA RS UA +# PARLIAMENTS = AT BE BG CZ DK EE ES ES-CT ES-GA ES-PV FI FR GB GR HR HU IS IT LV NL NO PL PT SE SI TR BA RS UA ##$JAVA-MEMORY## Set a java memory maxsize in GB JAVA-MEMORY = @@ -99,6 +102,9 @@ setup-parliament-newInParlaMint2: setup-parliament-CanaryIslands: make setup-parliament PARLIAMENT-NAME='Canary Islands' PARLIAMENT-CODE='ES-CN' LANG-LIST='es (Spanish)' +setup-parliament-Germany: + make setup-parliament PARLIAMENT-NAME='Germany' PARLIAMENT-CODE='DE' LANG-LIST='de (German)' + ## initTaxonomies-XX ## initialize taxonomies in folder ParlaMint-XX #### parameter LANG-CODE-LIST can contain space separated list of languages initTaxonomies-XX = $(addprefix initTaxonomies-, $(PARLIAMENTS)) diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-CHES.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-CHES.xml new file mode 100644 index 000000000..eed007cb4 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-CHES.xml @@ -0,0 +1,395 @@ + + + + CHES variables: Taxonomy of identifiers from the Chapel Hill Expert Survey (CHES) trend files: 1999-2019 Codebook and 2019 Codebook. + + + General indicators: General indicators used to identify basic characteristics of a political party, such as country of origin, number of experts evaluating a party, membership in the EU, etc. + + + Country of origin: Unique identifier for each country. For a full list of country identifiers, consult the 1999-2019 Chapel Hill Expert Survey (CHES) trend file Codebook. Note that ths category is not present in the ParlaMint CHES data, as the country is given explicitly for each corpus. + + + + Eastern-western origin: Origin of a political party, indicating whether the party originates from Central/Eastern Europe or whether it is one of the original 15 members of the EU (EU-15). Values: 0 = party from Central/Eastern Europe, 1 = party from EU-15. + + + + EU Membership status: political party membership status in European Union during a specific year. Values: 0 = not an EU member in that year, 1 = EU member in that year. + + + + Year of evaluation: year for which party experts were asked to evaluate: 1999, 2002, 2006, 2010, 2014, 2019. Note that ths category is not present in the ParlaMint CHES data, as the year is given explicitly for each variable value. + + + + Number of experts: number of experts who evaluated a party. + + + + Party identifier: unique identifier for each party. For full list of party identifiers, consult the 1999-2019 Chapel Hill Expert Survey (CHES) trend file Codebook. Note that ths category is present in the ParlaMint CHES data as the value of the top level state/@key attribute. + + + + Party abbreviation: the abbreviated name of the party. Note that one Party identifier can correspond to several Party abbreviations, depending on the year of the survey. For full list of Party identifiers matched with the Party identifiers consult the 1999-2019 Chapel Hill Expert Survey (CHES) trend file Codebook. Note that ths category is present in the ParlaMint CHES data also as the value of the top level state/@n attribute; when there are more than one Party appreviations, they are separated by space. + + + + Comparative Manifesto Project party identificator: political party identificator according to Comparative Manifesto Project. Data: https://visuals.manifesto-project.wzb.eu/mpdb-shiny/cmp_dashboard_dataset/; Codebook: https://manifesto-project.wzb.eu/down/data/2023a/codebooks/codebook_MPDataset_MPDS2023a.pdf + + + + Vote percentage: vote percentage received by the party in the national election most prior to the specified year. + + + + Seat share: share of parliamentary seats won by a political party in the national election prior to the specified year. + + + + Year of the previous national election: year of the national election most prior to the specified year. + + + + European Parliament Vote Percentage: percentage of votes acquired by political party in the European Parliament election most prior to specified year. + + + + Party family: categorization of political parties into distinct families, based on their shared characteristics, ideologies and affiliations or other factors. Initially inspired by the Hix and Lord (1997) framework, this classification differentiates parties with a focus on various factors. In this classification, confessional and agrarian parties are placed into separate categories. For parties in Central/Eastern Europe, the Derksen classification (now integrated into Wikipedia) is used as a foundation, complemented by two key methods: a) assessing membership or ties to international and EU party associations, and b) self-identification of a party. This classification is periodically updated to account for shifts in party ideologies or organizational changes. Values: 1 = Radical Right (RADRT), 2 = Conservatives (CON), 3 = Liberal (LIB), 4 = Christian-Democratic (CD), 5 = Socialist (SOC), 6 = Radical Left (RADLEFT), 7 = Green (GRREN), 8 = Regionalist (REG), 9 = No family (NOFAMILY), 10 = Confessional (CONFESS), 11 = Agrarian/Center (AGRARIAN/CENTER). + + + + Government participation: participation of a party in the government during a specific year. Values: 0 = Party not in government, 0.5 = Party in government for part of the year, 1 = Party in government full year. + + + + + European Integration: Indicators of political parties' positions on European integration. Unless otherwise indicated, questions are available for the period 1999 - 2019. + + + European integration orientation: overall orientation of the party leadership towards European integration in a specific year. Values range on a scale from 1 - 7, 1 = Strongly opposed, 7 = Strongly in favor. + + + + European integration standard deviation: standard deviation of expert placement of overall orientation of the party leadership towards European integration in 2019. + + + + European integration salience: relative salience of European integration in the party’s public stance in a specific year. Values range from 0 to 10, 0 = European Integration is of no importance, never mentioned, 10 = European Integration is the most important issue. + + + + European integration dissent: degree of dissent on European integration in a specific year. Values range on a scale from 0 to 10; 0 = Party was completely united, 10 = Party was extremely divided. + + + + European integration blur: how blurry was each party’s position on on European integration (Asked in 2019). Values range on a scale from 0 to 10; 0 = Not at all blurred, 10 = Extremely blurred. + + + + European integration benefit: position of the party leadership in a specific year on whether a specific country has benefited from being a member of the EU (Asked in 2010 and 2014). Values range on a scale from 1 - 3, 1 = Benefited, 2 = Neither benefited nor lost, 3 = Not benefited. + + + + + European Policies: Indicators of political parties' positions on European policies. Unless otherwise indicated, questions are available for the period 1999 - 2019. + + + European Parliament power: position of the party leadership in a specific year on the powers of the European Parliament (Not asked in 2019). Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU tax harmonization: position of the party leadership in a specific year on EU tax harmonization to reduce regime competition (only asked in 1999). Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU internal market: position of the party leadership in a specified year on the internal market (i.e. free movement of goods, services, capital and labor) (Not asked in 1999). Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU common employment policy: position of the party leadership in a specified year on a common employment policy (only asked in 1999 and 2002). Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU economic and budgetary policies: position of the party leadership in a specified year on EU authority over member states’ economic and budgetary policies. Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU agricultural spending: position of the party leadership in a specific year on EU’s agricultural spending. Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU cohesion policy: position of the party leadership in a specified year on EU cohesion or regional policy (e.g. the structural funds). Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU environmental policy: position of the party leadership in a specific year on a common policy on the environment. Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU political asylum policy: position of the party leadership in a specific year on a common policy on political asylum. Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU foreign and security policy: position of the party leadership in a specific year on EU foreign and security policy. Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + EU enlargement to Turkey: position of the party leadership in a specified year on EU enlargement to Turkey. Values range from 1 - 7, 1 = Strongly opposes, 7 = Strongly favors. + + + + + Ideological positions: Indicators of the political party's ideological stance on general, economic, social, and cultural issues. Unless otherwise indicated, questions are available for the period 1999 - 2019. + + + Orientation position: position of political party in terms of its overall ideological stance. Values on a scale from 0 to 10, 0 = extreme left, 5 = centre, 10 = extreme right + + + + Economic ideological position: position of political party in terms of its ideological stance on economic issues. Parties can be classified in terms of their stance on economic issues such as privatization, taxes, regulation, government spending, and the welfare state. Parties on the economic left want government to play an active role in the economy. Parties on the economic right want a reduced role for government. Values on a scale from 0 to 10: 0 = extreme left, 5 = center, 10 = extreme right. + + + + Economic ideological position salience: significance of economic issues in a political party's public stance during a specific year (only asked in 2014 and 2019). Values range on a scale from 0 to 10, 0 = No importance, 10 = Great importance + + + + Economic ideological position dissent: the level of internal disagreement within a political party regarding economic issues during a specific year, with responses collected in 2019 (only asked in 2019). Values on a scale from 0 to 10, 0 = Party was completely united, 10 = Party was extremely divided. + + + + Economic ideological position blurriness: how blurry was each party’s position on economic issues in a specific year (only asked in 2019). Values on a scale from 0 to 10;; 0 = Not at all blurred, 10 = Extremely blurred. + + + + Economic ideological position standard deviation: standard deviation of expert placement of the party in 2019 in terms of its ideological stance on economic issues. + + + + Social and cultural ideological position: ideological stance of a political party concerning social and cultural values. "Libertarian" or "postmaterialist" parties favor expanded personal freedoms, for example, abortion rights, divorce, and same-sex marriage. "Traditional" or "authoritarian" parties reject these ideas in favor of order, tradition, and stability, believing that the government should be a firm moral authority on social and cultural issues. Values range on a scale from 0 to 10; 0 = Libertarian/Postmaterialist, 5 = center, 10 = Traditional/Authoritarian + + + + Social and cultural ideological position salience: relative salience of libertarian/traditional issues in the party’s public stance in specific year (only asked in 2014 and 2019). Values range on a scale from 0 to 10; 0 = No importance, 10 = Great importance + + + + Social and cultural ideological position dissent: degree of dissent on libertarian/traditional issues in specific year (only asked in 2019). Values range on a scale from 0 to 10; 0 = Party was completely united, 10 = Party was extremely divided. + + + + Social and cultural ideological position blur: how blurry was each party’s position on libertarian/traditional issues in specific year (only asked in 2019). Values range on a scale from 0 to 10; 0 = Not at all blurred, 10 = Extremely blurred. + + + + Social and cultural ideological position standard deviation: standard deviation of expert placement of the party in 2019 in terms of their views on democratic freedoms and rights. + + + + + Policy dimensions: Indicators of the political party's stance on various policy issues, such as public services, taxes, market regulation, civil liberties, religious principles, etc. Unless otherwise indicated, questions are available for a period 2006 - 2019. + + + Pubic services vs. tax reduction: political party's position on improving public services vs. reducing taxes. Values range on a scale from 0 to 10; 0 = Strongly favors improving public services, 10 = Strongly favors reducing taxes. + + + + Pubic services vs. tax reduction salience: importance/salience of improving public services vs. reducing taxes (not asked in 2014 or 2019). Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Deregulation of markets: political party's perspective on the deregulation of markets. Values range on a scale from 0 to 10; 0 = Strongly opposes deregulation of markets, 10 = Strongly supports deregulation of markets. + + + + Deregulation of markets salience: importance/salience of deregulation of markets. Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Redistribution of wealth: political party's position on redistribution of wealth from the rich to the poor. Values range on a scale from 0 to 10; 0 = Strongly favors redistribution, 10 = Strongly opposes redistribution. + + + + Redistribution of wealth salience: importance/salience of redistribution. (Not asked in 2014) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + State intervention in the economy: position on state intervention in the economy (only asked in 2014 and 2019). Values range on a scale from 0 to 10; 0 = Fully in favor of state intervention; 10 = Fully opposed to state intervention. + + + + Civil Liberties vs. Law and Order: position on civil liberties vs. law and order. Values range on a scale from 0 to 10; 0 = Strongly promotes civil liberties, 10 = Strongly supports tough measures to fight crime. + + + + Civil Liberties vs. Law and Order salience: importance/salience of civil liberties vs. law and order. (Not asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Social lifestyle: position on social lifestyle (e.g. rights for homosexuals, gender equality). Values range on a scale from 0 to 10; 0 = Strongly supports liberal policies, 10 = Strongly opposes liberal policies. + + + + Social lifestyle salience: importance/salience of social lifestyle issues (e.g. homosexuality). (Not asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Religious principles: position on role of religious principles in politics. Values range on a scale from 0 to 10; 0 = Strongly opposes religious principles in politics, 10 = Strongly supports religious principles in politics. + + + + Religious principles salience: importance/salience of religious principles in politics. (Not asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Immigration policy: position on immigration policy. Values range on a scale from 0 to 10; 0 = Strongly favors a liberal policy on immigration, 10 = Strongly favors a restrictive policy on immigration. + + + + Immigration policy salience: importance/salience of immigration policy. (Not asked in 2014) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Immigration policy dissent: degree of dissent on immigration policy in a specified year. (only asked in 2019). Values range on a scale from 0 to 10; 0 = Party was completely united, 10 = Party was extremely divided. + + + + Integration of immigrants: position on integration of immigrants and asylum seekers (multiculturalism vs. assimilation). Values range on a scale from 0 to 10; 0 = Strongly favors multiculturalism, 10 = Strongly favors assimilation. + + + + Integration of immigrants salience: importance/salience of integration of immigrants and asylum seekers. (Not asked in 2014) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Integration of immigrants dissent: degree of dissent on immigrants and asylum seekers issues in a specified year (only asked in 2019). Values range on a scale from 0 to 10; 0 = Party was completely united, 10 = Party was extremely divided. + + + + Urban vs. Rural: position on urban vs. rural interests. Values range on a scale from 0 to 10; 0 = Strongly supports urban interests, 10 = Strongly supports rural interests. + + + + Urban vs. Rural salience: importance/salience of urban vs. rural interests. (Not asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Environmental sustainability: position towards environmental sustainability. (Asked in 2010, 2014, and 2019) Values range on a scale from 0 to 10; 0 = Strongly supports environmental protection even at the cost of economic growth, 10 = Strongly supports economic growth even at the cost of environmental protection. + + + + Environmental sustainability salience: importance/salience of environmental sustainability. (Asked in 2010 and 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Cosmopolitanism vs. nationalism: position on cosmopolitanism vs. nationalism (only asked in 2006). Values range on a scale from 0 to 10; 0 = Strongly advocates cosmopolitanism, 10 = Strongly advocates nationalism. + + + + Cosmopolitanism vs. nationalism salience: importance/salience of cosmopolitanism vs. nationalism (only asked in 2006). Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Trade liberalization vs. protectionism: position towards trade liberalization/protectionism (only asked in 2019). Values range on a scale from 0 to 10; 0 = Strongly favors trade liberalization, 10 = Strongly favors protection of domestic producers. + + + + Regional decentralization: position on political decentralization to regions/localities. Values range on a scale from 0 to 10; 0 = Strongly favors political decentralization, 10 = Strongly opposes political decentralization. + + + + Regional decentralization salience: importance/salience of decentralization. (Not asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + International security: position towards international security and peacekeeping missions. (Asked in 2010 and 2014) Values range on a scale from 0 to 10; 0 = Strongly favors troop deployment from their own country, 10 = Strongly opposes troop deployment from their own country. + + + + International security salience: importance/salience of international security and peacekeeping missions (only asked in 2010). Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + US power: position towards US power in world affairs (only asked in 2006). Values range on a scale from 0 to 10; 0 = Strongly opposes strong US leadership in world affairs, 10 = Strongly favors strong US leadership in world affairs. + + + + US power salience: importance/salience of US power in world affairs (only asked in 2006). Values range on a scale from 0 to 10; 0 = not important at all, 10 = extremely important. + + + + Ethnic minorities: position towards ethnic minorities. Values range on a scale from 0 to 10; 0 = Strongly supports more rights for ethnic minorities, 10 = Strongly opposes more rights for ethnic minorities. + + + + Ethnic minorities salience: importance/salience of ethnic minorities. (Not asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Cosmopolitanism vs. nationalism position: position towards cosmopolitanism vs. nationalism. (Asked in 2014 or 2019) Values range on a scale from 0 to 10; 0 = Strongly promotes cosmopolitan conceptions of society, 10 = Strongly promotes nationalist conceptions of society. + + + + Anti-Islam Rhetoric salience: salience of anti-Islam rhetoric for the party leadership. Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Russian interference salience: salience of Russian interference in domestic affairs for the party leadership. Values range on a scale from 0 to 10; 0 = No importance, 10 = Great importance. + + + + + Party characteristics: Indicators of a political party's positions on direct people participation, anti-elite sentiment, anti-corruption, and the role of members vs. leadership in policy decisions. + + + People vs. elite: position on people vs elected representatives. Some political parties take the position that ‘the people’ should have the final say on the most important issues, for example, by voting directly in referendums. At the opposite pole are political parties that believe that elected representatives should make the most important political decisions. (only asked in 2019). Values range on a scale from 0 to 10; 0 = Elected office holders should make the most important decisions, 10 = ‘The people’, not politicians, should make the most important decisions. + + + + Anti-elite salience: salience of anti-establishment and anti-elite rhetoric. (Asked in 2014 and 2019). Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Reducing political corruption salience: salience of reducing political corruption. (Asked in 2014 and 2019). Values range on a scale from 0 to 10; 0 = Not important at all, 10 = Extremely important. + + + + Members/Activists vs. Party leadership in policy choices: position on party leadership vs. members/activists making party policy choices. (only asked in 2019). Values range on a scale from 0 to 10; 0 = Members/activists have complete control over policy choices, 10 = Leadership has complete control over policy choices. + + + + + Turkey's EU Membership: Indicators of political parties' attitudes toward Turkey's EU membership. (only asked in 2019) + + + EU economic requirements: position on party leadership vs. members/activists making party policy choices (only asked in 2019). Values range on a scale from 1 to 7; 1 = Strongly opposed to fulfilling economic requirements, 7 = Strongly in favor of fulfilling economic requirements. + + + + EU political requirements: position on fulfilling the political requirements of EU membership. Values range on a scale from 1 to 7; 1 = Strongly opposed to fulfilling political requirements, 7 = Strongly in favor of fulfilling political requirements. + + + + EU good governance requirements: position on fulfilling the good governance requirements of EU membership. Values range on a scale from 1 to 7; 1 = Strongly opposed to fulfilling good governance requirements, 7 = Strongly in favor of fulfilling good governance requirements. + + + + + Most important issues: Entries for the next three questions are a summary of the expert responses to the Most Important Issue question. Each expert ranked one issue as the most important, one issue as the second most important, and one issue as the third most important issue. In this dataset, we aggregate these responses using a simple ordinal voting technique. For each party, an issue received 10 points if it is ranked as the #1 issue by an expert 5 points if it is ranked #2 by an expert, and 1 point if it is ranked #3 by an expert. After adding together the issue scores for all the experts for each individual party, we ranked each issue by the total number of points, yielding the MIP_ONE, MIP_TWO, and MIP_THREE variables. Tables with the Most important issue options are available in the 1999-2019 Chapel Hill Expert Survey (CHES) trend file Codebook (https://www.chesdata.eu/s/1999-2019_CHES_codebook.pdf) + + + Most important question: most important issue for the party over the course of specified year. + + + + Second most important question: second most important issue for the party over the course of specified year. + + + + Third most important question: third most important issue for the party over the course of specified year. + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-NER.ana.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-NER.ana.xml new file mode 100644 index 000000000..5504a79f1 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-NER.ana.xml @@ -0,0 +1,22 @@ + + + Named entities + Eigennamen + + person + Person + + + location + Ort + + + organization + Organisation + + + miscellaneous + Sonstiges + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-UD-SYN.ana.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-UD-SYN.ana.xml new file mode 100644 index 000000000..f34282bd9 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-UD-SYN.ana.xml @@ -0,0 +1,1687 @@ + + + + UD syntactic relations: All defined syntactic relations of the Universal Dependendencies project (GitHub commit) + + + acl: clausal modifier of noun (adnominal clause) + + + acl:adv: adverbs acting as amod + + + + acl:attr: attributive adnominal clause + + + + acl:cleft: clefted phrase modifier + + + + acl:cont: content clause as clausal modifier + + + + acl:crel: correlative clausal modifier + + + + acl:dpct: depictive clausal modifier + + + + acl:fixed: fixed clausal modifier + + + + acl:inf: acl:inf + + + + acl:pred: attributive clausal modifier + + + + acl:ptcp: attributive clausal modifier + + + + acl:relat: relational adnominal clause + + + + acl:relcl: relative clause modifier + + + + acl:subj: clausal modifier of noun (adnominal clause): the noun is the subject + + + + acl:tmod: clausal modifier of noun of time (adnominal clause) + + + + acl:tonp: nounization + + + + + advcl: adverbial clause modifier + + + advcl:abs: ablativus absolutus + + + + advcl:advers: adversal adverbial clause + + + + advcl:arelcl: adverbial clause headed by an adverbial relative clause + + + + advcl:cau: adverb with causative modality + + + + advcl:caus: causal adverbial clause + + + + advcl:ccomp: adverbial clausal modifier with elided speech verb + + + + advcl:cleft: clefted adverbial clause modifier + + + + advcl:cmp: comparative clause + + + + advcl:cmpr: comparative clause + + + + advcl:concess: concessive clause + + + + advcl:cond: conditional adverbial clause modifier + + + + advcl:consec: consecutive adverbial clausal modifier + + + + advcl:conv: adverbial clause headed by a converb + + + + advcl:coverb: adverbial coverb phrase + + + + advcl:dpct: clausal modifier of covert referent + + + + advcl:eval: advcl with evaluative modality + + + + advcl:fin: final clause + + + + advcl:lcl: advcl with locative to modality + + + + advcl:lto: advcl with locative to modality + + + + advcl:manner: adverbial clause of manner + + + + advcl:mcl: adverbial clause with modal modality + + + + advcl:objective: adverbial purpose clause modifier + + + + advcl:pred: adverbial secondary predication + + + + advcl:relcl: adverbial relative clause modifier + + + + advcl:svc: adverbial infinitive + + + + advcl:tcl: temporal adverbial clause + + + + + advmod: adverbial modifier + + + advmod:adj: adverbial modifier is an adjective + + + + advmod:arg: adverbial complement + + + + advmod:cau: adverb with causative modality + + + + advmod:cmp: comparative modifier of an adjective or adverb + + + + advmod:deg: advmod with degree modality + + + + advmod:det: adverbial modification by a determiner + + + + advmod:df: duration or frequency adverbial modifier + + + + advmod:dir: direction adverbial modifier + + + + advmod:emph: emphasizing word, intensifier + + + + advmod:eval: advmod with evaluative modality + + + + advmod:fixed: fixed adverbial modifier + + + + advmod:foc: advmod with focus modality + + + + advmod:freq: advmod with frequentative modality + + + + advmod:lfrom: advmod with source locative from modality + + + + advmod:lmod: locative adverbial modifier + + + + advmod:lmp: advmod with locative multipoint modality + + + + advmod:loc: an adverbial local marker + + + + advmod:locy: adverbial modifier where + + + + advmod:lto: advmod with locative to modality + + + + advmod:mmod: advmod with modal modality + + + + advmod:mode: adverbial modifier + + + + advmod:neg: negation word + + + + advmod:obl: contracted advmod and oblique nominal + + + + advmod:que: question suffix + + + + advmod:tfrom: adverbial modifier since when + + + + advmod:tlocy: adverbial modifier when + + + + advmod:tmod: advmod with temporal modality + + + + advmod:to: adverbial modifier where to + + + + advmod:tto: adverbial modifier till when + + + + advmod:verb: adverbial modifier is a verb + + + + + amod: adjectival modifier + + + amod:att: adjectival modifier + + + + amod:attlvc: participle in a light-verb construction + + + + amod:flat: adjectival part of a named entity + + + + + appos: appositional modifier + + + appos:nmod: An appositional modifier + + + + appos:trans: nominal translation pair + + + + + aux: auxiliary + + + aux:aff: auxiliary affix + + + + aux:aspect: aspect auxiliary + + + + aux:caus: causative auxiliary + + + + aux:clitic: mobile inflection auxiliary + + + + aux:cnd: conditional auxiliary + + + + aux:ex: existentials as auxiliary + + + + aux:exhort: exhortative auxiliary + + + + aux:imp: imperative marker auxiliary + + + + aux:nec: necessative auxiliary + + + + aux:neg: negative auxiliary + + + + aux:opt: optative auxiliary + + + + aux:part: auxiliary particle + + + + aux:pass: passive auxiliary + + + + aux:pot: potential auxiliary + + + + aux:q: question auxiliary + + + + aux:tense: tense auxiliary + + + + + case: case marking + + + case:acc: accusative case marker + + + + case:adv: case marking to form adverbs + + + + case:det: preposition with determiner + + + + case:gen: genitive case marker + + + + case:loc: postpositional localizer + + + + case:pred: predicative particles + + + + case:sim: standard marker in comparision + + + + case:voc: vocative particle + + + + + cc: coordinating conjunction + + + cc:nc: coordinated conjunct : non coordonant + + + + cc:preconj: preconjunct + + + + + ccomp: clausal complement + + + ccomp:cleft: required clausal dependent of the pronoun _to_ + + + + ccomp:obj: clausal complement (object) + + + + ccomp:obl: clausal complement (non-object) + + + + ccomp:pmod: clausal prepositional object + + + + ccomp:pred: clausal complement (predicative) + + + + ccomp:purp: a clausal complement that is a purposive + + + + ccomp:quote: a clausal complement consisting of direct speech + + + + ccomp:rel: relative clause as core argument + + + + ccomp:relcl: double pronoun construction or free relative acting as object + + + + ccomp:reported: reported speech from active verb of saying or digression in discursive form + + + + ccomp:speech: direct quotation + + + + + clf: classifier + + + clf:det: classifier used as determiner + + + + + compound: compound + + + compound:a: adjective compound + + + + compound:adj: adjective and adjective compound + + + + compound:affix: construct state modification + + + + compound:amod: Noun and adjective compound + + + + compound:apr: adjective and particle compound + + + + compound:atov: adjective and verb compound + + + + compound:coord: non-final member of a coordinate compound + + + + compound:dir: directional verb compound + + + + compound:ext: extent and descriptive verb compound + + + + compound:lvc: light verb construction + + + + compound:name: left-of-head member of multi-word name + + + + compound:nn: noun compound modifier + + + + compound:preverb: relation between verb and preverb + + + + compound:pron: Noun and pronoun compound + + + + compound:prt: phrasal verb particle + + + + compound:quant: verb-quantifier compound + + + + compound:redup: reduplicated compounds + + + + compound:smixut: construct state modification + + + + compound:svc: serial verb compounds + + + + compound:verbnoun: verb and noun compound + + + + compound:vmod: noun-verb compound + + + + compound:vo: verb-object compound + + + + compound:vv: verb-verb compound + + + + compound:z: compound with Z (POS tag) + + + + + conj: conjunct + + + conj:expl: explicative conjunct + + + + conj:extend: open/incomplete conjunct + + + + conj:redup: reduplicated conjunction + + + + conj:svc: coordination of serial verbs + + + + + cop: copula + + + cop:emph: emphasizing copula + + + + cop:expl: copula with expletive subject + + + + cop:locat: copula with a locative predicate + + + + cop:outer: outer copula + + + + cop:own: copula for posessive clauses + + + + + csubj: clausal subject + + + csubj:asubj: clausal subject: Adjective subject + + + + csubj:caus: clausal causative subject + + + + csubj:cleft: relative clause modifier + + + + csubj:cop: clausal copular subject + + + + csubj:outer: outer clause clausal subject + + + + csubj:pass: clausal passive subject + + + + csubj:pred: clausal subject of a secondary predicate (enhanced dependency) + + + + csubj:relcl: double pronoun construction or free relative acting as subject + + + + csubj:reported: reported speech from passive verb of saying + + + + csubj:vsubj: clausal subject: Adjective subject + + + + + dep: unspecified dependency + + + dep:aff: affix unspecified dependency + + + + dep:agr: agreement clitic + + + + dep:alt: alternative gender suffix + + + + dep:ana: anaphoric prefix dependency + + + + dep:aux: derivational auxiliary dependency + + + + dep:comp: unspecified dependency + + + + dep:conj: conjunct + + + + dep:cop: derivational copular dependency + + + + dep:der: derivational suffix + + + + dep:emo: emotional root dependency + + + + dep:infl: inflectional dependency + + + + dep:mark: derivational marker dependency + + + + dep:mod: modifier underspecified for the syntactic category of its head + + + + dep:pos: postural root dependency + + + + dep:repeat: repeating of word + + + + dep:ss: status suffix + + + + + det: determiner + + + det:adj: article of a prearticulated adjective + + + + det:clf: determiner classifier + + + + det:noun: article of a prearticulated noun + + + + det:numgov: pronominal quantifier governing the case of the noun + + + + det:nummod: pronominal quantifier agreeing in case with the noun + + + + det:pmod: pronoun determiner + + + + det:poss: possessive determiner + + + + det:predet: predeterminer + + + + det:pron: article of a prearticulated pronoun + + + + det:rel: relative determiner + + + + + discourse: discourse element + + + discourse:conn: discourse connective marker + + + + discourse:emo: emoticons, emojis + + + + discourse:filler: filler sound in spoken data + + + + discourse:intj: interjection + + + + discourse:q: discourse particle for questions + + + + discourse:sp: sentence particle + + + + + dislocated: dislocated elements + + + dislocated:advcl: dislocated adverbial clause + + + + dislocated:ccomp: dislocated complement clause + + + + dislocated:cleft: cleft constructions that lack a copula + + + + dislocated:csubj: dislocated clausal subject + + + + dislocated:nsubj: dislocated nominal subject + + + + dislocated:obj: dislocated object + + + + dislocated:obl: dislocated oblique argument + + + + dislocated:subj: dislocated subject + + + + dislocated:vo: dislocated object of verb-object compound + + + + + expl: expletive + + + expl:comp: expletive + + + + expl:impers: impersonal expletive + + + + expl:pass: reflexive pronoun used in reflexive passive + + + + expl:poss: possessively used reflexive clitic + + + + expl:pv: reflexive clitic with an inherently reflexive verb + + + + expl:subj: expletive subject + + + + + fixed: fixed multiword expression + + + + flat: flat expression + + + flat:abs: clausal absolutive + + + + flat:date: flat multiword expression: date + + + + flat:dist: distributive + + + + flat:foreign: foreign words + + + + flat:gov: partitive-like appositional element + + + + flat:name: names + + + + flat:num: compound number + + + + flat:number: flat multiword expression: number + + + + flat:range: range + + + + flat:redup: reduplication + + + + flat:repeat: repetition + + + + flat:sibl: siblings + + + + flat:time: flat multiword expression: time + + + + flat:title: parts of a title + + + + flat:vv: serial verbs + + + + + goeswith: goes with + + + + iobj: indirect object + + + iobj:agent: agentive indirect object + + + + iobj:appl: applied indirect object in applicative construction + + + + iobj:patient: patient object of a non-actor/patient-focused verb + + + + + list: list + + + + mark: marker + + + mark:adv: manner adverbializer + + + + mark:advmod: adverbial modifier confusable with a subordination marker + + + + mark:aff: affix marker + + + + mark:int: interrogative + + + + mark:pcomp: marker for the purpose clause + + + + mark:plur: independently written plural suffix + + + + mark:prt: particle + + + + mark:q: question particle + + + + mark:rel: adjectival, relativizer, and nominalizer 的 DE + + + + mark:sim: clausal comparison marker + + + + + neg: negation modifier + + + + nmod: nominal modifier + + + nmod:agent: agent of verbnouns in _cael_ constructions + + + + nmod:appos: nominal modifier apposition + + + + nmod:arg: nominal modifier used as an argument + + + + nmod:att: nominal modifier without postposition + + + + nmod:attlvc: object nominal in a nominalized light-verb construction + + + + nmod:attr: attributive nominal modifier + + + + nmod:bahuv: nominal bahuvriihi modifier + + + + nmod:cau: nominal modifier indicating the causee of a causative predicate + + + + nmod:comp: comparative modifier of an adjective or adverb + + + + nmod:desc: descriptor modifier in nominal + + + + nmod:det: determinative + + + + nmod:flat: nominal part of a named entity + + + + nmod:gen: genitive modifier + + + + nmod:gobj: genitive object + + + + nmod:gsubj: genitive subject + + + + nmod:lfrom: advmod with locative to modality + + + + nmod:lmod: noun modifier that describes the position/location relative to its parent noun + + + + nmod:npmod: noun phrase as adverbial modifier + + + + nmod:obj: nominative object + + + + nmod:obl: nominal modifier in oblique case + + + + nmod:part: nominal modifier indicating part-whole relations + + + + nmod:poss: possessive nominal modifier + + + + nmod:pre: preposed genitive + + + + nmod:pred: nominal first member in *bahuvrīhi* compound + + + + nmod:prep: prepositional pronouns + + + + nmod:prp: proprietive modifier of a noun + + + + nmod:relat: relational nominal modifier + + + + nmod:subj: nominative subject + + + + nmod:tmod: temporal modifier + + + + + nsubj: nominal subject + + + nsubj:advmod: fused subject pronoun and adverb + + + + nsubj:aff: nominal subject affix + + + + nsubj:bfoc: nominal subject of a beneficiary-focused verb + + + + nsubj:caus: causative nominal subject + + + + nsubj:cleft: nominal residual subject of a cleft sentence + + + + nsubj:contrast: nsubj:contrast + + + + nsubj:cop: nominal copular subject + + + + nsubj:expl: expletive subject + + + + nsubj:ifoc: nominal subject of an instrumental-focused verb + + + + nsubj:lfoc: nominal subject of a locative-focused verb + + + + nsubj:lvc: subject in a light-verb construction + + + + nsubj:nc: non-canonical nominal subject + + + + nsubj:nn: nominal subject: the predicate is noun + + + + nsubj:obj: fused subject and object pronoun + + + + nsubj:outer: outer clause nominal subject + + + + nsubj:pass: passive nominal subject + + + + nsubj:periph: postposed subject + + + + nsubj:pred: subject of a secondary predicate (enhanced dependency) + + + + nsubj:quasi: quasi subject + + + + nsubj:situation: nsubj:situation + + + + nsubj:x: controlling subject (enhanced dependency) + + + + nsubj:xsubj: The subject xcomp is the complement + + + + + nummod: numeric modifier + + + nummod:det: numeric is determiner + + + + nummod:entity: numeric modifier governed by a noun + + + + nummod:flat: numeral part of a named entity + + + + nummod:gov: numeric modifier governing the case of the noun + + + + + obj: object + + + obj:advmod: fused adverb and object pronoun + + + + obj:advneg: fused negation and object pronoun + + + + obj:agent: agentive object + + + + obj:appl: applied object in applicative construction + + + + obj:cau: direct object of an intransitive causative verb + + + + obj:caus: agentive object in causative construction + + + + obj:infx: infixed pronouns + + + + obj:lo: local object + + + + obj:lvc: object in a light-verb construction + + + + obj:obl: fused oblique and object pronoun + + + + obj:periph: preposed object + + + + + obl: oblique nominal + + + obl:ad: oblique adjunct + + + + obl:adj: oblique nominal: Auxiliary nouns for adjectives + + + + obl:adv: oblique nominal for adverbs + + + + obl:advmod: adverbial modifier confusable with an oblique dependent + + + + obl:agent: agent modifier + + + + obl:appl: applied oblique argument in non-canonical applicative construction + + + + obl:arg: oblique argument + + + + obl:benef: benefactive/malefactive oblique argument + + + + obl:cau: oblique with causative modality + + + + obl:cmp: standard-of-comparison modifier of an adjective or adverb + + + + obl:cmpr: comparative oblique argument + + + + obl:comp: oblique nominal with other preposition + + + + obl:dat: dative argument + + + + obl:depict: depictive oblique modifier + + + + obl:float: oblique as a floating modifier + + + + obl:freq: oblique with frequentative modality + + + + obl:goal: oblique argument expressing the goal + + + + obl:grad: oblique argument denoting the standard in a comparison + + + + obl:inst: oblique instrument + + + + obl:instr: oblique argument expressing the instrument + + + + obl:iobj: oblique nominal for verb means "to give" + + + + obl:lfrom: obl with source locative from modality + + + + obl:lmod: locative modifier + + + + obl:lmp: obl with locative to modality + + + + obl:lto: obl with locative to modality + + + + obl:lvc: oblique nominal in a light-verb construction + + + + obl:manner: oblique argument expressing the manner + + + + obl:mcl: obl indicating manner + + + + obl:mod: oblique modifier + + + + obl:npmod: noun phrase as adverbial modifier + + + + obl:obj: NP-related object + + + + obl:orphan: adpositional dependent with the elided noun + + + + obl:own: owner in possessive existential clauses + + + + obl:path: oblique argument expressing the path + + + + obl:patient: object in ZOENG construction + + + + obl:pmod: prepositional object + + + + obl:poss: oblique modifier specifying possessor + + + + obl:prep: prepositional pronouns + + + + obl:pronmod: pronominal modifier + + + + obl:sentcon: sentence-initial discourse connective + + + + obl:smod: oblique spatial modifier + + + + obl:soc: oblique argument expressing the accompaniment + + + + obl:source: oblique argument expressing the source + + + + obl:subj: NP-related subject + + + + obl:tmod: temporal modifier + + + + obl:verb: oblique verb + + + + obl:with: oblique nominal to answer "with whom" + + + + + orphan: orphan + + + orphan:missing: textual gap in the source + + + + + parataxis: parataxis + + + parataxis:appos: paratactic apposition + + + + parataxis:conj: coordinating parataxis + + + + parataxis:deletion: loosely connected clause because of deletion + + + + parataxis:discourse: paratactic discourse + + + + parataxis:dislocated: parataxis:dislocated + + + + parataxis:hashtag: paratactic hashtag + + + + parataxis:insert: parataxis insert + + + + parataxis:mod: parataxis:mod + + + + parataxis:newsent: new sentence attached to a node in the previous sentence + + + + parataxis:nsubj: paratactic nominal subject + + + + parataxis:obj: direct speech + + + + parataxis:parenth: parataxis parenthesical + + + + parataxis:rel: relative clause for clauses + + + + parataxis:rep: reported speech + + + + parataxis:reporting: interjected verb of saying + + + + parataxis:restart: loosely connected clause because of restart + + + + parataxis:rt: retweets + + + + parataxis:sentence: sentence + + + + parataxis:shared: parataxis with shared arguments + + + + parataxis:speaker: backgrounded actor initiating a direct discourse + + + + parataxis:speech: reported speech + + + + parataxis:trans: translation pair + + + + parataxis:url: URLs + + + + + punct: punctuation + + + + reparandum: overridden disfluency + + + + root: root + + + + vocative: vocative + + + vocative:cl: clausal vocative + + + + vocative:mention: Twitter mentions + + + + + xcomp: open clausal complement + + + xcomp:adj: open clausal complement for adjective + + + + xcomp:advmod: a verb play role as an advmod for another verb + + + + xcomp:cleft: required open dependent of the pronoun _to_ + + + + xcomp:dir: open clausal complement for adjective + + + + xcomp:ds: clausal complement with different subject + + + + xcomp:obj: infinitival objects + + + + xcomp:pred: predicate + + + + xcomp:relcl: double pronoun construction or free relative acting as complement clause + + + + xcomp:result: resultative + + + + xcomp:subj: infinitival and adverbial subjects + + + + xcomp:vcomp: open clausal complement for adjective + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-parla.legislature.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-parla.legislature.xml new file mode 100644 index 000000000..434529cdf --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-parla.legislature.xml @@ -0,0 +1,138 @@ + + + Legislature + Legislative + + Geo-political or administrative units + Geo-politische oder administrative Einheiten + + Supranational legislature + Supranational Gesetzgebung + + + National legislature + Nationale Gesetzgebungsebene + + + Regional legislature + Regionale Gesetzgebung + + + Local legislature + Lokale Gesetzgebung + + + + Organization + Organisation + + Chambers + Kammern + + Unicameralism + Einkammernsystem + + + Bicameralism + Zweikammernsystem + + Upper house + Bundesrat + + + Lower house + Nationalrat + + + + Multicameralism + Mehrkammernsystem + + Chamber + Kammern + + + + + Committee + Ausschuss + + Standing committee + Ständiger Ausschuss + + + Special committee + Spezialausschuss + + + Committee of inquiry + Untersuchungsausschuss + + + + + Legislative period: term of the parliament between general elections. + Gesetzgebungsperiode + + Legislative session: the period of time in which a legislature is convened for purpose of lawmaking, usually being one of two or more smaller divisions of the entire time between two elections. A session is a meeting or series of connected meetings devoted to a single order of business, program, agenda, or announced purpose. + Sitzungsperiode + + Meeting: Each meeting may be a separate session or part of a group of meetings constituting a session. The session/meeting may take one or more days. + Sitzung + + Types of meetings + Arten von Sitzungen + + Regular meeting + Reguläre Sitzung + + + Special meeting + Sondersitzung + + Extraordinary meeting + Sondersitzung + + + Urgent meeting + dirgende Sitzung + + + Ceremonial meeting + Festsitzung + + + Commemorative meeting + Gedenksitzung + + + Public presentation of opinions + Öffentliche Meinungsäußerung + + + Committee meeting + Ausschussitzung + + + + Continued meeting + Fortsetzende Sitzung + + + Public meeting + Öffentliche Sitzung + + + Executive meeting + Vorstandssitzung + + + + Sitting: sitting day + Sitzungstag + + + + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-politicalOrientation.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-politicalOrientation.xml new file mode 100644 index 000000000..e308bbc89 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-politicalOrientation.xml @@ -0,0 +1,78 @@ + + + Political orientation of political parties and parliamentary groups + Politische Orientierung der Parteien und Parlamentsklubs + + Left + Rechts + + + Centre + Mitte + + + Right + Rechts + + + Far-left + Extrem Links + + + Far-right + Extrem Rechts + + + Centre-left + Mitte-links + + + Centre-right + Mitte-rechts + + + Centre to centre-left + Mitte bis Mitte-links + + + Centre to centre-right + Mitte bis Mitte-rechts + + + Centre-left to left + Mitte-links bis Links + + + Centre-right to right + Mitte-rechts bis Rechts + + + Left to far-left + Links bis extrem Links + + + Right to far-right + Rechts bis extrem Rechts + + + Big tent: Big tent or catch-all refers to political parties that have members covering a broad spectrum of beliefs. + "Catch-all" Partei + + + Nonpartisanism: Nonpartisanism refers to a political stance that does not agree with the current political party system. + Überparteilichkeit + + + Pirate Party: Pirate Party refers to political parties that support civil rights, direct democracy, encourage innovation and creativity, free sharing of knowledge, information privacy, free speech, anti-corruption, net neutrality and oppose mass surveillance, censorship and Big Tech. + Piratenpartei + + + Single Issue Politics: Single Issue Politics refers to a political stance that is based on one essential policy area or idea. + Interessenpartei + + + Syncretic politics: Syncretic politics refers to politics that combine elements from across the conventional left–right political spectrum. + Synkretische Politik + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-sentiment.ana.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-sentiment.ana.xml new file mode 100644 index 000000000..63ae8b989 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-sentiment.ana.xml @@ -0,0 +1,42 @@ + + + Sentiment: 3 and 6 class sentiment labels following the Multilingual sentiment dataset of parliamentary debates ParlaSent 1.0; the values given in the category descriptions are as returned by the ParlaSent classifier trained on ParlaSent 1.0. + Stimmung: 3 und 6 teilige Stimmungsannotationen basierende auf dem mehrsprachigen Stimmungsdatenset von Parlemantesdebatten ParlaSent 1.0 folgen; die Werte in den Kategoriebeschreibungen sind die deutsche Übersetzung der Werte, die das Tool ParlaSent classifier, trainiert auf den ParlaSent 1.0 Daten ausgibt + + Negative: value < 1.5 + Negativ: value < 1.5 + + negative: value < 0.5 + negativ: value < 1.5 + + + mixed negative: interval [0.5, 1.5) + gemischt negativ: interval [0.5, 1.5) + + + + Neutral: interval [1.5, 3.5) + Neutral: interval [1.5, 3.5) + + neutral negative: interval [1.5, 2.5) + neutral negativ: interval [1.5, 2.5) + + + neutral positive: interval [2.5, 3.5) + neutral positiv: interval [2.5, 3.5) + + + + Positive: value >= 3.5 + Positiv: value >= 3.5 + + mixed positive: interval [3.5, 4.5) + gemischt positiv: interval [3.5, 4.5) + + + positive: value >= 4.5 + positiv: value >= 4.5 + + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-speaker_types.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-speaker_types.xml new file mode 100644 index 000000000..b8503cd35 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-speaker_types.xml @@ -0,0 +1,18 @@ + + + Types of speakers + Sprechertyp + + Chairperson: chairman of a sitting + PräsidentIn: PräsidentIn der Sitzung + + + Regular: a regular speaker at a sitting + Reguläre/r SprecherIn: Mitglied des Nationalrats oder der Regierung + + + Guest: a guest speaker at a sitting + Gast: GastsprecherIn + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-subcorpus.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-subcorpus.xml new file mode 100644 index 000000000..79fccb478 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-subcorpus.xml @@ -0,0 +1,18 @@ + + + Subcorpora + Subkorpora + + Reference: reference subcorpus, until 2020-01-30 + Reference + + + COVID: COVID subcorpus, from 2020-01-31 onwards, when WHO made the formal declaration of PHEIC, i.e. the Public Health Emergency of International Concern for COVID-19 + COVID: COVID Subkorpus, ab 2020-01-31 als WHO COVID-19 zu GNIT, d.h. zur gesundheitliche Notlage internationaler Tragweite erklärte + + + War: War in Ukraine subcorpus, from 2022-02-24 onwards, i.e. from Russia's full-scale invasion of Ukraine + Krieg: Krieg in der Ukranine Subkorpus, ab 2022-02-24, d.h. ab Russlands groß angelegter Invasion der Ukraine + + + diff --git a/Samples/ParlaMint-DE/ParlaMint-taxonomy-topic.xml b/Samples/ParlaMint-DE/ParlaMint-taxonomy-topic.xml new file mode 100644 index 000000000..c9a6e0248 --- /dev/null +++ b/Samples/ParlaMint-DE/ParlaMint-taxonomy-topic.xml @@ -0,0 +1,100 @@ + + + Topics: Comparative Agendas Project CAP major topic labels + Themen: Deutsche Übersetzungen der im Comparative Agendas Projekt (CAP) entstandenen Annotations-Tags für Themen + + Agriculture + Landwirtschaft + + + Civil Rights + Bürgerrechte + + + Culture + Kultur + + + Defense + Landesverteidigung + + + Domestic Commerce + Binnenhandel + + + Education + Bildung + + + Energy + Energuie + + + Environment + Umwelt + + + Foreign Trade + Handel + + + Government Operations + Regierungsgeschäfte + + + Health + Gesundheit + + + Housing + + + + + + Immigration + Einwanderung + + + International Affairs + internationale Angelegenheiten + + + Labor + Arbeit + + + Law and Crime + Recht und Kriminalität + + + Macroeconomics + Makroöknomie + + + Mix + Gemischt + + + Other + Anderes + + + Public Lands + Öffentlicher Grund und Boden + + + Social Welfare + Sozialwesen + + + Technology + Technologie + + + Transportation + Transportwesen + + + diff --git a/Samples/ParlaMint-DE/README.md b/Samples/ParlaMint-DE/README.md new file mode 100644 index 000000000..52f7e9895 --- /dev/null +++ b/Samples/ParlaMint-DE/README.md @@ -0,0 +1,2 @@ +# ParlaMint directory for samples of country DE (Germany) +## Languages: de (German) diff --git a/Scripts/check-links.xsl b/Scripts/check-links.xsl index 7a3a9c37c..4c0edabff 100644 --- a/Scripts/check-links.xsl +++ b/Scripts/check-links.xsl @@ -60,7 +60,7 @@ @inst | @lemmaRef | @location | @mergedIn | @mutual | @new | @next | @nymRef | @origin | @parent | @parts | @passive | @perf | @period | @prev | @ref | @rendition | @require | @resp | @sameAs | @scheme | @scriptRef | @select | - @since | @source | @spanTo | @start | @synch | @target | @targetEnd | + @since | @source | @spanTo | @synch | @target | @targetEnd | @toUnit | @toWhom | @unitRef | @uri | @url | @valueDatcat | @where | @who | @wit"> diff --git a/Scripts/parlamint-lib.xsl b/Scripts/parlamint-lib.xsl index 8f8bc2c91..bf4aad68d 100755 --- a/Scripts/parlamint-lib.xsl +++ b/Scripts/parlamint-lib.xsl @@ -47,9 +47,9 @@ - + - + diff --git a/TEI/Makefile b/TEI/Makefile index 93653668e..23f94f529 100644 --- a/TEI/Makefile +++ b/TEI/Makefile @@ -1,8 +1,8 @@ check: prepare val-schema html mount # Validate and compile ODD to RelaxNG, validate samples, make HTML -all: prepare val-schema xml-schemas val-samples html -xall: prepare val-schema xml-schemas val-samples html +all: prepare val-odd val-schema xml-schemas val-samples html +xall: prepare val-odd val-schema xml-schemas val-samples html # Builds html and pushes it deploy: html @@ -36,7 +36,9 @@ xml-schemas: val-schema #Stylesheets/bin/teitodtd ${SUBSET} ${ODD} ParlaMint.odd.dtd #Stylesheets/bin/teitoxsd ${SUBSET} ${ODD} ParlaMint.odd.xsd -val: val-schema val-samples +val: val-odd val-schema val-samples +val-odd: + $s -xsl:../Scripts/check-links.xsl tmp/ParlaMint.xml val-schema: $j tei_odds.rng tmp/ParlaMint.xml diff --git a/TEI/ParlaMint-schemaSpecs.editing.odd.xml b/TEI/ParlaMint-schemaSpecs.editing.odd.xml index 504f1c0b0..06ca973f3 100644 --- a/TEI/ParlaMint-schemaSpecs.editing.odd.xml +++ b/TEI/ParlaMint-schemaSpecs.editing.odd.xml @@ -2750,9 +2750,9 @@ ]]> + href="ParlaMint-SI-taxonomy-parla.legislature.xml"/>]]> ]]> + href="ParlaMint-SI-taxonomy.xml-speaker_types"/>]]> ... @@ -3142,9 +3142,9 @@ ]]> + href="ParlaMint-SI-listOrg.xml"/>]]> ]]> + href="ParlaMint-SI-listPerson.xml"/>]]> @@ -3287,9 +3287,9 @@ ]]> + href="ParlaMint-SI-listOrg.xml"/>]]> ]]> + href="ParlaMint-SI-listPerson.xml"/>]]> diff --git a/TEI/ParlaMint-schemaSpecs.odd.xml b/TEI/ParlaMint-schemaSpecs.odd.xml index ddb3aa2b7..27d4ce3a5 100644 --- a/TEI/ParlaMint-schemaSpecs.odd.xml +++ b/TEI/ParlaMint-schemaSpecs.odd.xml @@ -4606,9 +4606,9 @@ <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-taxonomy-parla.legislature.xml"/> + href="ParlaMint-SI-taxonomy-parla.legislature.xml"/> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-taxonomy.xml-speaker_types"/> + href="ParlaMint-SI-taxonomy.xml-speaker_types"/> ... @@ -5207,9 +5207,9 @@ <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-listOrg.xml"/> + href="ParlaMint-SI-listOrg.xml"/> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-listPerson.xml"/> + href="ParlaMint-SI-listPerson.xml"/> @@ -5466,9 +5466,9 @@ <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-listOrg.xml"/> + href="ParlaMint-SI-listOrg.xml"/> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-listPerson.xml"/> + href="ParlaMint-SI-listPerson.xml"/> diff --git a/TEI/ParlaMint.odd.rnc b/TEI/ParlaMint.odd.rnc index 232a12c8d..a384f4654 100644 --- a/TEI/ParlaMint.odd.rnc +++ b/TEI/ParlaMint.odd.rnc @@ -11,8 +11,8 @@ namespace xi = "http://www.w3.org/2001/XInclude" namespace xlink = "http://www.w3.org/1999/xlink" namespace xsl = "http://www.w3.org/1999/XSL/Transform" -# Schema generated from ODD source 2025-06-13T15:40:08Z. 2025-06-13. -# TEI Edition: P5 Version 4.10.0a. Last updated on 24th January 2025, revision 4e6b08ef5 +# Schema generated from ODD source 2025-09-30T13:46:56Z. 2025-09-30. +# TEI Edition: P5 Version 4.10.0a. Last updated on 10th June 2025, revision 78f0e8312 # TEI Edition Location: https://www.tei-c.org/Vault/P5/4.10.0a/ # @@ -236,6 +236,41 @@ tei_att.declarable.attribute.default = ## This element can only be selected explicitly, unless it is the only one of its kind, in which case it is selected if its parent is selected. "false" }? +sch:pattern [ + id = "declarable" + abstract = "true" + "\x{a}" ~ + " " + sch:rule [ + context = + "$tde[ ancestor::tei:teiHeader and following-sibling::$tde and not( preceding-sibling::$tde ) ]" + "\x{a}" ~ + " " + sch:report [ + test = "../child::$tde[ not( @xml:id ) ]" + "\x{a}" ~ + " When there is more than one " + sch:name [ ] + ", each must have an @xml:id\x{a}" ~ + " " + ] + "\x{a}" ~ + " " + sch:assert [ + test = + "count( ../child::$tde[ normalize-space( @default ) = ('1','true') ] ) eq 1" + "\x{a}" ~ + " When there is more than one " + sch:name [ ] + ", one and only one must have a @default of 'true'.\x{a}" ~ + " " + ] + "\x{a}" ~ + " " + ] + "\x{a}" ~ + " " +] tei_att.duration.w3c.attributes = tei_att.duration.w3c.attribute.dur tei_att.duration.w3c.attribute.dur = @@ -333,7 +368,7 @@ tei_att.ranging.attribute.confidence = }? sch:pattern [ id = - "parlamint-att.spanning-spanTo-spanTo-points-to-following-constraint-rule-5" + "parlamint-att.spanning-spanTo-spanTo-points-to-following-constraint-rule-6" "\x{a}" ~ " " sch:rule [ @@ -677,7 +712,7 @@ tei_model.pPart.data_sequenceRepeatable = tei_model.nameLike_sequenceRepeatable+ sch:pattern [ id = - "parlamint-att.calendarSystem-calendar-calendar_attr_on_empty_element-constraint-rule-6" + "parlamint-att.calendarSystem-calendar-calendar_attr_on_empty_element-constraint-rule-7" "\x{a}" ~ " " sch:rule [ @@ -712,7 +747,7 @@ tei_p = ((tei_ref | text)+) >> sch:pattern [ id = - "parlamint-p-abstractModel-structure-p-in-ab-or-p-constraint-rule-7" + "parlamint-p-abstractModel-structure-p-in-ab-or-p-constraint-rule-8" "\x{a}" ~ " " sch:rule [ @@ -734,7 +769,7 @@ tei_p = ] >> sch:pattern [ id = - "parlamint-p-abstractModel-structure-p-in-l-constraint-rule-8" + "parlamint-p-abstractModel-structure-p-in-l-constraint-rule-9" "\x{a}" ~ " " sch:rule [ @@ -764,7 +799,7 @@ tei_desc = (tei_term?, (text | tei_ref)+) >> sch:pattern [ id = - "parlamint-desc-deprecationInfo-only-in-deprecated-constraint-rule-9" + "parlamint-desc-deprecationInfo-only-in-deprecated-constraint-rule-10" "\x{a}" ~ " " sch:rule [ @@ -1008,7 +1043,7 @@ tei_ref = element ref { text >> sch:pattern [ - id = "parlamint-ref-refAtts-constraint-rule-10" + id = "parlamint-ref-refAtts-constraint-rule-11" "\x{a}" ~ " " sch:rule [ @@ -1161,7 +1196,16 @@ tei_bibl = ## (bibliographic citation) contains a loosely-structured bibliographic citation of which the sub-components may or may not be explicitly tagged. [3.12.1. Methods of Encoding Bibliographic References and Lists of References 2.2.7. The Source Description 16.3.2. Declarable Elements] element bibl { - tei_title+, (tei_edition? | tei_publisher? | tei_idno* | tei_date)+ + tei_title+, + ((tei_edition? | tei_publisher? | tei_idno* | tei_date)+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:bibl" ] + "\x{a}" ~ + " " + ] } tei_teiCorpus = [ @@ -1308,7 +1352,15 @@ tei_availability = ## (availability) supplies information about the availability of a text, for example any restrictions on its use or distribution, its copyright status, any licence applying to it, etc. [2.2.4. Publication, Distribution, Licensing, etc.] element availability { tei_licence, - tei_p+, + (tei_p+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:availability" ] + "\x{a}" ~ + " " + ], ## attribute status { @@ -1325,7 +1377,18 @@ tei_licence = tei_sourceDesc = ## (source description) describes the source(s) from which an electronic text was derived or generated, typically a bibliographic description in the case of a digitized text, or a phrase such as "born digital" for a text which has no previous existence. [2.2.7. The Source Description] - element sourceDesc { tei_bibl+, tei_recordingStmt? } + element sourceDesc { + tei_bibl+, + (tei_recordingStmt?) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:sourceDesc" ] + "\x{a}" ~ + " " + ] + } tei_encodingDesc = ## (encoding description) documents the relationship between an electronic text and the source or sources from which it was derived. [2.3. The Encoding Description 2.1.1. The TEI Header and Its Components] @@ -1340,32 +1403,78 @@ tei_encodingDesc = tei_projectDesc = ## (project description) describes in detail the aim or purpose for which an electronic file was encoded, together with any other relevant information concerning the process by which it was assembled or collected. [2.3.1. The Project Description 2.3. The Encoding Description 16.3.2. Declarable Elements] - element projectDesc { tei_p+ } + element projectDesc { + (tei_p+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:projectDesc" ] + "\x{a}" ~ + " " + ] + } tei_editorialDecl = ## (editorial practice declaration) provides details of editorial principles and practices applied during the encoding of a text. [2.3.3. The Editorial Practices Declaration 2.3. The Encoding Description 16.3.2. Declarable Elements] element editorialDecl { - (tei_correction - | tei_normalization - | tei_hyphenation - | tei_quotation - | tei_segmentation)+ + ((tei_correction + | tei_normalization + | tei_hyphenation + | tei_quotation + | tei_segmentation)+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:editorialDecl" ] + "\x{a}" ~ + " " + ] } tei_correction = ## (correction principles) states how and under what circumstances corrections have been made in the text. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements] - element correction { tei_p+ } + element correction { + (tei_p+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:correction" ] + "\x{a}" ~ + " " + ] + } tei_normalization = ## (normalization) indicates the extent of normalization or regularization of the original source carried out in converting it to electronic form. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements] - element normalization { tei_p+ } + element normalization { + (tei_p+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:normalization" ] + "\x{a}" ~ + " " + ] + } tei_quotation = ## (quotation) specifies editorial practice adopted with respect to quotation marks in the original. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements] element quotation { (tei_p+) >> sch:pattern [ - id = "parlamint-quotation-quotationContents-constraint-rule-11" + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:quotation" ] + "\x{a}" ~ + " " + ] + >> sch:pattern [ + id = "parlamint-quotation-quotationContents-constraint-rule-12" "\x{a}" ~ " " sch:rule [ @@ -1390,11 +1499,31 @@ tei_quotation = tei_hyphenation = ## (hyphenation) summarizes the way in which hyphenation in a source text has been treated in an encoded version of it. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements] - element hyphenation { tei_p+ } + element hyphenation { + (tei_p+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:hyphenation" ] + "\x{a}" ~ + " " + ] + } tei_segmentation = ## (segmentation) describes the principles according to which the text has been segmented, for example into sentences, tone-units, graphemic strata, etc. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements] - element segmentation { tei_p+ } + element segmentation { + (tei_p+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:segmentation" ] + "\x{a}" ~ + " " + ] + } tei_tagsDecl = ## (tagging declaration) provides detailed information about the tagging applied to a document. [2.3.4. The Tagging Declaration 2.3. The Encoding Description] @@ -1546,7 +1675,17 @@ tei_profileDesc = tei_langUsage = ## (language usage) describes the languages, sublanguages, registers, dialects, etc. represented within a text. [2.4.2. Language Usage 2.4. The Profile Description 16.3.2. Declarable Elements] - element langUsage { tei_language+ } + element langUsage { + (tei_language+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:langUsage" ] + "\x{a}" ~ + " " + ] + } tei_language = ## (language) characterizes a single language or sublanguage used within a text. [2.4.2. Language Usage] @@ -1576,7 +1715,17 @@ tei_language = tei_textClass = ## (text classification) groups information which describes the nature or topic of a text in terms of a standard classification scheme, thesaurus, etc. [2.4.3. The Text Classification] - element textClass { tei_catRef } + element textClass { + tei_catRef + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:textClass" ] + "\x{a}" ~ + " " + ] + } tei_catRef = ## (category reference) specifies one or more defined categories within some taxonomy or text typology. [2.4.3. The Text Classification] @@ -1683,7 +1832,7 @@ tei_div = | tei_u)+) >> sch:pattern [ id = - "parlamint-div-abstractModel-structure-div-in-l-constraint-rule-12" + "parlamint-div-abstractModel-structure-div-in-l-constraint-rule-13" "\x{a}" ~ " " sch:rule [ @@ -1704,7 +1853,7 @@ tei_div = ] >> sch:pattern [ id = - "parlamint-div-abstractModel-structure-div-in-ab-or-p-constraint-rule-13" + "parlamint-div-abstractModel-structure-div-in-ab-or-p-constraint-rule-14" "\x{a}" ~ " " sch:rule [ @@ -1745,12 +1894,30 @@ tei_particDesc = ## (participation description) describes the identifiable speakers and organisations in a ParlaMint corpus. This informations is given in the corpus root teiHeder. Note that the listPerson and listOrg elements are typically stored in separate files. [16.2. Contextual Information] element particDesc { - (tei_listOrg | tei_include), (tei_listPerson | tei_include) + ((tei_listOrg | tei_include), (tei_listPerson | tei_include)) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:particDesc" ] + "\x{a}" ~ + " " + ] } tei_settingDesc = ## (setting description) describes the setting or settings within which a language interaction takes place, or other places otherwise referred to in a text, edition, or metadata. [16.2. Contextual Information 2.4. The Profile Description] - element settingDesc { tei_setting+ } + element settingDesc { + (tei_setting+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:settingDesc" ] + "\x{a}" ~ + " " + ] + } tei_setting = ## describes one particular setting in which a language interaction takes place. [16.2.3. The Setting Description] @@ -1773,7 +1940,15 @@ tei_recording = ## (recording event) provides details of an audio or video recording event used as the source of a spoken text, either directly or from a public broadcast. [8.2. Documenting the Source of Transcribed Speech 16.3.2. Declarable Elements] element recording { - tei_media+, + (tei_media+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:recording" ] + "\x{a}" ~ + " " + ], ## the kind of recording. [ a:defaultValue = "audio" ] @@ -2359,7 +2534,15 @@ tei_listOrg = ## (list of organizations) contains a list of elements, each of which provides information about an identifiable organisation. [14.2.2. Organizational Names] element listOrg { - (tei_head*, tei_org+, tei_listRelation?), + (tei_head*, tei_org+, tei_listRelation?) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:listOrg" ] + "\x{a}" ~ + " " + ], tei_att.global.attribute.xmlid, tei_att.global.attribute.xmllang, empty @@ -2367,12 +2550,30 @@ tei_listOrg = tei_listEvent = ## (list of events) contains a list of descriptions, each of which provides information about an identifiable event. [14.3.1. Basic Principles] - element listEvent { tei_head*, tei_event* } + element listEvent { + (tei_head*, tei_event*) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:listEvent" ] + "\x{a}" ~ + " " + ] + } tei_listPerson = ## (list of persons) contains a list of descriptions, each of which provides information about an identifiable person or a group of people, for example the participants in a language interaction, or the people referred to in a historical source. [14.3.2. The Person Element 16.2. Contextual Information 2.4. The Profile Description 16.3.2. Declarable Elements] element listPerson { - (tei_head*, tei_person+), + (tei_head*, tei_person+) + >> sch:pattern [ + is-a = "declarable" + "\x{a}" ~ + " " + sch:param [ name = "tde" value = "tei:listPerson" ] + "\x{a}" ~ + " " + ], tei_att.global.attribute.xmlid, tei_att.global.attribute.xmllang, empty @@ -2525,7 +2726,7 @@ tei_relation = element relation { empty >> sch:pattern [ - id = "parlamint-relation-ref-or-key-or-name-constraint-rule-14" + id = "parlamint-relation-ref-or-key-or-name-constraint-rule-15" "\x{a}" ~ " " sch:rule [ @@ -2543,7 +2744,7 @@ tei_relation = " " ] >> sch:pattern [ - id = "parlamint-relation-active-mutual-constraint-rule-15" + id = "parlamint-relation-active-mutual-constraint-rule-16" "\x{a}" ~ " " sch:rule [ @@ -2561,7 +2762,7 @@ tei_relation = " " ] >> sch:pattern [ - id = "parlamint-relation-active-passive-constraint-rule-16" + id = "parlamint-relation-active-passive-constraint-rule-17" "\x{a}" ~ " " sch:rule [ @@ -2700,7 +2901,7 @@ tei_link = element link { empty >> sch:pattern [ - id = "parlamint-link-linkTargets3-constraint-rule-17" + id = "parlamint-link-linkTargets3-constraint-rule-18" "\x{a}" ~ " " sch:rule [ @@ -2798,7 +2999,7 @@ tei_s = | tei_pb)+, (tei_linkGrp?) >> sch:pattern [ - id = "parlamint-s-noNestedS-constraint-rule-18" + id = "parlamint-s-noNestedS-constraint-rule-19" "\x{a}" ~ " " sch:rule [ @@ -2966,7 +3167,7 @@ tei_att.global.linking.attribute.synch = }? tei_att.global.linking.attribute.next = - ## points to the next element of a virtual aggregate of which the current element is part. + ## (next) points to the next element of a virtual aggregate of which the current element is part. attribute next { xsd:anyURI { pattern = "\S+" } }? @@ -2998,7 +3199,7 @@ tei_att.global.source.attribute.source = }? sch:pattern [ id = - "parlamint-att.global.source-source-only_1_ODD_source-constraint-rule-19" + "parlamint-att.global.source-source-only_1_ODD_source-constraint-rule-20" "\x{a}" ~ " " sch:rule [ @@ -3098,7 +3299,7 @@ tei_att.linguistic.attribute.join = }? sch:pattern [ id = - "parlamint-att.measurement-att-measurement-unitRef-constraint-rule-20" + "parlamint-att.measurement-att-measurement-unitRef-constraint-rule-21" "\x{a}" ~ " " sch:rule [ @@ -3169,7 +3370,7 @@ tei_att.typed.attribute.subtype = xsd:token { pattern = "[^\p{C}\p{Z}]+" } }? sch:pattern [ - id = "parlamint-att.typed-subtypeTyped-constraint-rule-21" + id = "parlamint-att.typed-subtypeTyped-constraint-rule-22" "\x{a}" ~ " " sch:rule [ diff --git a/TEI/ParlaMint.odd.rng b/TEI/ParlaMint.odd.rng index 0a00d10d4..d12f92ad2 100644 --- a/TEI/ParlaMint.odd.rng +++ b/TEI/ParlaMint.odd.rng @@ -5,8 +5,8 @@ xmlns:xlink="http://www.w3.org/1999/xlink" datatypeLibrary="http://www.w3.org/2001/XMLSchema-datatypes" ns="http://www.tei-c.org/ns/1.0"> @@ -255,6 +255,25 @@ TEI Edition Location: https://www.tei-c.org/Vault/P5/4.10.0a/ + + + + When there is more than one , each must have an @xml:id + + + When there is more than one , one and only one must have a @default of 'true'. + + + @@ -398,7 +417,7 @@ TEI Edition Location: https://www.tei-c.org/Vault/P5/4.10.0a/ + id="parlamint-att.spanning-spanTo-spanTo-points-to-following-constraint-rule-6"> + id="parlamint-att.calendarSystem-calendar-calendar_attr_on_empty_element-constraint-rule-7"> + id="parlamint-p-abstractModel-structure-p-in-ab-or-p-constraint-rule-8"> + id="parlamint-p-abstractModel-structure-p-in-l-constraint-rule-9"> + id="parlamint-desc-deprecationInfo-only-in-deprecated-constraint-rule-10"> + id="parlamint-ref-refAtts-constraint-rule-11"> + + + @@ -2343,6 +2372,17 @@ Elements] + + + @@ -2368,6 +2408,17 @@ Elements] + + + @@ -2395,6 +2446,17 @@ Elements] + + + @@ -2409,6 +2471,17 @@ Elements] + + + @@ -2417,6 +2490,17 @@ Elements] + + + @@ -2425,6 +2509,17 @@ Elements] + + + @@ -2433,9 +2528,20 @@ Elements] + + + + id="parlamint-quotation-quotationContents-constraint-rule-12"> + + + @@ -2465,6 +2582,17 @@ Elements] + + + @@ -2663,6 +2791,17 @@ Elements] + + + @@ -2696,6 +2835,17 @@ Elements] (text classification) groups information which describes the nature or topic of a text in terms of a standard classification scheme, thesaurus, etc. [2.4.3. The Text Classification] + + + @@ -2822,7 +2972,7 @@ Elements] + id="parlamint-div-abstractModel-structure-div-in-l-constraint-rule-13"> + id="parlamint-div-abstractModel-structure-div-in-ab-or-p-constraint-rule-14"> + + + @@ -2888,6 +3048,16 @@ Elements] + + + @@ -2965,6 +3135,16 @@ Elements] + + + + + + @@ -3733,6 +3924,17 @@ Elements] + + + @@ -3746,6 +3948,17 @@ Elements] + + + @@ -3914,7 +4127,7 @@ Elements] + id="parlamint-relation-ref-or-key-or-name-constraint-rule-15"> + id="parlamint-relation-active-mutual-constraint-rule-16"> + id="parlamint-relation-active-passive-constraint-rule-17"> + id="parlamint-link-linkTargets3-constraint-rule-18"> + id="parlamint-s-noNestedS-constraint-rule-19"> - points to the next element of a virtual aggregate of which the current element is part. + (next) points to the next element of a virtual aggregate of which the current element is part. \S+ @@ -4472,7 +4685,7 @@ Elements] + id="parlamint-att.global.source-source-only_1_ODD_source-constraint-rule-20"> + id="parlamint-att.measurement-att-measurement-unitRef-constraint-rule-21"> + id="parlamint-att.typed-subtypeTyped-constraint-rule-22"> CLARIN - 2025-07-08 + 2025-09-30

This file is freely available and you are hereby authorised to copy, modify, and redistribute it in any way without further reference or permissions.

@@ -46,8 +46,42 @@ target="https://www.clarin.eu/">CLARIN
.

+ + + English + Bulgarian + Bosnian + Catalan + Czech + Danish + German + Greek + Spanish + Estonian + Basque + Finnish + French + Galician + Croatian + Hungarian + Icelandic + Italian + Latvian + Norwegian bokmål + Norwegian nynorsk + Dutch + Polish + Portuguese + Russian + Slovenian + Serbian + Swedish + Turkish + Ukrainian + + - Tomaž Erjavec: Correct some errata. + Tomaž Erjavec: Correct some errata found when making the PressMint ODD. Tomaž Erjavec: Add discussion on local taxonomy IDs and topic and sentiment taxonomies. Tomaž Erjavec: Allow ref/@xml:lang and several publisher/ref. Tomaž Erjavec: Allow ref/@xml:lang and several publisher/ref. @@ -81,7 +115,7 @@ ParlaMint corpora - 2025-07-08 + 2025-09-30

@@ -157,9 +191,9 @@ structure of a ParlaMint corpus, and introduces the distinction between the corpus root and corpus components; - Chapter 3 explains some general requirements - and the file-naming conventions a ParlaMint corpus has to meet; it also introduces - the top level elements and their attributes and the main pointing attributes; + Chapter 3 explains basic requirements + and the file conventions a ParlaMint corpus and introduces + the top level XML elements and their attributes incl. links; Chapter 4 concentrates on the stucture and encoding of the corpus metadata, such as the title information, documenting the source @@ -360,10 +394,9 @@ ... - The lingistically annotated version of the corpus is stored - separately, with the main directory and, as mentioned, the - corpus root and component filenames having the - additional suffix .ana, e.g. + The lingistically annotated version of the corpus is stored separately, with the main directory + and, as mentioned, the corpus root, its metadata files for linguistic annotation, and component filenames + having the additional suffix .ana, e.g. ParlaMint-BE.TEI.ana/ParlaMint-BE.ana.xml @@ -371,8 +404,8 @@ ParlaMint-BE.TEI.ana/ParlaMint-BE-listOrg.xml ParlaMint-BE.TEI.ana/ParlaMint-taxonomy-parla.legislature.xml ParlaMint-BE.TEI.ana/ParlaMint-taxonomy-speaker_types.xml - ParlaMint-taxonomy-NER.xml - ParlaMint-taxonomy-UD.xml + ParlaMint-taxonomy-NER.ana.xml + ParlaMint-taxonomy-UD.ana.xml ... ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-06-19.ana.xml ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-06-30.ana.xml @@ -420,8 +453,11 @@ into a single space. As the use of this character complicates (or breaks) further processing, esp. linguistic annotation, thise character must be substituted by the normal space character (U+0020). The same holds for other - variants of Unicode space characters (U+2000 - U+200A), which are, however, used much less - frequently. + variants of Unicode space characters (U+2000 - U+200A), which are, however, + used much less frequently. + + ZERO WIDTH NO-BREAK SPACE (U+FEFF), also used as the Byte Order Mark (BOM) + in Windows files should be removed. NON-BREAKING HYPHEN (U+2011), similarly to NO-BREAK SPACE, prevents a line break, in this case following its position. With a similar reasoning as @@ -466,10 +502,10 @@ or, and only if a two letter code does not exist for a language, the three-letter ISO 639-2/T code. For example, the code for Basque is 'eu'. ParlaMint - corpora will use at least two languages, i.e. the language that the - transcriptions are written in, which we will call the local - language and English, as the meta-language, which is (also) used in - the metadata. + corpora will use at least two languages (except for Great Britain), i.e. the + language that the transcriptions are written in, which we will call the + local language and English, as the meta-language, which is (also) + used in the metadata. Temporal, i.e. time-related information is typically stored in the when, from and to attributes of various @@ -504,10 +540,14 @@ special importance: xmlns determines the namespace of the element, and this should - always be the TEI namespace, i.e. http://www.tei-c.org/ns/1.0. Note - that all lower level elements inherit this namespace. - - xml:id is an attribute form the (implicitly assumed) XML + always be the TEI namespace, i.e. http://www.tei-c.org/ns/1.0 + (apart from the elements using the XInclude directive, cf. the Section on + Use of XInclude). + Note that lower level elements inherit the namespace of the superordinate element, + unless explicitly overridden, so it is only necessary to specify the TEI namespace on + the root element of a file. + + xml:id is an attribute from the (implicitly assumed) XML namespace, and gives the identifier of the element bearing it. The value of an ID should be unique in the corpus as a whole and should obey format requirements as defined by xml:lang is also a global attribute and gives the language - code of the text content of the element; for the corpus root this does not - (just) mean the content of its TEI header, but primarily the textual content - of its (XIncluded) components. The convention is that language of the text + code of the text content of the element; for the corpus root this + means the content of its TEI header, while for corpus components this is + the textual content of their TEI headers and text elements. + + The convention is that language of the text content of an element is determined by the value of the first xml:lang attribute on its ancestor axis. In cases where the content is multilingual, the language code should be of the majority language. When @@ -576,15 +618,18 @@ large number of different elements and that their value is a series of pointers, i.e. a white-space delimited sequence of references to the values of some xml:id attribute in the corpus or, in general, to an URI. + The three attributes are: ana serves to provide an analysis or to classify an element according to some pre-determined vocabulary. In ParlaMint the target element will - typically be a category in a taxonomy, an event or date, or an organisation. + typically be a category in a taxonomy, an event, or an organisation. + corresp points to items that correspond to the current element in some way, e.g. the (URL of a) media file to a page break. + ref provides an explicit reference to the full definition or identity for the entity being named. In ParlaMint it is used e.g. for connecting a person's affiliation with a particular organisation. The value @@ -637,7 +682,7 @@ The when attribute is used when the temporal information refers to a point in time, typically a date, and is used e.g. to give the date - when the corpus was published, or when a change in the corpus was made. + when the corpus was published, or the date when a session was held. The from and to attributes give the starting and ending date or time of an interval, e.g. the time period the corpus covers, or @@ -754,8 +799,8 @@ main or sub of their type attribute and the value of their xml:lang attribute.

-

The main title has a formulaic structure <Country name> - parliamentary corpus ParlaMint-<Country code> [ParlaMint], with an +

The main title has a formulaic structure <Country_name> + parliamentary corpus ParlaMint-<Country_code> [ParlaMint], with an equivalent structure for the local language. Note that the corpus stamp in square brackets can also be [ParlaMint.ana] for the linguistically annotated version of the corpus (as explained in the Chapter on Next come one or more responsibility statements, respStmt, each one containing one or more person names, persName, with an optional - ref attribute, giving the URL, where more information about the + ref attribute, giving the (typically ORCID) URL, where more information about the person can be found, and the responsibility element resp, which specifies what responsibility the statement is about.

@@ -869,8 +914,7 @@ corpus, while the minor version is reserved for e.g. correcting errata or other minor changes. We do not use the patch number. It should be noted that - at least so far - all the ParlaMint corpora were released together, so that they are all - of the same edition, i.e. have the same version number. At the time of writing, - the latest version is 2.1, with the next one planned to be 3.0.

+ of the same edition, i.e. have the same version number.

@@ -929,7 +973,7 @@ Creative Commons Attribution 4.0 International License.

- 11. 6. 2023 + 11. 6. 2023 @@ -943,7 +987,7 @@ contains the handle where the complete corpus corresponding to the specified version can be found. - The availability specifiers, via its licence element the + The availability specifies, via its licence element the fixed-value CC BY 4.0 URL, and in the following paragraph gives a prose description of the licence, including its URL via the target attribute of ref. As usual, the textual information is given in both languages. @@ -958,7 +1002,7 @@ Source description

The source description sourceDesc of the corpus root encodes - the original digital source of the ParlaMint corpus in the bibl element, + the digital source of the ParlaMint corpus in the bibl element, as shown in the following example: @@ -978,7 +1022,7 @@ Apart from the bi-lingual titles, it should also give in idno with the fixed type as URI the government URL where the - transcripts were first harvested from, while the dates of the earliest and latest + transcripts were harvested from, while the dates of the earliest and latest transcript in the corpus are indicated by the from and to attributes of the date element. As usual, the values of these attributes should be according to ISO 8601, while the textual content can be formatted @@ -1076,9 +1120,6 @@ Sciences and Digital Humanities based on the corpus data.

- - The description above is written for the CLARIN ParlaMint project and the English - language part can be used as is in the produced corpora for version 3.

@@ -1172,17 +1213,17 @@ + href="ParlaMint-taxonomy-parla.legislature.xml"/> + href="ParlaMint-taxonomy-speaker_types.xml"/> + href="ParlaMint-taxonomy-subcorpus.xml"/> + href="ParlaMint-taxonomy-politicalOrientation.xml"/> + href="ParlaMint-taxonomy-CHES.xml"/> ]]> As can be seen, three of the taxonomies are general ParlaMint taxonomies, while @@ -1246,7 +1287,7 @@

ParlaMint requires several taxonomies to be defined in the class declaration - of the corpus root (as well as a additionaly ones for the linguistically annotated + of the corpus root (as well as additional ones for the linguistically annotated corpus, as further described in the Section on Linguistic metadata). As mentioned, these taxonomies are defined globally and available as part of the data on the ParlaMint GitHub @@ -1289,7 +1330,7 @@

The profile description, profileDesc is the third main division of the metadata provided by the TEI header. It contains a description of non-bibliographic - aspects of the corpus, for example the list of speakers with their metadata. For + aspects of the corpus, for example, the list of speakers with their metadata. For the corpus root, it contains four elements, of which only the first, the settingDesc is used in corpus components. The elements are listed below: @@ -1488,7 +1529,8 @@ The revision description consists of a series of change elements, with the attribute when giving the date of the change, and the content containing the name of the person responsible for the change, and a - free-text description of the change.

+ free-text description of the change. Note that the change follow reverse + chronological order, i.e. the most recent changes are at the top.

@@ -2628,13 +2670,13 @@ The encoding of the linguistically annotated version differs from the plain-text one in the following: - All the corpus root and components file names + All the corpus root and component file names have the extension .ana.xml. - For example, if the plain-text root has the file name + For example, if the TEI plain-text root has the file name ParlaMint-CZ.xml, the linguistically annotated one should be - ParlaMint-CZ.ana.xml, or - ParlaMint-CZ_2016-04-13.xml and - ParlaMint-CZ_2016-04-13.ana.xml + ParlaMint-CZ.ana.xml, and if the component plain-text files is + ParlaMint-CZ_2016-04-13.xml the linguistically annotated one is + ParlaMint-CZ_2016-04-13.ana.xml. Because the file ID (i.e. the value of the top level element attribute xml:id, as explained in the Section on File names and directory structure), the previous point also means that the linguistically annotated files should have the top level ID suffixed with .ana, - e.g. <teiCorpus xml:id="ParlaMint-CZ.ana"> + e.g. <teiCorpus xml:id="ParlaMint-CZ.ana">.
The corpus stamp in the main title of the corpus root or components (cf. the Section on Title statement) which is @@ -2657,7 +2699,7 @@ The linguistically annotated version of the corpus should also have some added metadata in the TEI header of the corpus root, which is detailed in the - following Section on Metadata for linguistic + Section on Metadata for linguistic annotation.
@@ -2666,7 +2708,7 @@
Linguistic markup -

Linguistic annotation is added only to the text content of seg elements +

Linguistic annotation is added only to the immediate text content of seg elements inside the speeches, i.e u elements. For this text, ParlaMint requires the following additional markup to be present: @@ -2675,13 +2717,13 @@ sentences: what is a sentence; - lemmas: what is the base form of each word; + lemmas: the base form of each word; Universal Dependencies (UD) part-of-speech and morphological features, and, optionally, part-of-speech tags from a different (local) tagset; - named entities (NE): what is a name, categorised at least into the standard + named entities (NE): a name, categorised into the standard four NE classes; the UD dependency syntactic parse of the sentences; @@ -2866,7 +2908,7 @@ Westminster Hall - , + , ...

@@ -3023,7 +3065,7 @@
Application information for linguistic processing -

As the linguistic analysis of a ParlaMint will be performed by a tool, the +

As the linguistic analysis of a ParlaMint corpus will be performed by a tool, the information on which tool (or tools) have been used should be documented in the corpus root TEI header. This information is encoded in the appInfo element of the encodingDesc, as shown in the example below: @@ -3192,19 +3234,20 @@

Prefix definitions -

Pointing attributes, such as ana, take as their value a series of - references to the value of xml:id elements in an XML document. If this is - the same document, then the reference to the ID is the hash character, # - prefixed to the particular ID, e.g. #parla.uni, and if they are in another - XML document, then the hash is prefixed with the URL of the document, e.g. +

Pointing attributes, such as ana, take as their value a reference or + space-delimited series of references to a URL and/or the value of xml:id + elements. If the reference is to an ID, then it is prefixed the hash character, + #, e.g. #parla.uni, and if they are to an ID in another XML + document, then the hash follows the URL of the document, e.g. https://nl.ijs.si/ME/V6/msd/tables/msd-fslib2-sl.xml#Vmpr1p.

-

Because the complete URL tends to be long, which is especially inconvenient - when such references are given to every token in a corpus, TEI introduces the so - called Extended pointer syntax, whereby the reference to an ID can be given in the - form of a prefix, which is separated by a colon from the local part of the ID - reference, and the value of this prefix is determined via the prefixDef - element in the profileDesc of the TEI header.

+

Because complete URLs tend to be long, especially inconvenient when such + references are given to every token in a corpus, TEI introduces the so called + Abbreviated + Pointers, whereby a reference can be given in the form of a + prefix, which is separated by a colon from the local part of the ID reference, and + the value of this prefix is determined via the prefixDef element in the + encodingDesc of the TEI header.

ParlaMint uses this mechanism for all linguistic annotations with a closed vocabulary, in particular @@ -3224,7 +3267,7 @@ + replacementPattern="https://nl.ijs.si/ME/V6/msd/tables/msd-fslib-sl.xml#$1">

Private URIs with this prefix point to feature-structure elements defining the Slovenian MULTEXT-East Version 6 MSDs.

@@ -3233,10 +3276,10 @@ The specialised element for listing prefix definitions, listPrefixDef gives a series of prefix definitions, i.e. prefixDef elements. Each prefix definition defines its prefix as the value of the ident attribute, and then specifies - a regular expression that matches the part of the ID reference after the prefix in its + a regular expression that matches the part of the reference after the prefix in its matchPattern attribute, and its substitution as the value of the replacementPattern attribute. The first prefix definition thus defines the - ud-syn prefix, so for any ID reference with this prefix, + ud-syn prefix, so for any reference with this prefix, e.g. ud-syn:acl_relcl, the part after the prefix (acl_relcl) should be matched against (.+) and the result being the matched part (here the entire relation acl_relcl) substituted by @@ -3245,9 +3288,9 @@ course trivial, and hardly necessary, but was implemented so that all fixed-vocabulary linguistic analyses have the same treatment.

-

More to the point is the second example, where very short ID references, such as +

More to the point is the second example, where very short references, such as mte:Vmpr1p are transformed to - https://nl.ijs.si/ME/V6/msd/tables/msd-fslib2-sl.xml#Vmpr1p, as already + https://nl.ijs.si/ME/V6/msd/tables/msd-fslib-sl#Vmpr1p, as already explained in the Section on Word-level annotation.

@@ -3359,9 +3402,8 @@
- - The structure and encoding of ParlaMint corpora
The structure and encoding of ParlaMint corpora
2025-06-13

Table of contents

1. Introduction

This document is meant to serve as a reference for the encoding of ParlaMint corpora of parliamentary proceedings. In order for the ParlaMint corpora to be interoperable (i.e. so that the same scripts can be used to process them), their structure is fairly rigid, both in terms of file names and folder structure, as well as their TEI XML encoding. This is not to say that all the corpora have to contain exactly the same information because we distinguish obligatory information, which all the corpora should contain, from that which is optional, and present only in the corpora for which it has been possible to gather it from the corpus sources.

This document is a specialisation of Parla-CLARIN, itself a customisation the TEI Guidelines. But while Parla-CLARIN gives fairly general recommendations for encoding corpora of parliamentary proceedings, ParlaMint, as mentioned, is much stricter. This document gives very specific encoding recommendations without necessarily stating the reasons for their choice. It covers the overall structure of ParlaMint corpora, the metadata they contain, the encoding of transcriptions, and, for the linguistically annotated version, the encoding of word-level linguistic annotatios, syntactic dependencies and named entities.

The document is not meant as a tutorial on TEI or ParlaMint, but as a reference to elements, their nesting and attributes. Other sources can help in understanding the encoding and content of ParlaMint corpora:

  • Samples of ParlaMint corpora, available in the Samples/ directory of the ParlaMint GitHub repository; note that the samples in the main branch are supposed to be publication-ready, while those in the data branch are work in progress.
  • Two openly available papers detailing the results of the ParlaMint I and ParlaMint II projects:
    • The ParlaMint corpora of parliamentary proceedings. Language Resources & Evaluation (2022). DOI 10.1007/s10579-021-09574-0.
    • ParlaMint II: Advancing Comparable Parliamentary Corpora Across Europe. Language Resources & Evaluation (2024). DOI 10.1007/s10579-024-09798-w.
  • The Parla-CLARIN guidelines, which provide general guidelines for encoding parliamentary corpora in TEI; they also give links to the relevant chapters of the TEI Guidelines.

The rest of these recommendations are structured as follows:

  • Chapter 2 explains the overall XML structure of a ParlaMint corpus, and introduces the distinction between the corpus root and corpus components;
  • Chapter 3 explains some general requirements and the file-naming conventions a ParlaMint corpus has to meet; it also introduces the top level elements and their attributes and the main pointing attributes;
  • Chapter 4 concentrates on the stucture and encoding of the corpus metadata, such as the title information, documenting the source of the corpus, taxonomies used etc.;
  • Chapter 5 explains how and what information must be encoded about the persons giving the speeches and the (political) organisations they belong to;
  • Chapter 6 treats the encoding of the transcripts, including speeches and transcriber notes;
  • Chapter 7 details the addition of linguistic annotations to the corpus;
  • Chapter 8 introduces scripts to finalise, validate and convert a ParlaMint corpus to other formats;
  • Chapter 9 gives instructions on how to contribute samples of a ParlaMint corpus to GitHub;
  • Appendix A gives the formal specification of the Parla-CLARIN schema.

2. Overall corpus structure

2.1. XML structure

The parliamentary proceeding of one country of autonomous region constitute one ParlaMint corpus, which is stored as one XML document, with <teiCorpus> as its top-level element. It is composed of a <teiHeader>, giving the metadata for the corpus as a whole (further detailed in the Section on Corpus metadata), followed by a series of <TEI> elements that each contain one corpus component, as illustrated1 below:
+The structure and encoding of ParlaMint corpora
The structure and encoding of ParlaMint corpora
2025-09-30

Table of contents

1. Introduction

This document is meant to serve as a reference for the encoding of ParlaMint corpora of parliamentary proceedings. In order for the ParlaMint corpora to be interoperable (i.e. so that the same scripts can be used to process them), their structure is fairly rigid, both in terms of file names and folder structure, as well as their TEI XML encoding. This is not to say that all the corpora have to contain exactly the same information because we distinguish obligatory information, which all the corpora should contain, from that which is optional, and present only in the corpora for which it has been possible to gather it from the corpus sources.

This document is a specialisation of Parla-CLARIN, itself a customisation the TEI Guidelines. But while Parla-CLARIN gives fairly general recommendations for encoding corpora of parliamentary proceedings, ParlaMint, as mentioned, is much stricter. This document gives very specific encoding recommendations without necessarily stating the reasons for their choice. It covers the overall structure of ParlaMint corpora, the metadata they contain, the encoding of transcriptions, and, for the linguistically annotated version, the encoding of word-level linguistic annotatios, syntactic dependencies and named entities.

The document is not meant as a tutorial on TEI or ParlaMint, but as a reference to elements, their nesting and attributes. Other sources can help in understanding the encoding and content of ParlaMint corpora:

  • Samples of ParlaMint corpora, available in the Samples/ directory of the ParlaMint GitHub repository; note that the samples in the main branch are supposed to be publication-ready, while those in the data branch are work in progress.
  • Two openly available papers detailing the results of the ParlaMint I and ParlaMint II projects:
    • The ParlaMint corpora of parliamentary proceedings. Language Resources & Evaluation (2022). DOI 10.1007/s10579-021-09574-0.
    • ParlaMint II: Advancing Comparable Parliamentary Corpora Across Europe. Language Resources & Evaluation (2024). DOI 10.1007/s10579-024-09798-w.
  • The Parla-CLARIN guidelines, which provide general guidelines for encoding parliamentary corpora in TEI; they also give links to the relevant chapters of the TEI Guidelines.

The rest of these recommendations are structured as follows:

  • Chapter 2 explains the overall XML structure of a ParlaMint corpus, and introduces the distinction between the corpus root and corpus components;
  • Chapter 3 explains basic requirements and the file conventions a ParlaMint corpus and introduces the top level XML elements and their attributes incl. links;
  • Chapter 4 concentrates on the stucture and encoding of the corpus metadata, such as the title information, documenting the source of the corpus, taxonomies used etc.;
  • Chapter 5 explains how and what information must be encoded about the persons giving the speeches and the (political) organisations they belong to;
  • Chapter 6 treats the encoding of the transcripts, including speeches and transcriber notes;
  • Chapter 7 details the addition of linguistic annotations to the corpus;
  • Chapter 8 introduces scripts to finalise, validate and convert a ParlaMint corpus to other formats;
  • Chapter 9 gives instructions on how to contribute samples of a ParlaMint corpus to GitHub;
  • Appendix A gives the formal specification of the ParlaMint schema.

2. Overall corpus structure

2.1. XML structure

The parliamentary proceeding of one country or autonomous region constitute one ParlaMint corpus, which is stored as one XML document, with <teiCorpus> as its top-level element. It is composed of a <teiHeader>, giving the metadata for the corpus as a whole (further detailed in the Section on Corpus metadata), followed by a series of <TEI> elements that each contain one corpus component, as illustrated1 below:
             <!-- Corpus root --> <teiCorpus xmlns="http://www.tei-c.org/ns/1.0"> @@ -10,7 +10,7 @@           
Each corpus component should contain at most the transcripts for one day, although several components can contain the transcript for the same day, e.g. for different (types of) meetings. How and if these further subdivisions into separate components are realised is dependent on the corpus, as the granularity of parliamentary proceedings corpora, not to mention the national rules of structuring the workings of the parliament, differ substantially.
A corpus component will thus be rooted in the <TEI> element, which then contains its metadata in its own <teiHeader>, followed by the <text> element, which contains the transcription of the particular component, as illustrated below:
<TEI xmlns="http://www.tei-c.org/ns/1.0">  <teiHeader>...</teiHeader>  <text>...</text> -</TEI>

The <teiHeader> of a corpus component (further detailed in the Section on Corpus metadata) contains the metadata specific for this component (along with some redundant metadata about the provenance), and which should be unique in the corpus, i.e. the corpus component metadata should distinguish it from all the other components of the corpus.

2.2. Use of XInclude

The fact that a corpus is one XML document does not mean that it is also stored in one file. In fact, ParlaMint requires that each corpus component is stored in a separate file, with the corpus root, i.e. the top-level <teiCorpus>, also stored as one file. Furthermore, some parts of the corpus root metadata are also stored in separate files.

To enable one XML document to be composed of many files, we use the XInclude mechanism, and the corpus root uses this mechanism (i.e. the <include> elements in the XInclude namespace) to include its corpus component files, so a corpus root will be in fact encoded similarly to the following example:
+</TEI>

The <teiHeader> of a corpus component (further detailed in the Section on Corpus metadata) contains the metadata specific for this component (along with some redundant metadata about its provenance), and which should be unique in the corpus, i.e. the corpus component metadata should distinguish it from all the other components of the corpus.

2.2. Use of XInclude

The fact that a corpus is one XML document does not mean that it is also stored in one file. In fact, ParlaMint requires that each corpus component is stored in a separate file, with the corpus root, i.e. the top-level <teiCorpus>, also stored as one file. Furthermore, some parts of the corpus root metadata are also stored in separate files.

To enable one XML document to be composed of many files, we use the XInclude mechanism, and the corpus root uses this mechanism (i.e. the <include> elements in the XInclude namespace) to include its corpus component files, so a corpus root will be in fact encoded similarly to the following example:
           <!-- Corpus root file --> <teiCorpus xmlns="http://www.tei-c.org/ns/1.0" >  @@ -21,18 +21,18 @@       href="2014/ParlaMint-NL_2014-04-17.xml"/>  <!-- Corpus component file -->   ...                                            <!-- More corpus component files --> </teiCorpus> -          

Apart from corpus components, some parts of the overall corpus metadata (i.e. the <teiCorpus> <teiHeader> element) are also stored as separate files, and hence also included in the corpus root using the same XInclude mechanism as explained above.

2.3. File names and directory structure

ParlaMint has strict rules on how to name the various files that constitute a corpus, and how to collect them in directories.

The file names have the the following structure:

  • The corpus root file name should start with the string ParlaMint-, followed by the ISO 3166 country (or automous region) code (cf. Section on Standard values) e.g. ParlaMint-NL.xml or ParlaMint-ES-CT.
  • For machine-translated corpora the ISO 639 code of the language (cf. Section on Standard values) should follow the country code, e.g. ParlaMint-NL-en.xml.
  • A corpus component filename should start with the name of the root, followed by an underscore and the ISO 8601 formatted date of the transcript, for example ParlaMint-IS_2015-01-21-54.xml. In case a corpus component is further distinguished, so that there are are several components with the same date, the corpus compilers are free to extend the file name by a hyphen and any suffix containing only ASCII letters and numbers and the hyphen character, e.g. ParlaMint-NL_2018-10-30-eerstekamer-4.xml or ParlaMint-CZ_2016-04-13-ps2013-044-02-016-098.xml
  • Certain metadata elements from the corpus root <teiHeader> are stored in separate files, in particular the list of speakers, <listPerson>, the list of political parties and other organisations, <listOrg>, and the ParlaMint structural and linguistic taxonomies, i.e. <taxonomy> elements. The file names for such metadata files start with the name of the corpus root, followed by a hyphen, and then the name of the element, e.g. ParlaMint-BE-listPerson.xml. Where there are more files for instances of the same element name, as is the case for taxonomies, the filename should end with another hypen, followed by the ID of the particular element, e.g. ParlaMint-BE-taxonomy-UD-SYN.xml. Finally, some of the taxonomies are not corpus-specific, i.e. identical files are used by all ParlaMint corpora. In this case, the country or region code is ommitted, e.g. ParlaMint-taxonomy-parla.legislature.xml.
  • The file names of the corpus as a whole or corpus components that have been automatically converted from the source XML into some other format should have the same name as the corpus root or components, respectively, but with appropriate file extensions, e.g, ParlaMint-IS_2015-01-21-54.txt; this is further explained in the Section on Conversions.
  • As discussed in the Chapter on Linguistic annotation we distinguish the linguistically annotated version of the corpus from the ‘plain-text’ one, with the linguistic annotated version having the additional suffix .ana on the corpus root and components, e.g. ParlaMint-ES-CT.ana.xml or ParlaMint-IS_2015-01-21-54.ana.xml.

For distribution the complete XML corpus should be stored in a directory that has the same name prefix as the corpus root file. The directory then contains the corpus root file and its metadata files, while the corpus components should be in subdirectories, one per year, for example:

 ParlaMint-BE.TEI/ParlaMint-BE.xml
ParlaMint-BE.TEI/ParlaMint-BE-listPerson.xml
ParlaMint-BE.TEI/ParlaMint-BE-listOrg.xml
ParlaMint-BE.TEI/ParlaMint-taxonomy-parla.legislature.xml
ParlaMint-BE.TEI/ParlaMint-taxonomy-speaker_types.xml
...
ParlaMint-BE.TEI/2014/ParlaMint-BE_2014-06-19.xml
ParlaMint-BE.TEI/2014/ParlaMint-BE_2014-06-30.xml
ParlaMint-BE.TEI/2014/ParlaMint-BE_2014-07-17.xml
...
ParlaMint-BE.TEI/2015/ParlaMint-BE_2015-01-06-54.xml
ParlaMint-BE.TEI/2015/ParlaMint-BE_2015-01-07-54.xml
ParlaMint-BE.TEI/2015/ParlaMint-BE_2015-01-08-54.xml
...

The lingistically annotated version of the corpus is stored separately, with the main directory and, as mentioned, the corpus root and component filenames having the additional suffix .ana, e.g.

 ParlaMint-BE.TEI.ana/ParlaMint-BE.ana.xml
ParlaMint-BE.TEI.ana/ParlaMint-BE-listPerson.xml
ParlaMint-BE.TEI.ana/ParlaMint-BE-listOrg.xml
ParlaMint-BE.TEI.ana/ParlaMint-taxonomy-parla.legislature.xml
ParlaMint-BE.TEI.ana/ParlaMint-taxonomy-speaker_types.xml
ParlaMint-taxonomy-NER.xml
ParlaMint-taxonomy-UD.xml
...
ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-06-19.ana.xml
ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-06-30.ana.xml
ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-07-17.ana.xml
...
ParlaMint-BE.TEI.ana/2015/ParlaMint-BE_2015-01-06-54.ana.xml
ParlaMint-BE.TEI.ana/2015/ParlaMint-BE_2015-01-07-54.ana.xml
ParlaMint-BE.TEI.ana/2015/ParlaMint-BE_2015-01-08-54.ana.xml
...

3. General requirements

This section gives some general requirements a ParlaMint corpus has to meet, in particular those relating to the characters in a corpus, and the use of standards. It also details the structure of the file names of the ParlaMint root and component files, as well as the attributes expected on the <teiCorpus> and <TEI> tags.

3.1. Characters

The corpus should be encoded in Unicode, using the UTF-8 character encoding, at least for European languages. In cases where the original contains characters from the Unicode Private Use Area, these should, if possible, be given their closest Unicode equivalents or substituted by the Unicode replacement character U+FFFD. End-of-line hyphens, if present in the source files, should be removed, and the split words joined in order to enhance searching the corpus and to simplify linguistic processing.

The following characters, esp. prevalent when the source documents were in Word or HTML, deserve special mention:

  • TAB (U+0009) character helps the alignment of strings on successive lines. As ParlaMint is not interested in preserving the layout, all TAB chacters are substituted by space characters (U+0020).
  • NO-BREAK SPACE (U+00A0) prevents, with some applications, an automatic line break at its position and also collapsing such consecutive characters into a single space. As the use of this character complicates (or breaks) further processing, esp. linguistic annotation, these characters should be substituted by the normal space character (U+0020). The same holds for other variants of spaces (U+2000 - U+200A), which are, however, used much less frequently.
  • NON-BREAKING HYPHEN (U+2011), similarly to NO-BREAK SPACE, prevents a line break, in this case following its position. With a similar reasoning as above, this character should be substituted by the normal hyphen character ('-', U+002D).
  • SOFT HYPHEN (U+00AD) indicates that a word can be hyphenated at that point. Occurrences of this character should be removed from the corpus.

Text-bearing elements should also not start or end with space characters, and sequences of whitespace characters should be changed into a single space.

3.2. Standard values

Whenever possible, ParlaMint uses standards for information coding. In particular, the following information must be standardised:

  • As the identity of a ParlaMint corpus is determined by the country or region of the particular parliament, its code appears in many places. For specifying these codes, the ISO 3166 standard should be used, in particular ISO 3166-1 alpha-2 for the two letter codes of the countries (for national parliaments) and ISO 3166-2 for the names of country subdivision (for parliaments of autonomous provinces,). So, for example, the country code for Spain is "ES", while the code for the autonomous Basque community is "ES-PV". Note that we use the term regional parliaments for such cases.
  • The codes for the languages used in the corpora (i.e. the possible values of the xml:lang attribute) should follow BCP 47 (cf. also xml:lang in XML document schemas. Essentially, this means that the value for a language code should have two letters, following ISO 639-1 or, and only if a two letter code does not exist for a language, the three-letter ISO 639-2/T code. For example, the code for Basque is 'eu'. ParlaMint corpora will use at least two languages, i.e. the language that the transcriptions are written in, which we will call the local language and English, as the meta-language, which is (also) used in the metadata.
  • Temporal, i.e. time-related information is typically stored in the when, from and to attributes of various elements. To specify a date or time as the value of these attributes, formatting according to the ISO 8601 standard should be used, e.g. 2022-04-01 for the 1st of April 2022. More information on temporal attributes is given in the Section on Temporal attributes.

3.3. Attributes of top-level elements

The Chapter on Overall corpus structure introduced the top level elements of the corpus root file and of the component files (i.e. the <teiCorpus> and <TEI> elements), but did not elaborate on their attributes; these are presented in this section.

The corpus root has three required attributes, as shown below:
+          

Apart from corpus components, some parts of the overall corpus metadata (i.e. the <taxonomy> elements) are also stored as separate files, and hence also included in the corpus root using the same XInclude mechanism as explained above.

2.3. File names and directory structure

ParlaMint has strict rules on how to name the various files that constitute a corpus, and how to collect them in directories.

The file names have the the following structure:

  • The corpus root file name should start with the string ParlaMint-, followed by the ISO 3166 country (or automous region) code (cf. Section on Standard values) e.g. ParlaMint-NL.xml or ParlaMint-ES-CT.
  • For machine-translated corpora the ISO 639 code of the language (cf. Section on Standard values) should follow the country code, e.g. ParlaMint-NL-en.xml.
  • A corpus component filename should start with the name of the root, followed by an underscore and the ISO 8601 formatted date of the transcript, for example ParlaMint-IS_2015-01-21-54.xml. In case a corpus component is further distinguished, so that there are are several components with the same date, the corpus compilers are free to extend the file name by a hyphen and any suffix containing only ASCII letters and numbers and the hyphen character, e.g. ParlaMint-NL_2018-10-30-eerstekamer-4.xml or ParlaMint-CZ_2016-04-13-ps2013-044-02-016-098.xml
  • Certain metadata elements from the corpus root <teiHeader> are stored in separate files, in particular the list of speakers, <listPerson>, the list of political parties and other organisations, <listOrg>, and the ParlaMint structural and linguistic taxonomies, i.e. <taxonomy> elements. The file names for such metadata files start with the name of the corpus root, followed by a hyphen, and then the name of the element, e.g. ParlaMint-BE-listPerson.xml. Where there are more files for instances of the same element name, as is the case for taxonomies, the filename should end with another hypen, followed by the ID of the particular element, e.g. ParlaMint-BE-taxonomy-UD-SYN.xml. Finally, some of the taxonomies are not corpus-specific, i.e. identical files are used by all ParlaMint corpora. In this case, the country or region code is ommitted, e.g. ParlaMint-taxonomy-parla.legislature.xml.
  • The file names of the corpus as a whole or corpus components that have been automatically converted from the source XML into some other format should have the same name as the corpus root or components, respectively, but with appropriate file extensions, e.g, ParlaMint-IS_2015-01-21-54.txt; this is further explained in the Section on Conversions.
  • As discussed in the Chapter on Linguistic annotation we distinguish the linguistically annotated version of the corpus from the ‘plain-text’ one, with the linguistic annotated version having the additional suffix .ana on the corpus root and components, e.g. ParlaMint-ES-CT.ana.xml or ParlaMint-IS_2015-01-21-54.ana.xml.

For distribution the complete XML corpus should be stored in a directory that has the same name prefix as the corpus root file and extended with the format (e.g. TEI). The directory then contains the corpus root file and its metadata files, while the corpus components should be in subdirectories, one per year, for example:

 ParlaMint-BE.TEI/ParlaMint-BE.xml
ParlaMint-BE.TEI/ParlaMint-BE-listPerson.xml
ParlaMint-BE.TEI/ParlaMint-BE-listOrg.xml
ParlaMint-BE.TEI/ParlaMint-taxonomy-parla.legislature.xml
ParlaMint-BE.TEI/ParlaMint-taxonomy-speaker_types.xml
...
ParlaMint-BE.TEI/2014/ParlaMint-BE_2014-06-19-11.xml
ParlaMint-BE.TEI/2014/ParlaMint-BE_2014-06-30-11.xml
ParlaMint-BE.TEI/2014/ParlaMint-BE_2014-07-17-11.xml
...
ParlaMint-BE.TEI/2015/ParlaMint-BE_2015-01-06-54.xml
ParlaMint-BE.TEI/2015/ParlaMint-BE_2015-01-07-54.xml
ParlaMint-BE.TEI/2015/ParlaMint-BE_2015-01-08-54.xml
...

The lingistically annotated version of the corpus is stored separately, with the main directory and, as mentioned, the corpus root, its metadata files for linguistic annotation, and component filenames having the additional suffix .ana, e.g.

 ParlaMint-BE.TEI.ana/ParlaMint-BE.ana.xml
ParlaMint-BE.TEI.ana/ParlaMint-BE-listPerson.xml
ParlaMint-BE.TEI.ana/ParlaMint-BE-listOrg.xml
ParlaMint-BE.TEI.ana/ParlaMint-taxonomy-parla.legislature.xml
ParlaMint-BE.TEI.ana/ParlaMint-taxonomy-speaker_types.xml
ParlaMint-taxonomy-NER.ana.xml
ParlaMint-taxonomy-UD.ana.xml
...
ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-06-19.ana.xml
ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-06-30.ana.xml
ParlaMint-BE.TEI.ana/2014/ParlaMint-BE_2014-07-17.ana.xml
...
ParlaMint-BE.TEI.ana/2015/ParlaMint-BE_2015-01-06-54.ana.xml
ParlaMint-BE.TEI.ana/2015/ParlaMint-BE_2015-01-07-54.ana.xml
ParlaMint-BE.TEI.ana/2015/ParlaMint-BE_2015-01-08-54.ana.xml
...

3. General requirements

This section gives some general requirements a ParlaMint corpus has to meet, in particular those relating to the characters in a corpus, and the use of standards. It also details the structure of the file names of the ParlaMint root and component files, as well as the attributes expected on the <teiCorpus> and <TEI> tags.

3.1. Characters

The corpus should be encoded in Unicode, using the UTF-8 character encoding, at least for European languages. In cases where the original contains characters from the Unicode Private Use Area, these should, if possible, be given their closest Unicode equivalents or substituted by the Unicode replacement character U+FFFD. End-of-line hyphens, if present in the source files, should be removed, and the split words joined in order to enhance searching the corpus and to simplify linguistic processing.

The following characters, esp. prevalent when the source documents were in Word or HTML, deserve special mention:

  • TAB (U+0009) character helps the alignment of strings on successive lines. As ParlaMint is not interested in preserving the layout, all TAB chacters must be substituted by space characters (U+0020).
  • NO-BREAK SPACE (U+00A0) prevents, with some applications, an automatic line break at its position and also collapsing such consecutive characters into a single space. As the use of this character complicates (or breaks) further processing, esp. linguistic annotation, thise character must be substituted by the normal space character (U+0020). The same holds for other variants of Unicode space characters (U+2000 - U+200A), which are, however, used much less frequently.
  • ZERO WIDTH NO-BREAK SPACE (U+FEFF), also used as the Byte Order Mark (BOM) in Windows files should be removed.
  • NON-BREAKING HYPHEN (U+2011), similarly to NO-BREAK SPACE, prevents a line break, in this case following its position. With a similar reasoning as above, this character should be substituted by the normal hyphen character ('-', U+002D).
  • SOFT HYPHEN (U+00AD) indicates that a word can be hyphenated at that point. Occurrences of this character should be removed from the corpus.

Text-bearing elements should also not start or end with space characters, and sequences of whitespace characters should be changed into a single space.

3.2. Standard values

Whenever possible, ParlaMint uses standards for information coding. In particular, the following information must be standardised:

  • As the identity of a ParlaMint corpus is determined by the country or region of the particular parliament, its code appears in many places. For specifying these codes, the ISO 3166 standard should be used, in particular ISO 3166-1 alpha-2 for the two letter codes of the countries (for national parliaments) and ISO 3166-2 for the names of country subdivision (for parliaments of autonomous provinces,). So, for example, the country code for Spain is "ES", while the code for the autonomous Basque community is "ES-PV". Note that we use the term regional parliaments for such cases.
  • The codes for the languages used in the corpora (i.e. the possible values of the xml:lang attribute) should follow BCP 47 (cf. also xml:lang in XML document schemas. Essentially, this means that the value for a language code should have two letters, following ISO 639-1 or, and only if a two letter code does not exist for a language, the three-letter ISO 639-2/T code. For example, the code for Basque is 'eu'. ParlaMint corpora will use at least two languages (except for Great Britain), i.e. the language that the transcriptions are written in, which we will call the local language and English, as the meta-language, which is (also) used in the metadata.
  • Temporal, i.e. time-related information is typically stored in the when, from and to attributes of various elements. To specify a date or time as the value of these attributes, formatting according to the ISO 8601 standard should be used, e.g. 2022-04-01 for the 1st of April 2022. More information on temporal attributes is given in the Section on Temporal attributes.

3.3. Attributes of top-level elements

The Chapter on Overall corpus structure introduced the top level elements of the corpus root file and of the component files (i.e. the <teiCorpus> and <TEI> elements), but did not elaborate on their attributes; these are presented in this section.

The corpus root has three required attributes, as shown below:
             <teiCorpus xmlns="http://www.tei-c.org/ns/1.0"             xml:id="ParlaMint-FR"             xml:lang="fr"> -          
All three attributes can also be used on any other element, and are thus of special importance:
  • xmlns determines the namespace of the element, and this should always be the TEI namespace, i.e. http://www.tei-c.org/ns/1.0. Note that all lower level elements in the same file inherit this namespace, so it is not necessary (although it is not an error) for other elements to also define their namespace.
  • xml:id is an attribute form the (implicitly assumed) XML namespace, and gives the identifier for the corpus root or component. The value of an ID should be unique in the corpus as a whole and should obey format requirements as defined by W3C. For the corpus root, as well as for the components, it is required that this top level identifier is identical to the file name (without the file extension). The xml:id is a global attribute, so any element can have it. While this is not required, it is necessary for any element that is then referred to (via this same ID) by some other element, such as many elements in the <teiHeader>, as is explained in the Section on Corpus metadata. The subordinate elements in the transcription that have an ID (such as utterances and segments), are recommended to have the top level xml:id as a prefix and to indicate the element name in the ID. For example, if the top level ID is ParlaMint-GB_2021-01-06, the first utterance would have the ID ParlaMint-GB_2021-01-06-lords.u1 and the first segment ParlaMint-GB_2021-01-06-lords.seg1. The number of the element should not have leading zeros.
  • xml:lang is also a global attribute and gives the language code of the text content of the element; for the corpus root this does not (just) mean the content of its TEI header, but primarily the textual content of its XIncluded components. The convention is that language of the text content of an element is determined by the value of the first xml:lang attribute on its ancestor axis. In cases where the content is multilingual, the language code should be of the majority language. When the proportion of the languages is about equal, then the mul code for multiple languages can also be used.
A corpus component also has the same three required attributes, but additionally also the ana attribute:
+          
All three attributes can also be used on any other element, and are thus of special importance:
  • xmlns determines the namespace of the element, and this should always be the TEI namespace, i.e. http://www.tei-c.org/ns/1.0 (apart from the elements using the XInclude directive, cf. the Section on Use of XInclude). Note that lower level elements inherit the namespace of the superordinate element, unless explicitly overridden, so it is only necessary to specify the TEI namespace on the root element of a file.
  • xml:id is an attribute from the (implicitly assumed) XML namespace, and gives the identifier of the element bearing it. The value of an ID should be unique in the corpus as a whole and should obey format requirements as defined by W3C. For the corpus root, as well as for the components, it is required that this top level identifier is identical to the file name (without the file extension). The xml:id is a global attribute, so any element can have it. While this is not required, it is necessary for any element that is then referred to (via this same ID) by some other element, such as many elements in the <teiHeader>, as is explained in the Section on Corpus metadata. The subordinate elements in the transcription that have an ID (such as utterances and segments), are recommended to have the top level xml:id as a prefix and to indicate the element name in the ID. For example, if the top level ID is ParlaMint-GB_2021-01-06, the first utterance would have the ID ParlaMint-GB_2021-01-06-lords.u1 and the first segment ParlaMint-GB_2021-01-06-lords.seg1. The number of the element should not have leading zeros.
  • xml:lang is also a global attribute and gives the language code of the text content of the element; for the corpus root this means the content of its TEI header, while for corpus components this is the textual content of their TEI headers and <text> elements. The convention is that language of the text content of an element is determined by the value of the first xml:lang attribute on its ancestor axis. In cases where the content is multilingual, the language code should be of the majority language. When the proportion of the languages is about equal, then the mul code for multiple languages can also be used.
A corpus component also has the same three required attributes, but additionally also the ana attribute:
             <TEI xmlns="http://www.tei-c.org/ns/1.0"       xml:id="ParlaMint-FR_2017-07-04-E1001"       xml:lang="fr"       ana="#parla.sitting #reference"> -          
The same as for the corpus root, the component also sets the TEI namespace, and gives the language of its textual content, while its xml:id, of course, identifies the particular component. The ana attribute is a pointing attribute, and we introduce the these attributes in the next section.

3.4. Pointing attributes

The ParlaMint encoding uses pointing attributes for a number of purposes, e.g. for references to taxonomy categories, to speaker metadata, or to linguistic categories.

While a few elements have dedicated pointing attributes, there are three generally used ones. They share the characteristics that they are all used by a large number of different elements and that their value is a series of pointers, i.e. a white-space delimited sequence of references to the values of some xml:id attribute in the corpus or, in general, to an URI. The three attributes are:

  • ana serves to provide an analysis or to classify an element according to some pre-determined vocabulary. In ParlaMint the target element will typically be a category in a taxonomy, an event or date, or an organisation.
  • corresp points to items that correspond to the current element in some way, e.g. the (URL of a) media file to a page break.
  • ref provides an explicit reference to the full definition or identity for the entity being named. In ParlaMint it is used e.g. for connecting a person's affiliation with a particular organisation. The value of this attribute is often, but not always, an URL, e.g. for associating a place name with its GeoNames URL.
To illustrate, the example below gives some elements that contain one or more of these attributes:
<meeting ana="#parla.upper #parla.term #LEG.18">18 Legislatura</meeting> +          
The same as for the corpus root, the component also sets the TEI namespace, and gives the language of its textual content, while its xml:id, of course, identifies the particular component. The ana attribute is a pointing attribute, and we introduce the these attributes in the next section.

3.4. Pointing attributes

The ParlaMint encoding uses pointing attributes for a number of purposes, e.g. for references to taxonomy categories, to speaker metadata, or to linguistic categories.

While a few elements have dedicated pointing attributes, there are three generally used ones. They share the characteristics that they are all used by a large number of different elements and that their value is a series of pointers, i.e. a white-space delimited sequence of references to the values of some xml:id attribute in the corpus or, in general, to an URI. The three attributes are:

  • ana serves to provide an analysis or to classify an element according to some pre-determined vocabulary. In ParlaMint the target element will typically be a category in a taxonomy, an event, or an organisation.
  • corresp points to items that correspond to the current element in some way, e.g. the (URL of a) media file to a page break.
  • ref provides an explicit reference to the full definition or identity for the entity being named. In ParlaMint it is used e.g. for connecting a person's affiliation with a particular organisation. The value of this attribute is often, but not always, an URL, e.g. for associating a place name with its GeoNames URL.
To illustrate, the example below gives some elements that contain one or more of these attributes:
<meeting ana="#parla.upper #parla.term #LEG.18">18 Legislatura</meeting> ... <affiliation ref="#group.L-SP-PSd.Az"  role="memberana="#LEG.18from="2018-03-27"/> @@ -41,7 +41,7 @@ ... <link ana="ud-syn:det" - target="#ParlaMint-IT.seg1.2.6 #ParlaMint-IT.seg1.2.5"/>
The first example, with the <meeting> element classifies it (the definitions are given in the relevant taxonomy) as a meeting of the upper house, in the scope of a parlimentary term, specifically in the XVIII Legislative Term. The example with <affiliation> (again, the definitions are given the elements with the pointed-to ID) specifies that the (person that has this) affiliation is a member of the parliamentary group ‘Lega-Salvini Premier-Partito Sardo d'Azione’ in the scope of the XVIII Legislative Term. The <placeName> example gives the definition of Palermo in the GeoNames database via the used URL. Finally, the <link> example illustrates a Universal Dependencies determiner syntactic link between two tokens. The link uses the TEI extended pointer syntax, further explained in the Section on Prefix definitions.

It is often difficult to decide which of the attribute to use for a particular pointer, therefore examples of usage given with the relevant element should be always consulted.

3.5. Temporal attributes

ParlaMint makes a lot of use of temporal information, e.g. to determine when a session took place or the period when a certain person was an MP. As mentioned in the Section on Standard values, the ISO 8601 format should be used to specify the dates or times.

The following attributes are used to specify temporal information:

  • The when attribute is used when the temporal information refers to a point in time, typically a date, and is used e.g. to give the date when the corpus was published, or when a change in the corpus was made.
  • The from and to attributes give the starting and ending date or time of an interval, e.g. the time period the corpus covers, or the period when a person was an MP. If only one of the two attributes is present, then the assumption is that this interval extends at least to the start (if from is missing) or after the end (if to is missing) of time period that the particular ParlaMint corpus covers. Similary, if both attributes are missing, the assumption is that the interval covers the complete time period of the ParlaMint corpus.

It should be noted that, in ParlaMint, we do not support overlapping dates. For example, it is common for a term to end on a particular day, with the next term starting on the same day, and the same for a coalition and opposition. But encoding this overlap leads to problems of multivalued attributes (e.g. for a speaker to belong to a coalition and opposition on the same day), so the starting date of the next event or state should be, by convention, moved one day forward. ParlaMint thus introduces a small mistake in recording the facts but this is outweighted by the simplification in processing.

4. Corpus metadata

As mentioned, <teiCorpus> and <TEI> elements contain the obligatory <teiHeader> element, which stores the metadata to the corpus root or component. In this section we explain and give examples of the required and optional metadata that is contained in the <teiHeader>, proceeding through its various elements, and there distinguishing which parts and what content is appropriate for the corpus root, and which for a corpus component.

As a general remark, most metadata contains free text, and it is a requirement of ParlaMint that this data is given in the English language, to help researchers for other countries to understand it, and it is recommended to also give it in the local language in which the (main portion of) parliamentary transcripts is written, for a local researcher to be able to use it in their native tongue.

A ParlaMint <teiHeader> contains three obligatory elements: the file description, <fileDesc>, the encoding description, <encodingDesc>, and the profile description, <profileDesc>, and an optional revision description, <revisionDesc>:
<teiHeader>target="#ParlaMint-IT.seg1.2.6 #ParlaMint-IT.seg1.2.5"/>
The first example, with the <meeting> element classifies it (the definitions are given in the relevant taxonomy) as a meeting of the upper house, in the scope of a parlimentary term, specifically in the XVIII Legislative Term. The example with <affiliation> (again, the definitions are given the elements with the pointed-to ID) specifies that the (person that has this) affiliation is a member of the parliamentary group ‘Lega-Salvini Premier-Partito Sardo d'Azione’ in the scope of the XVIII Legislative Term. The <placeName> example gives the definition of Palermo in the GeoNames database via the used URL. Finally, the <link> example illustrates a Universal Dependencies determiner syntactic link between two tokens. The link uses the TEI extended pointer syntax, further explained in the Section on Prefix definitions.

It is often difficult to decide which of the attribute to use for a particular pointer, therefore examples of usage given with the relevant element should be always consulted.

3.5. Temporal attributes

ParlaMint makes a lot of use of temporal information, e.g. to determine when a session took place or the period when a certain person was an MP. As mentioned in the Section on Standard values, the ISO 8601 format should be used to specify the dates or times.

The following attributes are used to specify temporal information:

  • The when attribute is used when the temporal information refers to a point in time, typically a date, and is used e.g. to give the date when the corpus was published, or the date when a session was held.
  • The from and to attributes give the starting and ending date or time of an interval, e.g. the time period the corpus covers, or the period when a person was an MP. If only one of the two attributes is present, then the assumption is that this interval extends at least to the start (if from is missing) or after the end (if to is missing) of time period that the particular ParlaMint corpus covers. Similary, if both attributes are missing, the assumption is that the interval covers the complete time period of the ParlaMint corpus.

It should be noted that, in ParlaMint, we do not support overlapping dates. For example, it is common for a term to end on a particular day, with the next term starting on the same day, and the same for a coalition and opposition. But encoding this overlap leads to problems of multivalued attributes (e.g. for a speaker to belong to a coalition and opposition on the same day), so the starting date of the next event or state should be, by convention, moved one day forward. ParlaMint thus introduces a small mistake in recording the facts but this is outweighted by the simplification in processing.

4. Corpus metadata

As mentioned, <teiCorpus> and <TEI> elements contain the obligatory <teiHeader> element, which stores the metadata to the corpus root or component. In this section we explain and give examples of the required and optional metadata that is contained in the <teiHeader>, proceeding through its various elements, and there distinguishing which parts and what content is appropriate for the corpus root, and which for a corpus component.

As a general remark, most metadata contains free text, and it is a requirement of ParlaMint that this data is given in the English language, to help researchers for other countries to understand it, and it is recommended to also give it in the local language in which the (main portion of) parliamentary transcripts is written, for a local researcher to be able to use it in their native tongue.

A ParlaMint <teiHeader> contains three obligatory elements: the file description, <fileDesc>, the encoding description, <encodingDesc>, and the profile description, <profileDesc>, and an optional revision description, <revisionDesc>:
<teiHeader>  <fileDesc>...</fileDesc>  <encodingDesc>...</encodingDesc>  <profileDesc>...</profileDesc> @@ -75,7 +75,7 @@   <orgName>Slovenska raziskovalna infrastruktura CLARIN.SI</orgName>   <orgName xml:lang="en">The Slovenian research infrastructure CLARIN.SI</orgName>  </funder> -</titleStmt>
The title statement starts with two titles (one main, the other subordinate), both in English and the local language, with the appropriate language code possibly inherited from a superordinate element. They are distinguished by the value main or sub of their type attribute and the value of their xml:lang attribute.

The main title has a formulaic structure ‘<Country name> parliamentary corpus ParlaMint-<Country code> [ParlaMint]’, with an equivalent structure for the local language. Note that the corpus ‘stamp’ in square brackets can also be ‘[ParlaMint.ana]’ for the linguistically annotated version of the corpus (as explained in the Chapter on Linguistic annotation) or ‘[ParlaMint SAMPLE]’ for corpus data samples, as available on the ParlaMint GitHub repository.

The subordinate title, in contrast to the main one, is free text, and usually formed on the basis of the source of the corpus. As with the main one, it should be given in both languages.

After the titles come the specification of the particular sessions that the corpus contains, encoded as <meeting> elements: the two meeting elements in the above example state that the ParlaMint-SI corpus contains the meetings of the 7th and 8th terms of the lower house of the National Assembly of the Republic of Slovenia. The <meeting> elements can give, as the value of their n attribute, the numbers of the meetings that the corpus covers, and their text content can give a free-text description of the meetings in the local language.

The formal information on the meetings is given in the values of the corresp and ana attributes, which are pointing attributes, as already explained in the Section on Attributes of top-level elements. Here they refer to the definition of organisations further explained in the Section on Organisations and the categories of taxonomy elements, further explained in the Section on the Class declaration. The value of the corresp attribute points to the governmental body of which a particular meeting element is a meeting of (in this case the National Assembly of the Republic of Slovenia), while the ana attribute contains a space-delimited sequence of pointers: #parla.lower points to the definition of the lower house, #parla.term to the definition of a parliamentary term, and #DZ.7 to the definition of the seventh mandate.

Next come one or more responsibility statements, <respStmt>, each one containing one or more person names, <persName>, with an optional ref attribute, giving the URL, where more information about the person can be found, and the responsibility element <resp>, which specifies what responsibility the statement is about.

In a similar manner, the <funder> elements give information on the organisations which have financially contributed to the compilation of the corpus, with the names of the organisations given in the <orgName> elements.

A corpus component has a very similar title statement to the corpus root, except that certain elements specify the metadata of the component, rather than the complete corpus. The also contain some redundant metadata, in particular, the responsibility statement and the funder, as illustrated in the example below:
<titleStmt> +</titleStmt>
The title statement starts with two titles (one main, the other subordinate), both in English and the local language, with the appropriate language code possibly inherited from a superordinate element. They are distinguished by the value main or sub of their type attribute and the value of their xml:lang attribute.

The main title has a formulaic structure ‘<Country_name> parliamentary corpus ParlaMint-<Country_code> [ParlaMint]’, with an equivalent structure for the local language. Note that the corpus ‘stamp’ in square brackets can also be ‘[ParlaMint.ana]’ for the linguistically annotated version of the corpus (as explained in the Chapter on Linguistic annotation) or ‘[ParlaMint SAMPLE]’ for corpus data samples, as available on the ParlaMint GitHub repository.

The subordinate title, in contrast to the main one, is free text, and usually formed on the basis of the source of the corpus. As with the main one, it should be given in both languages.

After the titles come the specification of the particular sessions that the corpus contains, encoded as <meeting> elements: the two meeting elements in the above example state that the ParlaMint-SI corpus contains the meetings of the 7th and 8th terms of the lower house of the National Assembly of the Republic of Slovenia. The <meeting> elements can give, as the value of their n attribute, the numbers of the meetings that the corpus covers, and their text content can give a free-text description of the meetings in the local language.

The formal information on the meetings is given in the values of the corresp and ana attributes, which are pointing attributes, as already explained in the Section on Attributes of top-level elements. Here they refer to the definition of organisations further explained in the Section on Organisations and the categories of taxonomy elements, further explained in the Section on the Class declaration. The value of the corresp attribute points to the governmental body of which a particular meeting element is a meeting of (in this case the National Assembly of the Republic of Slovenia), while the ana attribute contains a space-delimited sequence of pointers: #parla.lower points to the definition of the lower house, #parla.term to the definition of a parliamentary term, and #DZ.7 to the definition of the seventh mandate.

Next come one or more responsibility statements, <respStmt>, each one containing one or more person names, <persName>, with an optional ref attribute, giving the (typically ORCID) URL, where more information about the person can be found, and the responsibility element <resp>, which specifies what responsibility the statement is about.

In a similar manner, the <funder> elements give information on the organisations which have financially contributed to the compilation of the corpus, with the names of the organisations given in the <orgName> elements.

A corpus component has a very similar title statement to the corpus root, except that certain elements specify the metadata of the component, rather than the complete corpus. The also contain some redundant metadata, in particular, the responsibility statement and the funder, as illustrated in the example below:
<titleStmt>  <title type="main">Slovenski parlamentarni korpus ParlaMint-SI, izredna seja 59 [ParlaMint]</title>  <title type="mainxml:lang="en">Slovenian parliamentary corpus ParlaMint-SI, Extraordinary Session 59 [ParlaMint]</title>  <title type="sub">Zapisi sej Državnega zbora Republike Slovenije, 7. mandat, 59. izredna seja, 13.4.2018</title> @@ -99,7 +99,7 @@  </funder> </titleStmt>
In the example it can be seen that the main title of a corpus component is simply an extension of the corpus root title, as it also gives the name of the particular meeting that the component contains, while the subordinate title is, again, free text. Both titles must be unique in the complete corpus.

The other difference is in the <meeting> elements, which here specify a particular meeting of the corpus component transcription. In the exmple above, this is an extraordinary meeting of the lower house in the seventh term of the National Assembly of the Republic of Slovenia.

4.1.2. Edition statement

ParlaMint corpora have their edition statement, <editionStmt> both in the corpus root and components. As illustrated below, the only element it contains is <edition>:
<editionStmt>  <edition>3.0</edition> -</editionStmt>
We use semantic versioning to specify the version of the corpus, i.e. giving the version number, where a new major version means substantial changes to the corpus, while the minor version is reserved for e.g. correcting errata or other minor changes. We do not use the patch number. It should be noted that - at least so far - all the ParlaMint corpora were released together, so that they are all of the same edition, i.e. have the same version number. At the time of writing, the latest version is 2.1, with the next one planned to be 3.0.

4.1.3. Extents

The <extent> element gives information on selected sizes of the complete corpus (in the corpus root) or of one corpus component, as illustrated below in the case of a corpus root extent:
<extent> +</editionStmt>
We use semantic versioning to specify the version of the corpus, i.e. giving the version number, where a new major version means substantial changes to the corpus, while the minor version is reserved for e.g. correcting errata or other minor changes. We do not use the patch number. It should be noted that - at least so far - all the ParlaMint corpora were released together, so that they are all of the same edition, i.e. have the same version number.

4.1.3. Extents

The <extent> element gives information on selected sizes of the complete corpus (in the corpus root) or of one corpus component, as illustrated below in the case of a corpus root extent:
<extent>  <measure unit="speechesquantity="75122"   xml:lang="sl">75.122 govorov</measure>  <measure unit="speechesquantity="75122" @@ -124,15 +124,15 @@   <ref target="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0        International License</ref>.</p>  </availability><date when="2021-06-11">11. 6. 2023</date> -</publicationStmt>
The <publisher> is, at least for the corpora produced in the scope of the CLARIN ParlaMint project, the CLARIN research infrastructure, and the element also gives the home page of the infrastructure. The ‘identifier number’ element, <idno>, specifies via its type and subtype attributes with fixed values URI and handle that the identifier is a handle, and contains the handle where the complete corpus corresponding to the specified version can be found. The <availability> specifiers, via its <licence> element the fixed-value CC BY 4.0 URL, and in the following paragraph gives a prose description of the licence, including its URL via the target attribute of <ref>. As usual, the textual information is given in both languages. Finally, the <date> gives the date of the release, where the when gives the date in the ISO 8601 format, while the textual content can give it according to the conventions used in the local language.

4.1.5. Source description

The source description <sourceDesc> of the corpus root encodes the original digital source of the ParlaMint corpus in the <bibl> element, as shown in the following example:
<sourceDesc><date when="2023-06-11">11. 6. 2023</date> +</publicationStmt>
The <publisher> is, at least for the corpora produced in the scope of the CLARIN ParlaMint project, the CLARIN research infrastructure, and the element also gives the home page of the infrastructure. The ‘identifier number’ element, <idno>, specifies via its type and subtype attributes with fixed values URI and handle that the identifier is a handle, and contains the handle where the complete corpus corresponding to the specified version can be found. The <availability> specifies, via its <licence> element the fixed-value CC BY 4.0 URL, and in the following paragraph gives a prose description of the licence, including its URL via the target attribute of <ref>. As usual, the textual information is given in both languages. Finally, the <date> gives the date of the release, where the when gives the date in the ISO 8601 format, while the textual content can give it according to the conventions used in the local language.

4.1.5. Source description

The source description <sourceDesc> of the corpus root encodes the digital source of the ParlaMint corpus in the <bibl> element, as shown in the following example:
<sourceDesc>  <bibl>   <title type="mainxml:lang="sl">Zapisi sej Državnega zbora Republike Slovenije</title>   <title type="mainxml:lang="en">Minutes of the National Assembly of the Republic of Slovenia</title>   <idno type="URI">https://www.dz-rs.si</idno>   <date from="2014-08-01to="2020-07-16">1.8.2014 - 16.7.2020</date>  </bibl> -</sourceDesc>
Apart from the bi-lingual <title>s, it should also give in <idno> with the fixed type as URI the government URL where the transcripts were first harvested from, while the dates of the earliest and latest transcript in the corpus are indicated by the from and to attributes of the <date> element. As usual, the values of these attributes should be according to ISO 8601, while the textual content can be formatted according to the local rules for writing dates.
For corpus components the source description is very similar to the one for the corpus root, except that the <title> can be modified to constrain the description to the exact meeting the component contains. The <date> element must, of course, specify the exact date when the meeting took place. If the transcription of the meeting is avilable on the Web, the <idno> should give this URL. Furthermore, if the audio or video of the meeting is available, this information can be given in the <recodingStmt>, as illustrated in the example below:
<sourceDesc> +</sourceDesc>
Apart from the bi-lingual <title>s, it should also give in <idno> with the fixed type as URI the government URL where the transcripts were harvested from, while the dates of the earliest and latest transcript in the corpus are indicated by the from and to attributes of the <date> element. As usual, the values of these attributes should be according to ISO 8601, while the textual content can be formatted according to the local rules for writing dates.
For corpus components the source description is very similar to the one for the corpus root, except that the <title> can be modified to constrain the description to the exact meeting the component contains. The <date> element must, of course, specify the exact date when the meeting took place. If the transcription of the meeting is avilable on the Web, the <idno> should give this URL. Furthermore, if the audio or video of the meeting is available, this information can be given in the <recodingStmt>, as illustrated in the example below:
<sourceDesc>  <bibl>   <title type="mainxml:lang="cs">Parlament České republiky, Poslanecká sněmovna</title>   <title type="mainxml:lang="en">Parliament of the Czech Republic, Chamber of Deputies</title> @@ -165,7 +165,7 @@    corpora linguistically; (3) make the corpora available for download    and through concordancers; and (4) build use cases in Political    Sciences and Digital Humanities based on the corpus data.</p> -</projectDesc>
The description above is written for the CLARIN ParlaMint project and the English language part can be used as is in the produced corpora for version 3.

4.2.2. Editorial declaration

The editorial declaration, <editorialDecl> is used only in the corpus root and contains prose descriptions of the editorial decision made in the process of compiling the corpus, along several dimensions, in particular what, if any types of <correction>, <normalization>, <quotation>, <hyphenation>, and <segmentation> was performed on the source texts of the corpus. The example below illustrates the use of these elements:
<editorialDecl> +</projectDesc>

4.2.2. Editorial declaration

The editorial declaration, <editorialDecl> is used only in the corpus root and contains prose descriptions of the editorial decision made in the process of compiling the corpus, along several dimensions, in particular what, if any types of <correction>, <normalization>, <quotation>, <hyphenation>, and <segmentation> was performed on the source texts of the corpus. The example below illustrates the use of these elements:
<editorialDecl>  <correction>   <p xml:lang="en">No correction of source texts was performed.</p>  </correction> @@ -199,17 +199,17 @@ </tagsDecl>
It should be noted that similar to the extents (as explained in the Section on Extents) the tag usage is inserted into the TEI headers in the finalisation of a corpus (cf. the Section on Validation and conversion) by a common script, so it is not necessary to compute it the process of developing a ParlaMint corpus.

4.2.4. Class declaration and taxonomies

The class declaration, <classDecl> is used only in the corpus root and contains only definitions of (most) controlled vocabularies used in ParlaMint corpora. These vocabularies, possibly hierarchically organised, are encoded using the <taxonomy> element.

The taxonomies themselves are stored in separate files, and are typically ParlaMint-wide, i.e. all corpora use the same taxonomies. The taxonomies are included in the document root with the XInclude directive, as illustrated below, for the case of the Czech corpus:
<classDecl>   <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" -      href="ParlaMint-taxonomy-parla.legislature.xml"/>    <!-- Common taxonomy on parliament legislature --> +      href="ParlaMint-taxonomy-parla.legislature.xml"/>    <!-- Common taxonomy of parliament legislature -->   <xi:include xmlns:xi="http://www.w3.org/2001/XInclude"       href="ParlaMint-CZ-taxonomy-meeting.parts.xml"/>     <!-- CZ-specific taxonomy with additional categories for meetings -->   <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" -      href="ParlaMint-taxonomy-speaker_types.xml"/>        <!-- Common taxonomy on types of speakers --> +      href="ParlaMint-taxonomy-speaker_types.xml"/>        <!-- Common taxonomy of types of speakers -->   <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" -      href="ParlaMint-taxonomy-subcorpus.xml"/>            <!-- Common taxonomy on predeterimentd subcorpora --> +      href="ParlaMint-taxonomy-subcorpus.xml"/>            <!-- Common taxonomy of predeterimentd subcorpora -->   <xi:include xmlns:xi="http://www.w3.org/2001/XInclude"  -      href="ParlaMint-taxonomy-politicalOrientation.xml"/> <!-- Common taxonomy on L-R political orientations of pol. parties --> +      href="ParlaMint-taxonomy-politicalOrientation.xml"/> <!-- Common taxonomy on political orientations of political parties -->   <xi:include xmlns:xi="http://www.w3.org/2001/XInclude"  -      href="ParlaMint-taxonomy-CHES.xml"/>                 <!-- Common taxonomy on CHES variables of pol. parties --> +      href="ParlaMint-taxonomy-CHES.xml"/>                 <!-- Common taxonomy of CHES variables of political parties --> </classDecl>
As can be seen, three of the taxonomies are general ParlaMint taxonomies, while two are corpus specific, and are distinguished by including the country code CZ (followed by hyphen) into the filename.
To illustrate the structure of a taxonomy element, we give below the simplest common taxonomy (and include only the descriptions in English), which contains the categories that define the three subcorpora of a ParlaMint corpus:
<classDecl> ... <taxonomy xml:id="ParlaMint-taxonomy-subcorpus"   xml:lang="mul"> @@ -240,7 +240,7 @@   <catDesc xml:lang="en">    <term>Agenda</term>: topic discussed during sitting</catDesc>  </category> -</taxonomy>

ParlaMint requires several taxonomies to be defined in the class declaration of the corpus root (as well as a additionaly ones for the linguistically annotated corpus, as further described in the Section on Linguistic metadata). As mentioned, these taxonomies are defined globally and available as part of the data on the ParlaMint GitHub repository, and there is a special procedure modifying them, in particular on how to insert translations of a new language.

The five obligatory taxonomies are:

  • The subcorpus taxonomy, already given in the example
  • The taxonomy of speaker types, which distinguishes e.g. the chair of a meeting, ordinary speakers, and guest speakers.
  • The legislature taxonomy, which gives the possible organisations of a parliament, and is by far the most complex one.
  • The political orientation taxonomy, which gives the values of the (mostly) left-right political orientation of political parties and parliamentary groups.
  • The CHES variables taxonomy, which gives the variables of the Chapel Hill Survey on political parties (cf. the Section on Political parties and parliamentary groups).
  • The CAP topics taxonomy, which gives major topic labels of the Comparative Agendas Project.

Furhtermore, there are several obligatory taxonomies which pertain to the linguistically analysed version of the corpus only, cf. the Section on Linguistic taxonomies.

4.3. Profile description

The profile description, <profileDesc> is the third main division of the metadata provided by the TEI header. It contains a description of non-bibliographic aspects of the corpus, for example the list of speakers with their metadata. For the corpus root, it contains four elements, of which only the first, the <settingDesc> is used in corpus components. The elements are listed below:
<profileDesc> +</taxonomy>

ParlaMint requires several taxonomies to be defined in the class declaration of the corpus root (as well as additional ones for the linguistically annotated corpus, as further described in the Section on Linguistic metadata). As mentioned, these taxonomies are defined globally and available as part of the data on the ParlaMint GitHub repository, and there is a special procedure modifying them, in particular on how to insert translations of a new language.

The five obligatory taxonomies are:

  • The subcorpus taxonomy, already given in the example
  • The taxonomy of speaker types, which distinguishes e.g. the chair of a meeting, ordinary speakers, and guest speakers.
  • The legislature taxonomy, which gives the possible organisations of a parliament, and is by far the most complex one.
  • The political orientation taxonomy, which gives the values of the (mostly) left-right political orientation of political parties and parliamentary groups.
  • The CHES variables taxonomy, which gives the variables of the Chapel Hill Survey on political parties (cf. the Section on Political parties and parliamentary groups).
  • The CAP topics taxonomy, which gives major topic labels of the Comparative Agendas Project.

Furhtermore, there are several obligatory taxonomies which pertain to the linguistically analysed version of the corpus only, cf. the Section on Linguistic taxonomies.

4.3. Profile description

The profile description, <profileDesc> is the third main division of the metadata provided by the TEI header. It contains a description of non-bibliographic aspects of the corpus, for example, the list of speakers with their metadata. For the corpus root, it contains four elements, of which only the first, the <settingDesc> is used in corpus components. The elements are listed below:
<profileDesc>  <settingDesc>...</settingDesc>  <textClass>...</textClass>  <particDesc>...</particDesc> @@ -292,7 +292,7 @@   <name>Tomaž Erjavec</name>: Finalized encoding.</change>  <change when="2021-05-28">   <name>Tomaž Erjavec</name>: Built corpus.</change> -</revisionDesc>
The revision description consists of a series of <change> elements, with the attribute when giving the date of the change, and the content containing the <name> of the person responsible for the change, and a free-text description of the change.

5. Speakers and their organisations

ParlaMint places considerable emphasis of including in the corpora significant information about the persons giving the speeches contained in the transcriptions. This is why, even though this information is encoded in the <particDesc> element of the <teiHeader> of the corpus root (cf. the Chapter on Participant description) we treat it here in a separate Chapter. Below we first discuss the information on persons, including how they are affiliated with (political) organisation, and then explain the encoding of these organisations.

5.1. Speakers

The information on speakers is given in the <listPerson> element of the <particDesc> element (cf. the Section on Participant description). This element contains the series of <person> elements, each of which gives information on an individual speaker, as the example below illustrates:
<listPerson> +</revisionDesc>
The revision description consists of a series of <change> elements, with the attribute when giving the date of the change, and the content containing the <name> of the person responsible for the change, and a free-text description of the change. Note that the <change> follow reverse chronological order, i.e. the most recent changes are at the top.

5. Speakers and their organisations

ParlaMint places considerable emphasis of including in the corpora significant information about the persons giving the speeches contained in the transcriptions. This is why, even though this information is encoded in the <particDesc> element of the <teiHeader> of the corpus root (cf. the Chapter on Participant description) we treat it here in a separate Chapter. Below we first discuss the information on persons, including how they are affiliated with (political) organisation, and then explain the encoding of these organisations.

5.1. Speakers

The information on speakers is given in the <listPerson> element of the <particDesc> element (cf. the Section on Participant description). This element contains the series of <person> elements, each of which gives information on an individual speaker, as the example below illustrates:
<listPerson>  <person xml:id="AccettoMatej">   <persName>    <surname>Accetto</surname> @@ -585,7 +585,7 @@ <u who="#JeremyCorbyn"  ana="#regular #interruptingxml:id="GB001.8.4">Traitor!</u> <u who="#BorisJohnsonana="#regular" - xml:id="GB001.8.5prev="#GB001.8.3">Because England does not want any dealings with the European Union.</u>
As can be seen, the split is indicated by the use of the next attribute on the first part of the split utterance and by the prev attribute of the next part of the split utterance, while the fact that an utterance interrupts another one is signaled by the addition of #interrupting to the ana attribute. The values of the next and prev attributes are pointers to the next of previous identifiers of the appropriate part of the split utterance.4.

In the example the speaker of the interrupting speech has also been identified and marked in the who attribute; in cases where this is not possible, this attribute can be omitted. As mentioned this speaker should also have the value #interrupting in their ana attribute; this value comes from the appropriate category of the ParlaMint speaker type taxonomy. In case the speaker is identified, and their status can be determined, ana should also contain the type proper of the speaker, i.e. whether they are #chair, #regular or #guest speaker is assumed.

7. Linguistic annotation

This section introduces the ParlaMint linguistic annotation. An important note is that a linguistically annotated ParlaMint corpus is stored separately from its base (or plain-text) version, i.e. the version that has been discussed in the preceding sections. The encoding of the linguistically annotated version differs from the plain-text one in the following:

  • All the corpus root and components file names the extension .ana.xml. For example, if the plain-text root has the file name ParlaMint-CZ.xml, the linguistically annotated one should be ParlaMint-CZ.ana.xml, or ParlaMint-CZ_2016-04-13.xml and ParlaMint-CZ_2016-04-13.ana.xml
  • Because the file ID (i.e. the value of the top level element attribute xml:id, as explained in the Section on Attributes of top-level elements) should be the same as the file name (cf. the Section on File names and directory structure), the previous point also means that the linguistically annotated files should have the top level ID suffixed with .ana, e.g. <teiCorpus xml:id="ParlaMint-CZ.ana">
  • The corpus stamp in the main title of the corpus root or components (cf. the Section on Title statement) which is [ParlaMint] in the plain-text version, should be [ParlaMint.ana] for the linguistically annotated version.
  • All the plain text of the utterance segments (i.e. the text immediately contained by the <seg> elements) of the plain-text version should be linguistically annotated on the specified levels, as is further explained in the following Section on Linguistic markup.
  • The linguistically annotated version of the corpus should also have some added metadata in the TEI header of the corpus root, which is detailed in the following Section on Metadata for linguistic annotation.

7.1. Linguistic markup

Linguistic annotation is added only to the text content of <seg> elements inside the speeches, i.e <u> elements. For this text, ParlaMint requires the following additional markup to be present:

  • tokens: what is a word, and what is punctuation, with preserved information on inter-token spaces;
  • sentences: what is a sentence;
  • lemmas: what is the base form of each word;
  • Universal Dependencies (UD) part-of-speech and morphological features, and, optionally, part-of-speech tags from a different (local) tagset;
  • named entities (NE): what is a name, categorised at least into the standard four NE classes;
  • the UD dependency syntactic parse of the sentences;
  • USAS semantic annotations on words and phrases but only for the machine translated corpora.

Below, we explain the encoding of each of these levels.

7.1.1. Word-level annotation

Basic linguistic annotation comprises tokenisation, sentence segmentation, part-of-speech tagging and lemmatisation, and this mark-up is illustrated in the example below:
<s>xml:id="GB001.8.5prev="#GB001.8.3">Because England does not want any dealings with the European Union.</u>
As can be seen, the split is indicated by the use of the next attribute on the first part of the split utterance and by the prev attribute of the next part of the split utterance, while the fact that an utterance interrupts another one is signaled by the addition of #interrupting to the ana attribute. The values of the next and prev attributes are pointers to the next of previous identifiers of the appropriate part of the split utterance.4.

In the example the speaker of the interrupting speech has also been identified and marked in the who attribute; in cases where this is not possible, this attribute can be omitted. As mentioned this speaker should also have the value #interrupting in their ana attribute; this value comes from the appropriate category of the ParlaMint speaker type taxonomy. In case the speaker is identified, and their status can be determined, ana should also contain the type proper of the speaker, i.e. whether they are #chair, #regular or #guest speaker is assumed.

7. Linguistic annotation

This section introduces the ParlaMint linguistic annotation. An important note is that a linguistically annotated ParlaMint corpus is stored separately from its base (or plain-text) version, i.e. the version that has been discussed in the preceding sections. The encoding of the linguistically annotated version differs from the plain-text one in the following:

  • All the corpus root and component file names have the extension .ana.xml. For example, if the TEI plain-text root has the file name ParlaMint-CZ.xml, the linguistically annotated one should be ParlaMint-CZ.ana.xml, and if the component plain-text files is ParlaMint-CZ_2016-04-13.xml the linguistically annotated one is ParlaMint-CZ_2016-04-13.ana.xml.
  • Because the file ID (i.e. the value of the top level element attribute xml:id, as explained in the Section on Attributes of top-level elements) should be the same as the file name (cf. the Section on File names and directory structure), the previous point also means that the linguistically annotated files should have the top level ID suffixed with .ana, e.g. <teiCorpus xml:id="ParlaMint-CZ.ana">.
  • The corpus stamp in the main title of the corpus root or components (cf. the Section on Title statement) which is [ParlaMint] in the plain-text version, should be [ParlaMint.ana] for the linguistically annotated version.
  • All the plain text of the utterance segments (i.e. the text immediately contained by the <seg> elements) of the plain-text version should be linguistically annotated on the specified levels, as is further explained in the following Section on Linguistic markup.
  • The linguistically annotated version of the corpus should also have some added metadata in the TEI header of the corpus root, which is detailed in the Section on Metadata for linguistic annotation.

7.1. Linguistic markup

Linguistic annotation is added only to the immediate text content of <seg> elements inside the speeches, i.e <u> elements. For this text, ParlaMint requires the following additional markup to be present:

  • tokens: what is a word, and what is punctuation, with preserved information on inter-token spaces;
  • sentences: what is a sentence;
  • lemmas: the base form of each word;
  • Universal Dependencies (UD) part-of-speech and morphological features, and, optionally, part-of-speech tags from a different (local) tagset;
  • named entities (NE): a name, categorised into the standard four NE classes;
  • the UD dependency syntactic parse of the sentences;
  • USAS semantic annotations on words and phrases but only for the machine translated corpora.

Below, we explain the encoding of each of these levels.

7.1.1. Word-level annotation

Basic linguistic annotation comprises tokenisation, sentence segmentation, part-of-speech tagging and lemmatisation, and this mark-up is illustrated in the example below:
<s>  <w msd="UPosTag=DET|Case=Gen|Gender=Neut|Number=Sing|PronType=Dem"   lemma="ta">Tega</w>  <w msd="UPosTag=PRON|PronType=Prs|Reflex=Yes|Variant=Short" @@ -651,7 +651,7 @@  <w join="rightlemma="Hall"   msd="UPosTag=PROPN|Number=Sing">Hall</w> </name> -<w lemma=",msd="UPosTag=PUNCT">,</w> +<pc msd="UPosTag=PUNCT">,</pc> ...
ParlaMint also supports more complex NE annotation schemes, such as the one used for Czech data, which introduces very detailed NE types and also allows for nested named entities. The example below gives such a case, where the top level and ParlaMint-compatible person name also contains nested names:
<name type="PERana="ne:p">  <name ana="ne:pf"> @@ -695,7 +695,7 @@   ana="sem:Z1">Prime</w>  <w lemma="Ministerfunction="Z1mf,Z3c"   ana="sem:Z1">Minister</w> -</phr>
The first thing to note is that all the tokens inside a MWE receive identical semantic markup as its encompassing MWE <phr> element. Second, the function attribute gives the USAS tags exactly as output by the tool used for semantic tagging, which includes not only the tag computed to be the most appropriate for the given context, but also all the other lexically possible tags, with the comma being the separator. The first tag is then used to compute the values of the ana attribute. Here the conjunctive USAS tag (here the delimiter is the slash) is transformed into a series of references to the USAS taxomomy (cf. the Section on Linguistic taxonomies) while also removing tag qualifiers. For example, the tag G1.1/S2mf is transformed into sem:G1.1 sem:S2 i.e. the mf (for male and female) qualifiers are removed from S2mf. Note also that the prefix sem (cf. the Section on Prefix definitions) is used for pointing into the USAS taxonomy. With this set-up it is possible to encode the exact USAS tags as well as their ParlaMint categories, which give the not only the tag, but also its gloss.

7.2. Metadata for linguistic annotation

What kind of metadata a plain-text ParlaMint corpus should contain was explained in the Section on Corpus metadata and in this section we detail what additions must be made to the metadata for the linguistically annotated version. Note that the changes for this version have been already explained at the start of this Chapter. In short, there are three additional parts that should be added to the <teiHeader> of the corpus root, namely a description of the tool(s) used to linguistically annotate the corpus, two additional taxonomies (one for named entities, and one for UD syntactic relations) and the definition of the prefix expansions for UD syntactic relations. These descriptions should also serve as the point of departure for those that want to introduce their own prefixes and taxonomies for defining additional and corpus-specific part-of-speech tagging schemes or named entity classes.

7.2.1. Application information for linguistic processing

As the linguistic analysis of a ParlaMint will be performed by a tool, the information on which tool (or tools) have been used should be documented in the corpus root TEI header. This information is encoded in the <appInfo> element of the <encodingDesc>, as shown in the example below:
<appInfo> +</phr>
The first thing to note is that all the tokens inside a MWE receive identical semantic markup as its encompassing MWE <phr> element. Second, the function attribute gives the USAS tags exactly as output by the tool used for semantic tagging, which includes not only the tag computed to be the most appropriate for the given context, but also all the other lexically possible tags, with the comma being the separator. The first tag is then used to compute the values of the ana attribute. Here the conjunctive USAS tag (here the delimiter is the slash) is transformed into a series of references to the USAS taxomomy (cf. the Section on Linguistic taxonomies) while also removing tag qualifiers. For example, the tag G1.1/S2mf is transformed into sem:G1.1 sem:S2 i.e. the mf (for male and female) qualifiers are removed from S2mf. Note also that the prefix sem (cf. the Section on Prefix definitions) is used for pointing into the USAS taxonomy. With this set-up it is possible to encode the exact USAS tags as well as their ParlaMint categories, which give the not only the tag, but also its gloss.

7.2. Metadata for linguistic annotation

What kind of metadata a plain-text ParlaMint corpus should contain was explained in the Section on Corpus metadata and in this section we detail what additions must be made to the metadata for the linguistically annotated version. Note that the changes for this version have been already explained at the start of this Chapter. In short, there are three additional parts that should be added to the <teiHeader> of the corpus root, namely a description of the tool(s) used to linguistically annotate the corpus, two additional taxonomies (one for named entities, and one for UD syntactic relations) and the definition of the prefix expansions for UD syntactic relations. These descriptions should also serve as the point of departure for those that want to introduce their own prefixes and taxonomies for defining additional and corpus-specific part-of-speech tagging schemes or named entity classes.

7.2.1. Application information for linguistic processing

As the linguistic analysis of a ParlaMint corpus will be performed by a tool, the information on which tool (or tools) have been used should be documented in the corpus root TEI header. This information is encoded in the <appInfo> element of the <encodingDesc>, as shown in the example below:
<appInfo>  <application version="1.0ident="classla">   <label>CLASSLA</label>   <desc xml:lang="en">Linguistic processing performed with with CLASSLA trained for @@ -806,16 +806,16 @@  </category> ... -</taxonomy>
The current USAS taxonomy covers a subset of all USAS semantic tags and was derived from the official list of USAS semantic subcategories. A category in the taxonomy can contain up to one positive (USAS = '+', taxonomy = 'p') or negative (USAS = '-', taxonomy = 'n') modifier. Other USAS modifiers (i.e. regex [mfnci%@]) are not retained. The taxonomy includes 455 categories, each with its USAS code and gloss.

7.2.3. Prefix definitions

Pointing attributes, such as ana, take as their value a series of references to the value of xml:id elements in an XML document. If this is the same document, then the reference to the ID is the hash character, # prefixed to the particular ID, e.g. #parla.uni, and if they are in another XML document, then the hash is prefixed with the URL of the document, e.g. https://nl.ijs.si/ME/V6/msd/tables/msd-fslib2-sl.xml#Vmpr1p.

Because the complete URL tends to be long, which is especially inconvenient when such references are given to every token in a corpus, TEI introduces the so called Extended pointer syntax, whereby the reference to an ID can be given in the form of a prefix, which is separated by a colon from the local part of the ID reference, and the value of this prefix is determined via the <prefixDef> element in the <profileDesc> of the TEI header.

ParlaMint uses this mechanism for all linguistic annotations with a closed vocabulary, in particular for the Universal Dependencies syntactic relations, for the optional and corpus-specific analytical part-of-speech tags (c.f. the Sections on Syntactic parses and Word-level annotation), and for semantic annotation in the machine translated corpora (c.f. the Section on Semantic annotation). The example below illustrates the prefix definitions for the obligatory UD syntactic relations and for the optional MULTEXT-East tags:
<listPrefixDef> +</taxonomy>
The current USAS taxonomy covers a subset of all USAS semantic tags and was derived from the official list of USAS semantic subcategories. A category in the taxonomy can contain up to one positive (USAS = '+', taxonomy = 'p') or negative (USAS = '-', taxonomy = 'n') modifier. Other USAS modifiers (i.e. regex [mfnci%@]) are not retained. The taxonomy includes 455 categories, each with its USAS code and gloss.

7.2.3. Prefix definitions

Pointing attributes, such as ana, take as their value a reference or space-delimited series of references to a URL and/or the value of xml:id elements. If the reference is to an ID, then it is prefixed the hash character, #, e.g. #parla.uni, and if they are to an ID in another XML document, then the hash follows the URL of the document, e.g. https://nl.ijs.si/ME/V6/msd/tables/msd-fslib2-sl.xml#Vmpr1p.

Because complete URLs tend to be long, especially inconvenient when such references are given to every token in a corpus, TEI introduces the so called Abbreviated Pointers, whereby a reference can be given in the form of a prefix, which is separated by a colon from the local part of the ID reference, and the value of this prefix is determined via the <prefixDef> element in the <encodingDesc> of the TEI header.

ParlaMint uses this mechanism for all linguistic annotations with a closed vocabulary, in particular for the Universal Dependencies syntactic relations, for the optional and corpus-specific analytical part-of-speech tags (c.f. the Sections on Syntactic parses and Word-level annotation), and for semantic annotation in the machine translated corpora (c.f. the Section on Semantic annotation). The example below illustrates the prefix definitions for the obligatory UD syntactic relations and for the optional MULTEXT-East tags:
<listPrefixDef>  <prefixDef ident="ud-syn"   matchPattern="(.+)replacementPattern="#$1">   <p xml:lang="en">Private URIs with this prefix point to elements giving their name. In this document they are simply local references into the UD-SYN taxonomy categories in the corpus root TEI header.</p>  </prefixDef>  <prefixDef ident="mtematchPattern="(.+)" -  replacementPattern="http://nl.ijs.si/ME/V6/msd/tables/msd-fslib-sl.xml#$1"> +  replacementPattern="https://nl.ijs.si/ME/V6/msd/tables/msd-fslib-sl.xml#$1">   <p xml:lang="en">Private URIs with this prefix point to feature-structure elements defining the Slovenian MULTEXT-East Version 6 MSDs.</p>  </prefixDef> -</listPrefixDef>
The specialised element for listing prefix definitions, <listPrefixDef> gives a series of prefix definitions, i.e. <prefixDef> elements. Each prefix definition defines its prefix as the value of the ident attribute, and then specifies a regular expression that matches the part of the ID reference after the prefix in its matchPattern attribute, and its substitution as the value of the replacementPattern attribute. The first prefix definition thus defines the ud-syn prefix, so for any ID reference with this prefix, e.g. ud-syn:acl_relcl, the part after the prefix (acl_relcl) should be matched against (.+) and the result being the matched part (here the entire relation acl_relcl) substituted by #$1, i.e. by the hash character followed by the original value, so that ud-syn:acl_relcl gives #acl_relcl. This substitution is of course trivial, and hardly necessary, but was implemented so that all fixed-vocabulary linguistic analyses have the same treatment.

More to the point is the second example, where very short ID references, such as mte:Vmpr1p are transformed to https://nl.ijs.si/ME/V6/msd/tables/msd-fslib2-sl.xml#Vmpr1p, as already explained in the Section on Word-level annotation.

Finally, each prefix definition also contains a possibly bi-lingual paragraph explaining the definition.

8. Translations of corpora

The ParlaMint machine translated corpora are encoded simliarly to corpora in their source language, i.e. they have an identically structured corpus root and components. The most obvious differences are the following:

  • The filenames are extended with the target language code, as described in the Section on Filenames, e.g. the linguistically analysed English translation of the Latvian corpus would have the root filename ParlaMint-LV-en.ana.xml and a component filename could be ParlaMint-LV_2014-11-04.ana.xml.
  • The top-level xml:id of the corpus root and of components must also be changed accordingly to e.g. ParlaMint-LV-en.ana or ParlaMint-LV_2014-11-04.ana.
  • The stamp in the main titles should be changed too by adding the target langauge suffix, e.g. to [ParlaMint-en.ana].
  • The top level xml:lang in the root and components should be set to the target language, e.g. xml:id="en".
The structure of the <text>, inluding the transcriber comments and the linguistic analysis is encoded the same as as for the corpora in the source language. The two differences are that the aligned elements are linked to their corresponding elements in the source corpus, and, to simplify processing, the transcriber comments are moved outside sentences in the (rare) cases where they appeared inside them in the original language corpus. In the current machine translated corpora the alignment are given to utterances, segments, sentences, and transcriber comments, which have, furthermore, always a 1-1 mapping to the corresponding source element. Therefore the alignment is trivial, simply specifiying the same xml:id value of the element the source corpus, as illustrated in the following example:
<div type="debateSectionxml:lang="en"> +</listPrefixDef>
The specialised element for listing prefix definitions, <listPrefixDef> gives a series of prefix definitions, i.e. <prefixDef> elements. Each prefix definition defines its prefix as the value of the ident attribute, and then specifies a regular expression that matches the part of the reference after the prefix in its matchPattern attribute, and its substitution as the value of the replacementPattern attribute. The first prefix definition thus defines the ud-syn prefix, so for any reference with this prefix, e.g. ud-syn:acl_relcl, the part after the prefix (acl_relcl) should be matched against (.+) and the result being the matched part (here the entire relation acl_relcl) substituted by #$1, i.e. by the hash character followed by the original value, so that ud-syn:acl_relcl gives #acl_relcl. This substitution is of course trivial, and hardly necessary, but was implemented so that all fixed-vocabulary linguistic analyses have the same treatment.

More to the point is the second example, where very short references, such as mte:Vmpr1p are transformed to https://nl.ijs.si/ME/V6/msd/tables/msd-fslib-sl#Vmpr1p, as already explained in the Section on Word-level annotation.

Finally, each prefix definition also contains a possibly bi-lingual paragraph explaining the definition.

8. Translations of corpora

The ParlaMint machine translated corpora are encoded simliarly to corpora in their source language, i.e. they have an identically structured corpus root and components. The most obvious differences are the following:

  • The filenames are extended with the target language code, as described in the Section on Filenames, e.g. the linguistically analysed English translation of the Latvian corpus would have the root filename ParlaMint-LV-en.ana.xml and a component filename could be ParlaMint-LV_2014-11-04.ana.xml.
  • The top-level xml:id of the corpus root and of components must also be changed accordingly to e.g. ParlaMint-LV-en.ana or ParlaMint-LV_2014-11-04.ana.
  • The stamp in the main titles should be changed too by adding the target langauge suffix, e.g. to [ParlaMint-en.ana].
  • The top level xml:lang in the root and components should be set to the target language, e.g. xml:id="en".
The structure of the <text>, inluding the transcriber comments and the linguistic analysis is encoded the same as as for the corpora in the source language. The two differences are that the aligned elements are linked to their corresponding elements in the source corpus, and, to simplify processing, the transcriber comments are moved outside sentences in the (rare) cases where they appeared inside them in the original language corpus. In the current machine translated corpora the alignment are given to utterances, segments, sentences, and transcriber comments, which have, furthermore, always a 1-1 mapping to the corresponding source element. Therefore the alignment is trivial, simply specifiying the same xml:id value of the element the source corpus, as illustrated in the following example:
<div type="debateSectionxml:lang="en">  <note type="speaker"   xml:id="ParlaMint-LV_2019-01-31-PT13-516.ana.note1"   corresp="mt-src:ParlaMint-LV_2019-01-31-PT13-516.ana.note1">Head of the sitting.</note> @@ -851,18 +851,18 @@    (<ref target="https://github.com/UKPLab/EasyNMT">https://github.com/UKPLab/EasyNMT</ref>)    with OPUS-MT model bat    (<ref target="https://github.com/Helsinki-NLP/Opus-MT">https://github.com/Helsinki-NLP/Opus-MT</ref>)</desc> -</application>
This element should be given in the corpus root, together with all the other information on applications inside the application information (<appInfo>) element.

9. Validation and conversion

The chapter explains how to validate and finalise a ParlaMint corpus, and introduces scripts for converting a ParlaMint corpus to other, derived formats.

9.1. Validating ParlaMint corpora

The XML structure of ParlaMint corpora can be validated via RelaxNG schemas, which exist in two versions, one that was produced as a customisation of the TEI Guidelines, and a set of schemas that were made from scratch for ParlaMint.

The TEI customisation is written as a TEI ODD document, which is, in fact, the XML version of this document, and is available in the TEI/ directory of the ParlaMint GitHub repository. The XML contains not only the prose guidelines, but also the formal specification of the TEI schema, which is given in the Appendix A. In the XML it contains the formal schema specification, while in the on-line version this is converted to a reference to all the elements, attributes and classes used in ParlaMint corpora. The ODD document is not immediately useful for XML validation, but has to be converted with TEI XSLT stylesheets first in order to obtain a RelaxNG schema, and this schema is also available in the same directory under the name of ParlaMint.rng (in RelaxNG XML syntax) and ParlaMint.rnc (in RelaxNG compact syntax). This schema should be used to check that ParlaMint component files validate against TEI.

However, it is difficult to constrain a TEI ODD-derived XML schema to allow only the kinds of nestings and attributes that should appear in a ParlaMint corpus, so this schema allows (and lists Appendix A) nesting of elements, as well as attributes that are in fact forbidden in ParlaMint corpora.

For this reason, we have also developed a set of RelaxNG schemas from scratch, which do allow only those elements, attributes and content models that are in fact valid for a ParlaMint corpus. There are all together four such schemas, one for a "plain-text" corpus root, one for its corpus components, one for the linguistically annotated corpus root, and one for its components. These schemas can be found in the Schema/ directory of the ParlaMint GitHub repository, with the README file giving instructions on how to use them.

Validating with XML schemas checks the formal structure of XML files but is less successful in validating other aspects of conformance, such as the textual content or linking of pointer attributes. For this reason, we have also developed an XSLT script that assumes a schema-validated ParlaMint file on its input, and checks various other aspects of conformance. These validation scripts can be found in the Scripts/ directory of the ParlaMint GitHub repository, with the README file listing them.

It should be noted that it is not necessary to run the validation scripts directly, as the validation can be performed by the main Makefile of the project. The Makefile is self-documenting, i.e. to see how to use it, please run make help in the top level directory of the ParlaMint project.

While each contributor of a corpus should validate their files with the ParlaMint schemas and validation script, there also exist further stages of validation, which are also applied to ParlaMint corpora:

  • The corpora are converted to derived formats, in particular, the linguistically annotated version of the corpus to CoNLL-U and to the so called vertical format for CQP-type concordancers. The Universal Dependencies project provides a program for validating the formatting and linguistic analyses in CoNLL-U files, and this validation is used on the CoNLL-U files derived from their XML source, up to level 2 conformance. The vertical files, on the other hand, are first compiled with manatee (the back end of (no)Sketch Engine) and this compilation can also expose various errors.
  • The last stage in validation is ‘human validation’ where e.g. simply looking at various produced metadata files or at the concordances of a corpus exposes errors.

9.2. Finalisation of corpora

While the vast majority of converting source encodings into the ParlaMint corpus format is left to the compilers of a corpus, there are a few metadata elements that can be produced by a common script on the basis of nearly finished corpora, which then results in the final version of the corpus for a particular release. This includes setting the date, edition and handle under which the corpus will be distributed, and also calculating the size of the corpus (cf. the Sections on Extents and on Tags declaration). The script for finalisation can be found in the Scripts/ directory of the ParlaMint GitHub repository and the README file briefly explains its function; more comments can be found in the script itself.

9.3. Conversions

A TEI encoded document is, in general, not meant to be used directly by software programs, rather, it serves as an interchange and storage format. The ParlaMint project has produced various scripts to down-convert the XML encoded corpora to other formats and they can be found in the Scripts/ directory of the ParlaMint GitHub repository, with the README file listing them and explaining their function. In short, the scripts convert the ParlaMint XML to plain text, to CoNLL-U, and to vertical format. There is also a script that takes a ParlaMint corpus and makes from it a sample for inclusion to the ParlaMint GitHub repository.

10. Contributing to ParlaMint

The ParlaMint GitHub repository contains these guidelines, the ParlaMint XML schemas, the scripts used to validate, finalise and convert the ParlaMint TEI XML corpora to derived formats, and samples of the ParlaMint corpora. There are four main branches in the repository:

  • main is the default branch used for the synchronisation of other branches. It is also used for releasing sample files that correspond to published corpora.
  • data serves as a pushing place for new sample files in ./Data/ParlaMint-XX directories.
  • devel: development of scripts and documentation.

The validation procedure for corpora is explained in the Section on Validating ParlaMint corpora, while the technical aspects of contributing corpora is further explained in the CONTRIBUTING file of the repository.

11. Acknowledgements

The work on these recommendations was funded by the CLARIN Research Infrastructure for Language Resources and Tools.

Appendix A Formal specification

Appendix A.1 Elements

Appendix A.1.1 <TEI>

<TEI> (TEI document) contains a single TEI-conformant document, combining a single TEI header with one or more members of the model.resource class. Multiple <TEI> elements may be combined within a <TEI> (or <teiCorpus>) element. [4. Default Text Structure 16.1. Varieties of Composite Text]
Moduletextstructure — Formal specification
Attributes
xml:id
StatusRequired
DatatypeID
xml:lang
StatusRequired
Datatypeteidata.language
ana
StatusRequired
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Contained by
core: teiCorpus
May contain
header: teiHeader
textstructure: text
Note

As with all elements in the TEI scheme (except <egXML>) this element is in the TEI namespace (see 5.7.2. Namespaces). Thus, when it is used as the outermost element of a TEI document, it is necessary to specify the TEI namespace on it. This is customarily achieved by including http://www.tei-c.org/ns/1.0 as the value of the XML namespace declaration (xmlns), without indicating a prefix, and then not using a prefix on TEI elements in the rest of the document. For example: <TEI version="4.8.1" xml:lang="it" xmlns="http://www.tei-c.org/ns/1.0">.

ExampleExample of ParlaMint corpus component:
<TEI xml:id="ParlaMint-GB_2015-01-06-commons" +</application>
This element should be given in the corpus root, together with all the other information on applications inside the application information (<appInfo>) element.

9. Validation and conversion

The chapter explains how to validate and finalise a ParlaMint corpus, and introduces scripts for converting a ParlaMint corpus to other, derived formats.

9.1. Validating ParlaMint corpora

The XML structure of ParlaMint corpora can be validated via RelaxNG schemas, which exist in two versions, one that was produced as a customisation of the TEI Guidelines, and a set of schemas that were made from scratch for ParlaMint.

The TEI customisation is written as a TEI ODD document, which is, in fact, the XML version of this document, and is available in the TEI/ directory of the ParlaMint GitHub repository. The XML contains not only the prose guidelines, but also the formal specification of the TEI schema, which is given in the Appendix A. In the XML it contains the formal schema specification, while in the on-line version this is converted to a reference to all the elements, attributes and classes used in ParlaMint corpora. The ODD document is not immediately useful for XML validation, but has to be converted with TEI XSLT stylesheets first in order to obtain a RelaxNG schema, and this schema is also available in the same directory under the name of ParlaMint.rng (in RelaxNG XML syntax) and ParlaMint.rnc (in RelaxNG compact syntax). This schema should be used to check that ParlaMint component files validate against TEI.

However, it is difficult to constrain a TEI ODD-derived XML schema to allow only the kinds of nestings and attributes that should appear in a ParlaMint corpus, so this schema allows (and lists Appendix A) nesting of elements, as well as attributes that are in fact forbidden in ParlaMint corpora.

For this reason, we have also developed a set of RelaxNG schemas from scratch, which do allow only those elements, attributes and content models that are in fact valid for a ParlaMint corpus. There are all together four such schemas, one for a "plain-text" corpus root, one for its corpus components, one for the linguistically annotated corpus root, and one for its components. These schemas can be found in the Schema/ directory of the ParlaMint GitHub repository, with the README file giving instructions on how to use them.

Validating with XML schemas checks the formal structure of XML files but is less successful in validating other aspects of conformance, such as the textual content or linking of pointer attributes. For this reason, we have also developed an XSLT script that assumes a schema-validated ParlaMint file on its input, and checks various other aspects of conformance. These validation scripts can be found in the Scripts/ directory of the ParlaMint GitHub repository, with the README file listing them.

It should be noted that it is not necessary to run the validation scripts directly, as the validation can be performed by the main Makefile of the project. The Makefile is self-documenting, i.e. to see how to use it, please run make help in the top level directory of the ParlaMint project.

While each contributor of a corpus should validate their files with the ParlaMint schemas and validation script, there also exist further stages of validation, which are also applied to ParlaMint corpora:

  • The corpora are converted to derived formats, in particular, the linguistically annotated version of the corpus to CoNLL-U and to the so called vertical format for CQP-type concordancers. The Universal Dependencies project provides a program for validating the formatting and linguistic analyses in CoNLL-U files, and this validation is used on the CoNLL-U files derived from their XML source, up to level 2 conformance. The vertical files, on the other hand, are first compiled with manatee (the back end of (no)Sketch Engine) and this compilation can also expose various errors.
  • The last stage in validation is ‘human validation’ where e.g. simply looking at various produced metadata files or at the concordances of a corpus exposes errors.

9.2. Finalisation of corpora

While the vast majority of converting source encodings into the ParlaMint corpus format is left to the compilers of a corpus, there are a few metadata elements that can be produced by a common script on the basis of nearly finished corpora, which then results in the final version of the corpus for a particular release. This includes setting the date, edition and handle under which the corpus will be distributed, and also calculating the size of the corpus (cf. the Sections on Extents and on Tags declaration). The script for finalisation can be found in the Scripts/ directory of the ParlaMint GitHub repository and the README file briefly explains its function; more comments can be found in the script itself.

9.3. Conversions

A TEI encoded document is, in general, not meant to be used directly by software programs, rather, it serves as an interchange and storage format. The ParlaMint project has produced various scripts to down-convert the XML encoded corpora to other formats and they can be found in the Scripts/ directory of the ParlaMint GitHub repository, with the README file listing them and explaining their function. In short, the scripts convert the ParlaMint XML to plain text, to CoNLL-U, and to vertical format. There is also a script that takes a ParlaMint corpus and makes from it a sample for inclusion to the ParlaMint GitHub repository.

10. Contributing to ParlaMint

The ParlaMint GitHub repository contains these guidelines, the ParlaMint XML schemas, the scripts used to validate, finalise and convert the ParlaMint TEI XML corpora to derived formats, and samples of the ParlaMint corpora. There are four main branches in the repository:

  • main is the default branch used for the synchronisation of other branches. It is also used for releasing sample files that correspond to published corpora.
  • data serves as a pushing place for new sample files in ./Data/ParlaMint-XX directories.
  • devel: development of scripts and documentation.

The validation procedure for corpora is explained in the Section on Validating ParlaMint corpora, while the technical aspects of contributing corpora is further explained in the CONTRIBUTING file of the repository.

11. Acknowledgements

The work on these recommendations was funded by the CLARIN Research Infrastructure for Language Resources and Tools.

Appendix A Formal specification

Appendix A.1 Elements

Appendix A.1.1 <TEI>

<TEI> (TEI document) contains a single TEI-conformant document, combining a single TEI header with one or more members of the model.resource class. Multiple <TEI> elements may be combined within a <TEI> (or <teiCorpus>) element. [4. Default Text Structure 16.1. Varieties of Composite Text]
Moduletextstructure — Formal specification
Attributes
xml:id
StatusRequired
DatatypeID
xml:lang
StatusRequired
Datatypeteidata.language
ana
StatusRequired
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Contained by
core: teiCorpus
May contain
header: teiHeader
textstructure: text
Note

As with all elements in the TEI scheme (except <egXML>) this element is in the TEI namespace (see 5.7.2. Namespaces). Thus, when it is used as the outermost element of a TEI document, it is necessary to specify the TEI namespace on it. This is customarily achieved by including http://www.tei-c.org/ns/1.0 as the value of the XML namespace declaration (xmlns), without indicating a prefix, and then not using a prefix on TEI elements in the rest of the document. For example: <TEI version="4.8.1" xml:lang="it" xmlns="http://www.tei-c.org/ns/1.0">.

ExampleExample of ParlaMint corpus component:
<TEI xml:id="ParlaMint-GB_2015-01-06-commons"  xml:lang="enana="#parla.sitting #reference" xmlns="http://www.tei-c.org/ns/1.0">  <teiHeader>...</teiHeader>  <text ana="#reference">   <body>...</body>  </text> -</TEI>
Content model
+</TEI>
Content model
 <content>
  <elementRef key="teiHeader"/>
  <elementRef key="text"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element TEI
 {
    tei_att.global.linking.attribute.corresp,
@@ -871,16 +871,16 @@
    attribute ana { list { + } },
    tei_teiHeader,
    tei_text
-}

Appendix A.1.2 <addName>

<addName> (additional name) contains an additional name component, such as a nickname, epithet, or alias, or any other descriptive phrase used within a personal name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: persName
May containCharacter data only
Example
<persName> +}

Appendix A.1.2 <addName>

<addName> (additional name) contains an additional name component, such as a nickname, epithet, or alias, or any other descriptive phrase used within a personal name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: persName
May containCharacter data only
Example
<persName>  <surname>Möderndorfer</surname>  <forename>Jani</forename>  <addName>Janko</addName> -</persName>
Content model
+</persName>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
-element addName { tei_att.global.attribute.xmllang, text }

Appendix A.1.3 <affiliation>

<affiliation> (affiliation) optionally contains the role name corresponding to the affiliation, the name of the organisation that the person is affiliated with, and notes giving further informal information about the affiliation. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
role
StatusRequired
Legal values are:
head
minister
member
academician
alternateOfDelegation
associateMember
candidateChairman
constitutionalJudge
deputyHead
deputyMinister
ministerDelegate
nonAttachedMember
observer
ombudsman
prosecutorGeneral
publicDefenderOfRights
replacement
representative
secretary
secretaryGeneral
secretaryOfState
verifier
vicePublicDefenderOfRights
Member of
Contained by
namesdates: person
May contain
core: note
namesdates: orgName roleName
Note

If included, the name of an organization may be tagged using either the <name> element as above, or the more specific <orgName> element.

Example
<person xml:id="AdamKalous.1979"> +
Schema Declaration
+element addName { tei_att.global.attribute.xmllang, text }

Appendix A.1.3 <affiliation>

<affiliation> (affiliation) optionally contains the role name corresponding to the affiliation, the name of the organisation that the person is affiliated with, and notes giving further informal information about the affiliation. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
role
StatusRequired
Legal values are:
head
minister
member
academician
alternateOfDelegation
associateMember
candidateChairman
constitutionalJudge
deputyHead
deputyMinister
ministerDelegate
nonAttachedMember
observer
ombudsman
prosecutorGeneral
publicDefenderOfRights
replacement
representative
secretary
secretaryGeneral
secretaryOfState
verifier
vicePublicDefenderOfRights
Member of
Contained by
namesdates: person
May contain
core: note
namesdates: orgName roleName
Note

If included, the name of an organization may be tagged using either the <name> element as above, or the more specific <orgName> element.

Example
<person xml:id="AdamKalous.1979">  <persName>   <surname>Kalous</surname>   <forename>Adam</forename> @@ -922,7 +922,7 @@   to="2021-10-21T00:00:00">   <roleName xml:lang="en">MP</roleName>  </affiliation> -</person>
Example
<p>The affiliation element can also include an <att>ana</att> attribute, which points to the appropriate legislative period when the person was affiliated with the specified organisation:</p> +</person>
Example
<p>The affiliation element can also include an <att>ana</att> attribute, which points to the appropriate legislative period when the person was affiliated with the specified organisation:</p> <person xml:id="BahŽibertAnja">  <persName>   <surname>Bah</surname> @@ -943,7 +943,7 @@   from="2018-06-22ana="#DZ.8">   <roleName xml:lang="en">MP</roleName>  </affiliation> -</person>
Content model
+</person>
Content model
 <content>
  <elementRef key="roleName" minOccurs="0"
   maxOccurs="unbounded"/>
@@ -952,7 +952,7 @@
  <elementRef key="note" minOccurs="0"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element affiliation
 {
    tei_att.global.analytic.attribute.ana,
@@ -990,13 +990,13 @@
    tei_roleName*,
    tei_orgName*,
    tei_note*
-}

Appendix A.1.4 <appInfo>

<appInfo> (application information) records information about an application which has edited the TEI file. [2.3.11. The Application Information Element]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
header: application
Example
<appInfo> +}

Appendix A.1.4 <appInfo>

<appInfo> (application information) records information about an application which has edited the TEI file. [2.3.11. The Application Information Element]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
header: application
Example
<appInfo>  <application version="4.0"   ident="stanford-corenlp">   <label>Stanford CoreNLP</label>   <desc>Tokenisation, POS tagging, NER and dependency parsed using Stanford CoreNLP <ref target="https://stanfordnlp.github.io/CoreNLP/">https://stanfordnlp.github.io/CoreNLP/</ref>.</desc>  </application> -</appInfo>
Example
<appInfo> +</appInfo>
Example
<appInfo>  <application version="1.0"   ident="reldi-tokeniser">   <label>ReLDI tokeniser</label> @@ -1009,13 +1009,13 @@   ident="janes-ner">   <label>NER system for South Slavic languages</label>  </application> -</appInfo>
Content model
+</appInfo>
Content model
 <content>
  <elementRef key="application"
   minOccurs="1" maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element appInfo { tei_application+ }

Appendix A.1.5 <application>

<application> provides information about an application which has acted upon the document. [2.3.11. The Application Information Element]
Moduleheader — Formal specification
Attributes
identsupplies an identifier for the application, independent of its version number or display name.
StatusRequired
Datatypeteidata.name
versionsupplies a version number for the application, independent of its identifier or display name.
StatusRequired
Datatypeteidata.versionNumber
Contained by
header: appInfo
May contain
core: desc label
Example
<appInfo> +
Schema Declaration
+element appInfo { tei_application+ }

Appendix A.1.5 <application>

<application> provides information about an application which has acted upon the document. [2.3.11. The Application Information Element]
Moduleheader — Formal specification
Attributes
identsupplies an identifier for the application, independent of its version number or display name.
StatusRequired
Datatypeteidata.name
versionsupplies a version number for the application, independent of its identifier or display name.
StatusRequired
Datatypeteidata.versionNumber
Contained by
header: appInfo
May contain
core: desc label
Example
<appInfo>  <application version="1"   ident="app-stanza">   <label>Stanza</label> @@ -1033,48 +1033,55 @@   <desc xml:lang="en">    <ref target="http://conllu2teixml">CoNLL-U 2 TEI XML</ref>: converter from CoNLL-U format to (ParlaClarin/ParlaMint) Tei XML Format</desc>  </application> -</appInfo>
Example
<appInfo> +</appInfo>
Example
<appInfo>  <application version="4.0"   ident="stanford-corenlp">   <label>Stanford CoreNLP</label>   <desc>Tokenisation, POS tagging, NER and dependency parsed using Stanford CoreNLP <ref target="https://stanfordnlp.github.io/CoreNLP/">https://stanfordnlp.github.io/CoreNLP/</ref>.</desc>  </application> -</appInfo>
Content model
+</appInfo>
Content model
 <content>
  <elementRef key="label"/>
  <elementRef key="desc" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element application
 {
    attribute ident { text },
    attribute version { text },
    tei_label,
    tei_desc+
-}

Appendix A.1.6 <availability>

<availability> (availability) supplies information about the availability of a text, for example any restrictions on its use or distribution, its copyright status, any licence applying to it, etc. [2.2.4. Publication, Distribution, Licensing, etc.]
Moduleheader — Formal specification
Attributes
status
StatusRequired
Legal values are:
free
Contained by
May contain
core: p
header: licence
Note

A consistent format should be adopted

Example
<availability status="free"> +}

Appendix A.1.6 <availability>

<availability> (availability) supplies information about the availability of a text, for example any restrictions on its use or distribution, its copyright status, any licence applying to it, etc. [2.2.4. Publication, Distribution, Licensing, etc.]
Moduleheader — Formal specification
Attributes
status
StatusRequired
Legal values are:
free
Contained by
May contain
core: p
header: licence
Note

A consistent format should be adopted

Example
<availability status="free">  <licence>http://creativecommons.org/licenses/by/4.0/</licence>  <p xml:lang="hr">Ovaj rad je dostupan pod <ref target="http://creativecommons.org/licenses/by/4.0/">međunarodnom licencom Creative Commons Imenovanje 4.0</ref>  </p>  <p xml:lang="en">This work is licensed under the <ref target="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ref>  </p> -</availability>
Content model
+</availability>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:availability"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="licence"/>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element availability { attribute status { "free" }, tei_licence, tei_p+ }

Appendix A.1.7 <bibl>

<bibl> (bibliographic citation) contains a loosely-structured bibliographic citation of which the sub-components may or may not be explicitly tagged. [3.12.1. Methods of Encoding Bibliographic References and Lists of References 2.2.7. The Source Description 16.3.2. Declarable Elements]
Modulecore — Formal specification
Member of
Contained by
header: sourceDesc
May contain
Note

Contains phrase-level elements, together with any combination of elements from the model.biblPart class

Example
<bibl> +
Schema Declaration
+element availability { attribute status { "free" }, tei_licence, tei_p+ }

Appendix A.1.7 <bibl>

<bibl> (bibliographic citation) contains a loosely-structured bibliographic citation of which the sub-components may or may not be explicitly tagged. [3.12.1. Methods of Encoding Bibliographic References and Lists of References 2.2.7. The Source Description 16.3.2. Declarable Elements]
Modulecore — Formal specification
Member of
Contained by
header: sourceDesc
May contain
Note

Contains phrase-level elements, together with any combination of elements from the model.biblPart class

Example
<bibl>  <title type="main">Minutes of the National Assembly of the Republic of Bulgaria</title>  <date when="2020-03-11">2020-03-11</date> -</bibl>
Example
<bibl> +</bibl>
Example
<bibl>  <title type="mainxml:lang="en">https://www.tbmm.gov.tr/tutanak/donem24/yil2/bas/b013m.htm</title>  <edition xml:lang="en">Official session record</edition>  <publisher xml:lang="en">The Turkish Parliament</publisher>  <idno type="URI">https://www.tbmm.gov.tr/</idno>  <date when="2011-10-27">2011-10-27</date> -</bibl>
Content model
+</bibl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:bibl"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="title" minOccurs="1"
   maxOccurs="unbounded"/>
@@ -1090,40 +1097,40 @@
    maxOccurs="1"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element bibl
 {
    tei_title+,
    ( tei_edition? | tei_publisher? | tei_idno* | tei_date )+
-}

Appendix A.1.8 <birth>

<birth> (birth) contains information about a person's birth, obligatorily its date and optionaly the place. Note that there can be several placeNames, all referring to the same place, but written in different languages or scripts. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
whensupplies the value of the date or time in a standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusRequired
Datatypeteidata.temporal.w3c
Contained by
namesdates: person
May contain
namesdates: placeName
Example
<person xml:id="ReinerŽeljkon="1291"> ... +}

Appendix A.1.8 <birth>

<birth> (birth) contains information about a person's birth, obligatorily its date and optionaly the place. Note that there can be several placeNames, all referring to the same place, but written in different languages or scripts. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
whensupplies the value of the date or time in a standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusRequired
Datatypeteidata.temporal.w3c
Contained by
namesdates: person
May contain
namesdates: placeName
Example
<person xml:id="ReinerŽeljkon="1291"> ... <birth when="1953-05-28"/> -</person>
Content model
+</person>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <elementRef key="placeName" minOccurs="0"
    maxOccurs="unbounded"/>
  </alternate>
 </content>
-    
Schema Declaration
-element birth { attribute when { text }, ( tei_placeName* ) }

Appendix A.1.9 <body>

<body> (text body) contains the whole body of a single unitary text, excluding any front or back matter. [4. Default Text Structure]
Moduletextstructure — Formal specification
Contained by
textstructure: text
May contain
textstructure: div
Example
<body> +
Schema Declaration
+element birth { attribute when { text }, ( tei_placeName* ) }

Appendix A.1.9 <body>

<body> (text body) contains the whole body of a single unitary text, excluding any front or back matter. [4. Default Text Structure]
Moduletextstructure — Formal specification
Contained by
textstructure: text
May contain
textstructure: div
Example
<body>  <div type="debateSection">...</div>  <div type="debateSection">...</div> ... -</body>
Content model
+</body>
Content model
 <content>
  <elementRef key="div" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element body { tei_div+ }

Appendix A.1.10 <catDesc>

<catDesc> (category description) describes some category within a taxonomy or text typology, either in the form of a brief prose description or in terms of the situational parameters used by the TEI formal <textDesc>. [2.3.7. The Classification Declaration]
Moduleheader — Formal specification
Attributes
Contained by
header: category
May contain
core: ref term
character data
Example
<category xml:id="parla.organisation"> +
Schema Declaration
+element body { tei_div+ }

Appendix A.1.10 <catDesc>

<catDesc> (category description) describes some category within a taxonomy or text typology, either in the form of a brief prose description or in terms of the situational parameters used by the TEI formal <textDesc>. [2.3.7. The Classification Declaration]
Moduleheader — Formal specification
Attributes
Contained by
header: category
May contain
core: ref term
character data
Example
<category xml:id="parla.organisation">  <catDesc xml:lang="en">   <term>Organisation</term>  </catDesc>  <catDesc xml:lang="bg">   <term>Организация</term>  </catDesc> -</category>
Content model
+</category>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="term"/>
@@ -1134,12 +1141,12 @@
   </alternate>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element catDesc
 {
    tei_att.global.attribute.xmllang,
    ( tei_term, ( text | tei_ref )+ )
-}

Appendix A.1.11 <catRef>

<catRef> (category reference) specifies one or more defined categories within some taxonomy or text typology. [2.4.3. The Text Classification]
Moduleheader — Formal specification
Attributes
targetspecifies the destination of the reference by supplying one or more URI References.
Derived fromatt.pointing
StatusRequired
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
schemeidentifies the classification scheme within which the set of categories concerned is defined, for example by a <taxonomy> element, or by some other resource.
StatusRequired
Datatypeteidata.pointer
Contained by
header: textClass
May containEmpty element
Note

The scheme attribute needs to be supplied only if more than one taxonomy has been declared.

Example
<textClass> +}

Appendix A.1.11 <catRef>

<catRef> (category reference) specifies one or more defined categories within some taxonomy or text typology. [2.4.3. The Text Classification]
Moduleheader — Formal specification
Attributes
targetspecifies the destination of the reference by supplying one or more URI References.
Derived fromatt.pointing
StatusRequired
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
schemeidentifies the classification scheme within which the set of categories concerned is defined, for example by a <taxonomy> element, or by some other resource.
StatusRequired
Datatypeteidata.pointer
Contained by
header: textClass
May containEmpty element
Note

The scheme attribute needs to be supplied only if more than one taxonomy has been declared.

Example
<textClass>  <catRef scheme="#parla.legislature"   target="#parla.uni"/> </textClass> @@ -1154,34 +1161,34 @@    <term>Unicameralism</term>   </catDesc>  </category> -</taxonomy>
Content model
+</taxonomy>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element catRef
 {
    attribute target { list { + } },
    attribute scheme { text },
    empty
-}

Appendix A.1.12 <category>

<category> (category) contains an individual descriptive category, possibly nested within a superordinate category, within a user-defined taxonomy. [2.3.7. The Classification Declaration]
Moduleheader — Formal specification
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
Derived fromatt.global
StatusRequired
DatatypeID
ana(analysis) indicates one or more elements containing interpretations of the element on which the ana attribute appears.
Derived fromatt.global.analytic
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Contained by
May contain
Example
<category xml:id="parla.session"> +}

Appendix A.1.12 <category>

<category> (category) contains an individual descriptive category, possibly nested within a superordinate category, within a user-defined taxonomy. [2.3.7. The Classification Declaration]
Moduleheader — Formal specification
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
Derived fromatt.global
StatusRequired
DatatypeID
ana(analysis) indicates one or more elements containing interpretations of the element on which the ana attribute appears.
Derived fromatt.global.analytic
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Contained by
May contain
Example
<category xml:id="parla.session">  <catDesc xml:lang="en">   <term>Session</term>: A parliamentary year, which always begins on the first Tuesday in October at 12.00 o’clock noon and ends on the same date at the same time the following year. However, parliamentary work at Christiansborg is organised in such a way that it primarily takes place from October to June.</catDesc> -</category>
Example
<category xml:id="parla.term"> +</category>
Example
<category xml:id="parla.term">  <catDesc xml:lang="nl">   <term>Zittingsperiode</term>  </catDesc>  <catDesc xml:lang="en">   <term>Legislative period</term>  </catDesc> -</category>
Content model
+</category>
Content model
 <content>
  <elementRef key="catDesc" minOccurs="1"
   maxOccurs="unbounded"/>
  <elementRef key="category" minOccurs="0"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element category
 {
    tei_att.global.attribute.n,
@@ -1189,12 +1196,12 @@
    attribute ana { list { + } }?,
    tei_catDesc+,
    tei_category*
-}

Appendix A.1.13 <change>

<change> (change) documents a change or set of changes made during the production of a source document, or during the revision of an electronic file. [2.6. The Revision Description 2.4.1. Creation 12.7. Identifying Changes and Revisions]
Moduleheader — Formal specification
Attributes
Contained by
header: revisionDesc
May contain
core: name
character data
Note

The who attribute may be used to point to any other element, but will typically specify a <respStmt> or <person> element elsewhere in the header, identifying the person responsible for the change and their role in making it.

It is recommended that changes be recorded with the most recent first. The status attribute may be used to indicate the status of a document following the change documented.

Example
<revisionDesc> +}

Appendix A.1.13 <change>

<change> (change) documents a change or set of changes made during the production of a source document, or during the revision of an electronic file. [2.6. The Revision Description 2.4.1. Creation 12.7. Identifying Changes and Revisions]
Moduleheader — Formal specification
Attributes
Contained by
header: revisionDesc
May contain
core: name
character data
Note

The who attribute may be used to point to any other element, but will typically specify a <respStmt> or <person> element elsewhere in the header, identifying the person responsible for the change and their role in making it.

It is recommended that changes be recorded with the most recent first. The status attribute may be used to indicate the status of a document following the change documented.

Example
<revisionDesc>  <change when="2021-01-28">   <name>Tommaso Agnoloni</name>: Generated corpus in ParlaMint.</change>  <change when="2021-02-26">   <name>Tommaso Agnoloni</name>, <name>Francesca Frontini</name>: Corpus revision, fixing</change> -</revisionDesc>
Content model
+</revisionDesc>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -1202,19 +1209,19 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element change
 {
    tei_att.global.attribute.n,
    tei_att.global.attribute.xmllang,
    tei_att.datable.w3c.attribute.when,
    ( tei_name | text )+
-}

Appendix A.1.14 <classDecl>

<classDecl> (classification declarations) contains taxonomies defining classificatory codes used elsewhere in the text. Note that the taxonomies are in ParlaMint typically stored in separate files. [2.3.7. The Classification Declaration 2.3. The Encoding Description]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
derived-module-parlamint: include
header: taxonomy
Example
<classDecl> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-taxonomy-parla.legislature.xml"/> +}

Appendix A.1.14 <classDecl>

<classDecl> (classification declarations) contains taxonomies defining classificatory codes used elsewhere in the text. Note that the taxonomies are in ParlaMint typically stored in separate files. [2.3.7. The Classification Declaration 2.3. The Encoding Description]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
derived-module-parlamint: include
header: taxonomy
Example
<classDecl> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" + href="ParlaMint-SI-taxonomy-parla.legislature.xml"/> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-taxonomy.xml-speaker_types"/> + href="ParlaMint-SI-taxonomy.xml-speaker_types"/> ... -</classDecl>
Content model
+</classDecl>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -1222,18 +1229,22 @@
   <elementRef key="include"/>
  </alternate>
 </content>
-    
Schema Declaration
-element classDecl { ( tei_taxonomy | tei_include )+ }

Appendix A.1.15 <correction>

<correction> (correction principles) states how and under what circumstances corrections have been made in the text. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Note

May be used to note the results of proof reading the text against its original, indicating (for example) whether discrepancies have been silently rectified, or recorded using the editorial tags described in section 3.5. Simple Editorial Changes.

Example
<editorialDecl> +
Schema Declaration
+element classDecl { ( tei_taxonomy | tei_include )+ }

Appendix A.1.15 <correction>

<correction> (correction principles) states how and under what circumstances corrections have been made in the text. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Note

May be used to note the results of proof reading the text against its original, indicating (for example) whether discrepancies have been silently rectified, or recorded using the editorial tags described in section 3.5. Simple Editorial Changes.

Example
<editorialDecl>  <correction>   <p>No correction of source texts was performed.</p>  </correction> -</editorialDecl>
Content model
+</editorialDecl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:correction"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element correction { tei_p+ }

Appendix A.1.16 <date>

<date> (date) contains a date in any format. [3.6.4. Dates and Times 2.2.4. Publication, Distribution, Licensing, etc. 2.6. The Revision Description 3.12.2.4. Imprint, Size of a Document, and Reprint Information 16.2.3. The Setting Description 14.4. Dates]
Modulecore — Formal specification
Attributes
Member of
Contained by
analysis: s
corpus: setting
May contain
analysis: pc w
core: date
character data
ExampleThe element <date> gives the date in the when attribute in the ISO 8601 format, while the textual content is not constrained:
<date when="2021-06-08">2021-06-08</date>
ExampleThe textual content can be given according to the conventions used in the local language:
<date when="2018-04-13xml:lang="sl">13.4.2018</date>
Content model
+    
Schema Declaration
+element correction { tei_p+ }

Appendix A.1.16 <date>

<date> (date) contains a date in any format. [3.6.4. Dates and Times 2.2.4. Publication, Distribution, Licensing, etc. 2.6. The Revision Description 3.12.2.4. Imprint, Size of a Document, and Reprint Information 16.2.3. The Setting Description 14.4. Dates]
Modulecore — Formal specification
Attributes
Member of
Contained by
analysis: s
corpus: setting
May contain
analysis: pc w
core: date
character data
ExampleThe element <date> gives the date in the when attribute in the ISO 8601 format, while the textual content is not constrained:
<date when="2021-06-08">2021-06-08</date>
ExampleThe textual content can be given according to the conventions used in the local language:
<date when="2018-04-13xml:lang="sl">13.4.2018</date>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -1243,7 +1254,7 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element date
 {
    tei_att.global.attribute.xmlid,
@@ -1254,15 +1265,15 @@
    tei_att.datable.w3c.attribute.to,
    tei_att.typed.attributes,
    ( tei_w | tei_pc | tei_date | text )+
-}

Appendix A.1.17 <death>

<death> (death) contains information about a person's death, obligatorily its date and optionaly the place. Note that there can be several placeNames, all referring to the same place, but written in different languages or scripts. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
whensupplies the value of the date or time in a standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusRequired
Datatypeteidata.temporal.w3c
Contained by
namesdates: person
May contain
namesdates: placeName
Example
<death when="2020-12-29"/>
Content model
+}

Appendix A.1.17 <death>

<death> (death) contains information about a person's death, obligatorily its date and optionaly the place. Note that there can be several placeNames, all referring to the same place, but written in different languages or scripts. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
whensupplies the value of the date or time in a standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusRequired
Datatypeteidata.temporal.w3c
Contained by
namesdates: person
May contain
namesdates: placeName
Example
<death when="2020-12-29"/>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <elementRef key="placeName" minOccurs="0"
    maxOccurs="unbounded"/>
  </alternate>
 </content>
-    
Schema Declaration
-element death { attribute when { text }, ( tei_placeName* ) }

Appendix A.1.18 <desc>

<desc> (description) contains a short description of the purpose, function, or use of its parent element, or when the parent is a documentation element, describes or defines the object being documented. [23.4.1. Description of Components]
Modulecore — Formal specification
Attributes
Member of
Contained by
core: gap
namesdates: org
May contain
core: ref term
character data
Note

When used in a specification element such as <elementSpec>, TEI convention requires that this be expressed as a finite clause, begining with an active verb.

Example
<p>Example of <gi>desc</gi> elements for transcriber comments:</p> +
Schema Declaration
+element death { attribute when { text }, ( tei_placeName* ) }

Appendix A.1.18 <desc>

<desc> (description) contains a short description of the purpose, function, or use of its parent element, or when the parent is a documentation element, describes or defines the object being documented. [23.4.1. Description of Components]
Modulecore — Formal specification
Attributes
Member of
Contained by
core: gap
namesdates: org
May contain
core: ref term
character data
Note

When used in a specification element such as <elementSpec>, TEI convention requires that this be expressed as a finite clause, begining with an active verb.

Example
<p>Example of <gi>desc</gi> elements for transcriber comments:</p> <gap reason="inaudible">  <desc>speaker spoke too quietly, not understood</desc> </gap> @@ -1280,7 +1291,7 @@ <incident type="action">  <desc>minute of silence</desc> -</incident>
ExampleExample of <desc> elements used as a part of taxonomy:
<taxonomy xml:id="parla.legislature"> +</incident>
ExampleExample of <desc> elements used as a part of taxonomy:
<taxonomy xml:id="parla.legislature">  <desc xml:lang="sl">   <term>Zakonodajna oblast</term>  </desc> @@ -1289,7 +1300,7 @@  </desc> ... -</taxonomy>
ExampleElement <desc> can also be used to describe tool(s) used to linguistically annotate the corpus:
<application version="1.0" +</taxonomy>
ExampleElement <desc> can also be used to describe tool(s) used to linguistically annotate the corpus:
<application version="1.0"  ident="reldi-tokeniser">  <label>ReLDI tokeniser</label>  <desc xml:lang="en">Tokenisation and sentence segmentation with ReLDI tokeniser, available from <ref target="https://github.com/clarinsi/reldi-tokeniser">https://github.com/clarinsi/reldi-tokeniser</ref>.</desc> @@ -1300,7 +1311,7 @@ that is being deprecated: that is, only an element that has a @validUntil attribute should have a child <desc type="deprecationInfo">.</sch:assert> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef minOccurs="0" key="term"/>
@@ -1311,12 +1322,12 @@
   </alternate>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element desc
 {
    tei_att.global.attribute.xmllang,
    ( tei_term?, ( text | tei_ref )+ )
-}

Appendix A.1.19 <div>

<div> (text division) contains division of the body a corpus component. [4.1. Divisions of the Body]
Moduletextstructure — Formal specification
Attributes
type
StatusRequired
Legal values are:
debateSection
General purpose text division for all parts of parliamentary proceedings. It should include at least one utterance. If needed, the @subtype attribute can be used for additional content classification.
commentSection
A special purpose text division used as a container for transcriber comments. Should not contain any utterances. If needed, the @subtype attribute can be used for additional content classification.
Contained by
textstructure: body
May contain
Example
<div type="debateSection"> +}

Appendix A.1.19 <div>

<div> (text division) contains division of the body a corpus component. [4.1. Divisions of the Body]
Moduletextstructure — Formal specification
Attributes
type
StatusRequired
Legal values are:
debateSection
General purpose text division for all parts of parliamentary proceedings. It should include at least one utterance. If needed, the @subtype attribute can be used for additional content classification.
commentSection
A special purpose text division used as a container for transcriber comments. Should not contain any utterances. If needed, the @subtype attribute can be used for additional content classification.
Contained by
textstructure: body
May contain
Example
<div type="debateSection">  <head>Devolution of Power (Cities)</head>  <u xml:id="ParlaMint-GB_2015-01-06-commons.u1">...</u>  <u xml:id="ParlaMint-GB_2015-01-06-commons.u2">...</u> @@ -1330,7 +1341,7 @@ <sch:rule context="tei:div"> <sch:report test="(ancestor::tei:p or ancestor::tei:ab) and not(ancestor::tei:floatingText)"> Abstract model violation: p and ab may not contain higher-level structural elements such as div, unless div is a descendant of floatingText. </sch:report> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <elementRef key="head" minOccurs="0"
   maxOccurs="unbounded"/>
@@ -1345,7 +1356,7 @@
   <elementRef key="u"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element div
 {
    tei_att.global.attribute.xmlid,
@@ -1364,20 +1375,20 @@
     | tei_pb
     | tei_u
    )+
-}

Appendix A.1.20 <edition>

<edition> (edition) describes the particularities of one edition of a text. [2.2.2. The Edition Statement]
Moduleheader — Formal specification
Attributes
Contained by
core: bibl
header: editionStmt
May containCharacter data only
Example
<edition>2.1</edition>
Content model
+}

Appendix A.1.20 <edition>

<edition> (edition) describes the particularities of one edition of a text. [2.2.2. The Edition Statement]
Moduleheader — Formal specification
Attributes
Contained by
core: bibl
header: editionStmt
May containCharacter data only
Example
<edition>2.1</edition>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
-element edition { tei_att.global.attribute.xmllang, text }

Appendix A.1.21 <editionStmt>

<editionStmt> (edition statement) groups information relating to one edition of a text. [2.2.2. The Edition Statement 2.2. The File Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
header: edition
Example
<editionStmt> +
Schema Declaration
+element edition { tei_att.global.attribute.xmllang, text }

Appendix A.1.21 <editionStmt>

<editionStmt> (edition statement) groups information relating to one edition of a text. [2.2.2. The Edition Statement 2.2. The File Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
header: edition
Example
<editionStmt>  <edition>2.1</edition> -</editionStmt>
Content model
+</editionStmt>
Content model
 <content>
  <elementRef key="edition" minOccurs="1"
   maxOccurs="1"/>
 </content>
-    
Schema Declaration
-element editionStmt { tei_edition }

Appendix A.1.22 <editorialDecl>

<editorialDecl> (editorial practice declaration) provides details of editorial principles and practices applied during the encoding of a text. [2.3.3. The Editorial Practices Declaration 2.3. The Encoding Description 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
Example
<editorialDecl> +
Schema Declaration
+element editionStmt { tei_edition }

Appendix A.1.22 <editorialDecl>

<editorialDecl> (editorial practice declaration) provides details of editorial principles and practices applied during the encoding of a text. [2.3.3. The Editorial Practices Declaration 2.3. The Encoding Description 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
Example
<editorialDecl>  <correction>   <p>No correction of source texts was performed.</p>  </correction> @@ -1393,7 +1404,11 @@  <segmentation>   <p>The texts are segmented into utterances (contributions) and segments (corresponding to paragraphs in the source transcription).</p>  </segmentation> -</editorialDecl>
Content model
+</editorialDecl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:editorialDecl"/> +</sch:pattern>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -1404,7 +1419,7 @@
   <elementRef key="segmentation"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element editorialDecl
 {
    (
@@ -1414,11 +1429,11 @@
     | tei_quotation
     | tei_segmentation
    )+
-}

Appendix A.1.23 <education>

<education> (education) contains a description of the educational experience of a person. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: person
May containCharacter data only
Example
<education>Bachelor of Science, Electrical and Information Technology Engineer</education>
Content model
+}

Appendix A.1.23 <education>

<education> (education) contains a description of the educational experience of a person. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: person
May containCharacter data only
Example
<education>Bachelor of Science, Electrical and Information Technology Engineer</education>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element education
 {
    tei_att.global.attribute.n,
@@ -1427,12 +1442,12 @@
    tei_att.datable.w3c.attribute.from,
    tei_att.datable.w3c.attribute.to,
    text
-}

Appendix A.1.24 <email>

<email> (electronic mail address) contains an email address identifying a location to which email messages can be delivered. [3.6.2. Addresses]
Modulecore — Formal specification
Attributes
Member of
Contained by
core: unit
May contain
analysis: pc w
character data
Note

The format of a modern Internet email address is defined in RFC 2822

ExampleThe element can be used for fine-grained Named Entities which include e-mail addresses:
<email ana="ne:me" +}

Appendix A.1.24 <email>

<email> (electronic mail address) contains an email address identifying a location to which email messages can be delivered. [3.6.2. Addresses]
Modulecore — Formal specification
Attributes
Member of
Contained by
core: unit
May contain
analysis: pc w
character data
Note

The format of a modern Internet email address is defined in RFC 2822

ExampleThe element can be used for fine-grained Named Entities which include e-mail addresses:
<email ana="ne:me"  xml:id="ParlaMint-CZ_2014-12-09-ps2013-023-05-003-133.ne87">  <w xml:id="ParlaMint-CZ_2014-12-09-ps2013-023-05-003-133.u4.p9.s3.w13"   lemma="namraza@cd.cz"   msd="UPosTag=NOUN|Case=Gen|Gender=Fem|Number=Plur|Polarity=Pos">namraza@cd.cz</w> -</email>
Content model
+</email>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -1441,19 +1456,19 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element email
 {
    tei_att.global.attribute.xmlid,
    tei_att.global.attribute.xmllang,
    tei_att.global.analytic.attribute.ana,
    ( tei_w | tei_pc | text )+
-}

Appendix A.1.25 <encodingDesc>

<encodingDesc> (encoding description) documents the relationship between an electronic text and the source or sources from which it was derived. [2.3. The Encoding Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Contained by
header: teiHeader
May contain
ExampleGeneral structure of an encoding description:
<encodingDesc> +}

Appendix A.1.25 <encodingDesc>

<encodingDesc> (encoding description) documents the relationship between an electronic text and the source or sources from which it was derived. [2.3. The Encoding Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Contained by
header: teiHeader
May contain
ExampleGeneral structure of an encoding description:
<encodingDesc>  <projectDesc>...</projectDesc>  <editorialDecl>...</editorialDecl>  <tagsDecl>...</tagsDecl>  <classDecl>...</classDecl> -</encodingDesc>
ExampleStructure of an encoding description for unannotated corpus root:
<encodingDesc> +</encodingDesc>
ExampleStructure of an encoding description for unannotated corpus root:
<encodingDesc>  <projectDesc>   <p xml:lang="sl">    <ref target="https://www.clarin.eu/content/parlamint">ParlaMint</ref> @@ -1476,7 +1491,7 @@   </namespace>  </tagsDecl>  <classDecl>...</classDecl> -</encodingDesc>
ExampleExample of encoding description of an annotated corpus root. The structure includes two additional elements, <listPrefixDef> and <appInfo>.
<encodingDesc> +</encodingDesc>
ExampleExample of encoding description of an annotated corpus root. The structure includes two additional elements, <listPrefixDef> and <appInfo>.
<encodingDesc>  <projectDesc>... </projectDesc>  <editorialDecl>...</editorialDecl>  <tagsDecl>...</tagsDecl> @@ -1491,10 +1506,10 @@  <appInfo>   <application>...</application>  </appInfo> -</encodingDesc>
ExampleExample of encoding description of a corpus component (annotated or unannotated). In contrast to the corpus root, the encoding description of a corpus component contains only two elements, namely, the <projectDesc> and the <tagsDecl>.
<encodingDesc> +</encodingDesc>
ExampleExample of encoding description of a corpus component (annotated or unannotated). In contrast to the corpus root, the encoding description of a corpus component contains only two elements, namely, the <projectDesc> and the <tagsDecl>.
<encodingDesc>  <projectDesc>...</projectDesc>  <tagsDecl>...</tagsDecl> -</encodingDesc>
Content model
+</encodingDesc>
Content model
 <content>
  <elementRef key="projectDesc"/>
  <elementRef key="editorialDecl"
@@ -1507,7 +1522,7 @@
  <elementRef key="appInfo" minOccurs="0"
   maxOccurs="1"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element encodingDesc
 {
    tei_projectDesc,
@@ -1516,7 +1531,7 @@
    tei_classDecl?,
    tei_listPrefixDef?,
    tei_appInfo?
-}

Appendix A.1.26 <equipment>

<equipment> (equipment) provides technical details of the equipment and media used for an audio or video recording used as the source for a spoken text. [8.2. Documenting the Source of Transcribed Speech 16.3.2. Declarable Elements]
Modulespoken — Formal specification
Attributes
Contained by
May contain
core: p
Example
<equipment> +}

Appendix A.1.26 <equipment>

<equipment> (equipment) provides technical details of the equipment and media used for an audio or video recording used as the source for a spoken text. [8.2. Documenting the Source of Transcribed Speech 16.3.2. Declarable Elements]
Modulespoken — Formal specification
Attributes
Contained by
May contain
core: p
Example
<equipment>  <p>"Hi-8" 8 mm NTSC camcorder with integral directional    microphone and windshield and stereo digital sound    recording channel. @@ -1524,49 +1539,52 @@ </equipment>
Example
<equipment>  <p>8-track analogue transfer mixed down to 19 cm/sec audio    tape for cassette mastering</p> -</equipment>
Content model
+</equipment>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:equipment"/> +</sch:pattern>
Content model
 <content>
  <classRef key="model.pLike" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element equipment
 {
    tei_att.global.attributes,
    tei_att.declarable.attributes,
    tei_model.pLike+
-}

Appendix A.1.27 <equipment>

<equipment> (equipment) provides technical details of the equipment and media used for an audio or video recording used as the source for a spoken text.
Modulespoken — Formal specification
Attributes
Contained by
May contain
core: p
Example
<equipment> +}

Appendix A.1.27 <equipment>

<equipment> (equipment) provides technical details of the equipment and media used for an audio or video recording used as the source for a spoken text.
Modulespoken — Formal specification
Attributes
Contained by
May contain
core: p
Example
<equipment>  <p>"Hi-8" 8 mm NTSC camcorder with integral directional    microphone and windshield and stereo digital sound    recording channel.  </p> -</equipment>
Content model
+</equipment>
Content model
 <content>
  <classRef key="model.pLike" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element equipment
 {
    tei_att.global.attributes,
    tei_att.declarable.attributes,
    tei_model.pLike+
-}

Appendix A.1.28 <event>

<event> (event) contains data relating to any kind of significant event associated with a person, place, or organisation. [14.3.1. Basic Principles]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: listEvent org
May contain
core: label
Example
<event xml:id="PoGB.55from="2010-05-18" +}

Appendix A.1.28 <event>

<event> (event) contains data relating to any kind of significant event associated with a person, place, or organisation. [14.3.1. Basic Principles]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: listEvent org
May contain
core: label
Example
<event xml:id="PoGB.55from="2010-05-18"  to="2015-03-30">  <label>Fifty-fifth Parliament of the United Kingdom</label> -</event>
Example
<org xml:id="government.HR" +</event>
Example
<org xml:id="government.HR"  role="government">  <orgName xml:lang="hrfull="yes">Vlada Republike Hrvatske</orgName>  <orgName xml:lang="enfull="yes">Government of the Republic of Croatia</orgName>  <event from="1990-05-30">   <label xml:lang="en">existence</label>  </event> -</org>
Content model
+</org>
Content model
 <content>
  <elementRef key="label" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element event
 {
    tei_att.global.attribute.xmlid,
@@ -1574,7 +1592,7 @@
    tei_att.datable.w3c.attribute.from,
    tei_att.datable.w3c.attribute.to,
    tei_label+
-}

Appendix A.1.29 <extent>

<extent> (extent) describes the approximate size of a text stored on some carrier medium or of some other object, digital or non-digital, specified in any convenient units. [2.2.3. Type and Extent of File 2.2. The File Description 3.12.2.4. Imprint, Size of a Document, and Reprint Information 11.7.1. Object Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
core: measure
Example
<extent> +}

Appendix A.1.29 <extent>

<extent> (extent) describes the approximate size of a text stored on some carrier medium or of some other object, digital or non-digital, specified in any convenient units. [2.2.3. Type and Extent of File 2.2. The File Description 3.12.2.4. Imprint, Size of a Document, and Reprint Information 11.7.1. Object Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
core: measure
Example
<extent>  <measure unit="speechesquantity="75122"   xml:lang="sl">75.122 govorov</measure>  <measure unit="speechesquantity="75122" @@ -1583,29 +1601,29 @@   xml:lang="sl">20.190.034 besed</measure>  <measure unit="wordsquantity="20190034"   xml:lang="en">20,190,034 words</measure> -</extent>
Content model
+</extent>
Content model
 <content>
  <elementRef key="measure" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element extent { tei_measure+ }

Appendix A.1.30 <figure>

<figure> (figure) groups elements representing or containing graphic information such as an illustration, formula, or figure. [15.4. Specific Elements for Graphic Images]
Modulefigures — Formal specification
Member of
Contained by
namesdates: person
May contain
Example
<figure> +
Schema Declaration
+element extent { tei_measure+ }

Appendix A.1.30 <figure>

<figure> (figure) groups elements representing or containing graphic information such as an illustration, formula, or figure. [15.4. Specific Elements for Graphic Images]
Modulefigures — Formal specification
Member of
Contained by
namesdates: person
May contain
Example
<figure>  <graphic url="https://www.psp.cz/eknih/cdrom/2017ps/eknih/2017ps/poslanci/i6497.jpg"/> -</figure>
Content model
+</figure>
Content model
 <content>
  <elementRef key="head" minOccurs="0"
   maxOccurs="1"/>
  <elementRef key="graphic" minOccurs="1"
   maxOccurs="1"/>
 </content>
-    
Schema Declaration
-element figure { tei_head?, tei_graphic }

Appendix A.1.31 <fileDesc>

<fileDesc> (file description) contains a full bibliographic description of an electronic file. [2.2. The File Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Contained by
header: teiHeader
May contain
Note

The major source of information for those seeking to create a catalogue entry or bibliographic citation for an electronic file. As such, it provides a title and statements of responsibility together with details of the publication or distribution of the file, of any series to which it belongs, and detailed bibliographic notes for matters not addressed elsewhere in the header. It also contains a full bibliographic description for the source or sources from which the electronic text was derived.

ExampleBasic structure of the <fileDesc> element:
<fileDesc> +
Schema Declaration
+element figure { tei_head?, tei_graphic }

Appendix A.1.31 <fileDesc>

<fileDesc> (file description) contains a full bibliographic description of an electronic file. [2.2. The File Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Contained by
header: teiHeader
May contain
Note

The major source of information for those seeking to create a catalogue entry or bibliographic citation for an electronic file. As such, it provides a title and statements of responsibility together with details of the publication or distribution of the file, of any series to which it belongs, and detailed bibliographic notes for matters not addressed elsewhere in the header. It also contains a full bibliographic description for the source or sources from which the electronic text was derived.

ExampleBasic structure of the <fileDesc> element:
<fileDesc>  <titleStmt>...</titleStmt>  <editionStmt>...</editionStmt>  <extent>...</extent>  <publicationStmt>...</publicationStmt>  <sourceDesc>...</sourceDesc> -</fileDesc>
ExampleExample of the <fileDesc> element in a corpus root:
<fileDesc> +</fileDesc>
ExampleExample of the <fileDesc> element in a corpus root:
<fileDesc>  <titleStmt>   <title type="mainxml:lang="en">Dutch parliamentary corpus ParlaMint-NL [ParlaMint]</title>   <title type="mainxml:lang="nl">Corpus van het Nederlandse Parlement ParlaMint-NL [ParlaMint]</title> @@ -1668,7 +1686,7 @@    <date from="2014-04-16to="2020-10-14">2014-04-16 - 2020-10-14</date>   </bibl>  </sourceDesc> -</fileDesc>
ExampleExample of the <fileDesc> element in a corpus component:
<fileDesc> +</fileDesc>
ExampleExample of the <fileDesc> element in a corpus component:
<fileDesc>  <titleStmt>   <title type="mainxml:lang="en">Dutch parliamentary corpus ParlaMint-NL, Lower House 2014-04-16 [ParlaMint]</title>   <title type="mainxml:lang="nl">Corpus van het Nederlandse parlement ParlaMint-NL, Tweede Kamer 2014-04-16 [ParlaMint]</title> @@ -1721,7 +1739,7 @@    <date when="2014-04-16">2014-04-16</date>   </bibl>  </sourceDesc> -</fileDesc>
Content model
+</fileDesc>
Content model
 <content>
  <elementRef key="titleStmt"/>
  <elementRef key="editionStmt"/>
@@ -1729,7 +1747,7 @@
  <elementRef key="publicationStmt"/>
  <elementRef key="sourceDesc"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element fileDesc
 {
    tei_titleStmt,
@@ -1737,38 +1755,38 @@
    tei_extent,
    tei_publicationStmt,
    tei_sourceDesc
-}

Appendix A.1.32 <forename>

<forename> (forename) contains a forename, given or baptismal name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: persName
May containCharacter data only
Example
<persName> +}

Appendix A.1.32 <forename>

<forename> (forename) contains a forename, given or baptismal name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: persName
May containCharacter data only
Example
<persName>  <surname>Bongiorno</surname>  <forename>Giulia</forename> -</persName>
Content model
+</persName>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
-element forename { tei_att.global.attribute.xmllang, text }

Appendix A.1.33 <funder>

<funder> (funding body) specifies the name of an individual, institution, or organisation responsible for the funding of a project or text. [2.2.1. The Title Statement]
Moduleheader — Formal specification
Contained by
header: titleStmt
May contain
core: ref
namesdates: orgName
Note

Funders provide financial support for a project; they are distinct from sponsors (see element <sponsor>), who provide intellectual support and authority.

Example
<funder> +
Schema Declaration
+element forename { tei_att.global.attribute.xmllang, text }

Appendix A.1.33 <funder>

<funder> (funding body) specifies the name of an individual, institution, or organisation responsible for the funding of a project or text. [2.2.1. The Title Statement]
Moduleheader — Formal specification
Contained by
header: titleStmt
May contain
core: ref
namesdates: orgName
Note

Funders provide financial support for a project; they are distinct from sponsors (see element <sponsor>), who provide intellectual support and authority.

Example
<funder>  <orgName xml:lang="es">CLARIN infraestructura de investigación científica</orgName>  <orgName xml:lang="en">The CLARIN research infrastructure</orgName> -</funder>
Content model
+</funder>
Content model
 <content>
  <elementRef key="orgName" minOccurs="1"
   maxOccurs="unbounded"/>
  <elementRef key="ref" minOccurs="0"
   maxOccurs="1"/>
 </content>
-    
Schema Declaration
-element funder { tei_orgName+, tei_ref? }

Appendix A.1.34 <gap>

<gap> (gap) indicates a point where material has been omitted in a transcription, whether for editorial reasons described in the TEI header, as part of sampling practice, or because the material is illegible, invisible, or inaudible. [3.5.3. Additions, Deletions, and Omissions]
Modulecore — Formal specification
Attributes
reason
StatusRecommended
Legal values are:
inaudible
editorial
foreign
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Note

The <gap>, <unclear>, and <del> core tag elements may be closely allied in use with the <damage> and <supplied> elements, available when using the additional tagset for transcription of primary sources. See section 12.3.3.2. Use of the gap, del, damage, unclear, and supplied Elements in Combination for discussion of which element is appropriate for which circumstance.

The <gap> tag simply signals the editors decision to omit or inability to transcribe a span of text. Other information, such as the interpretation that text was deliberately erased or covered, should be indicated using the relevant tags, such as <del> in the case of deliberate deletion.

Example
<gap reason="inaudible"> +
Schema Declaration
+element funder { tei_orgName+, tei_ref? }

Appendix A.1.34 <gap>

<gap> (gap) indicates a point where material has been omitted in a transcription, whether for editorial reasons described in the TEI header, as part of sampling practice, or because the material is illegible, invisible, or inaudible. [3.5.3. Additions, Deletions, and Omissions]
Modulecore — Formal specification
Attributes
reason
StatusRecommended
Legal values are:
inaudible
editorial
foreign
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Note

The <gap>, <unclear>, and <del> core tag elements may be closely allied in use with the <damage> and <supplied> elements, available when using the additional tagset for transcription of primary sources. See section 12.3.3.2. Use of the gap, del, damage, unclear, and supplied Elements in Combination for discussion of which element is appropriate for which circumstance.

The <gap> tag simply signals the editors decision to omit or inability to transcribe a span of text. Other information, such as the interpretation that text was deliberately erased or covered, should be indicated using the relevant tags, such as <del> in the case of deliberate deletion.

Example
<gap reason="inaudible">  <desc>microphone muted</desc> -</gap>
Example
<gap reason="editorial"> +</gap>
Example
<gap reason="editorial">  <desc xml:lang="de">Zitierte Druckfassung entfernt</desc>  <desc xml:lang="en">Quoted printed matter omited</desc> -</gap>
Example
<gap reason="foreign"> +</gap>
Example
<gap reason="foreign">  <desc xml:lang="und">Huliniahuanngittunga</desc> -</gap>
Content model
+</gap>
Content model
 <content>
  <elementRef key="desc" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element gap
 {
    tei_att.global.attribute.xmlid,
@@ -1777,23 +1795,23 @@
    tei_att.global.linking.attribute.corresp,
    attribute reason { "inaudible" | "editorial" | "foreign" }?,
    tei_desc+
-}

Appendix A.1.35 <graphic>

<graphic> (graphic) indicates the location of a graphic or illustration, either forming part of a text, or providing an image of it. [3.10. Graphics and Other Non-textual Components 12.1. Digital Facsimiles]
Modulecore — Formal specification
Attributes
Member of
Contained by
figures: figure
May containEmpty element
Note

The mimeType attribute should be used to supply the MIME media type of the image specified by the url attribute.

Within the body of a text, a <graphic> element indicates the presence of a graphic component in the source itself. Within the context of a <facsimile> or <sourceDoc> element, however, a <graphic> element provides an additional digital representation of some part of the source being encoded.

Example
<figure> +}

Appendix A.1.35 <graphic>

<graphic> (graphic) indicates the location of a graphic or illustration, either forming part of a text, or providing an image of it. [3.10. Graphics and Other Non-textual Components 12.1. Digital Facsimiles]
Modulecore — Formal specification
Attributes
Member of
Contained by
figures: figure
May containEmpty element
Note

The mimeType attribute should be used to supply the MIME media type of the image specified by the url attribute.

Within the body of a text, a <graphic> element indicates the presence of a graphic component in the source itself. Within the context of a <facsimile> or <sourceDoc> element, however, a <graphic> element provides an additional digital representation of some part of the source being encoded.

Example
<figure>  <graphic url="https://www.dekamer.be//site/wwwroot/images/cv/06595.gif"/> -</figure>
Content model
+</figure>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element graphic
 {
    tei_att.media.attribute.scale,
    tei_att.resourced.attributes,
    empty
-}

Appendix A.1.36 <head>

<head> (heading) contains any type of heading, for example the title of a section, or the heading of a list, glossary, manuscript description, etc. [4.2.1. Headings and Trailers]
Modulecore — Formal specification
Attributes
Contained by
figures: figure
textstructure: div
May containCharacter data only
Note

The <head> element is used for headings at all levels; software which treats (e.g.) chapter headings, section headings, and list titles differently must determine the proper processing of a <head> element based on its structural position. A <head> occurring as the first element of a list is the title of that list; one occurring as the first element of a <div1> is the title of that chapter or section.

ExampleThe most common use for the <head> element is to mark the headings of sections:
<div type="debateSection"> +}

Appendix A.1.36 <head>

<head> (heading) contains any type of heading, for example the title of a section, or the heading of a list, glossary, manuscript description, etc. [4.2.1. Headings and Trailers]
Modulecore — Formal specification
Attributes
Contained by
figures: figure
textstructure: div
May containCharacter data only
Note

The <head> element is used for headings at all levels; software which treats (e.g.) chapter headings, section headings, and list titles differently must determine the proper processing of a <head> element based on its structural position. A <head> occurring as the first element of a list is the title of that list; one occurring as the first element of a <div1> is the title of that chapter or section.

ExampleThe most common use for the <head> element is to mark the headings of sections:
<div type="debateSection">  <head>Regulation of Health and Social Care Professions Etc. Bill [HL]</head> ... -</div>
ExampleThe <head> element may also be used to give the title to specialised lists:
<listEvent> +</div>
ExampleThe <head> element may also be used to give the title to specialised lists:
<listEvent>  <head xml:lang="nl">Zittingsperiode</head>  <head xml:lang="en">Legislative period</head>  <event to="2007-05-02from="2003-06-05" @@ -1803,11 +1821,11 @@  </event> ... -</listEvent>
Content model
+</listEvent>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element head
 {
    tei_att.global.attribute.xmlid,
@@ -1815,35 +1833,39 @@
    tei_att.global.linking.attribute.corresp,
    tei_att.typed.attribute.type,
    text
-}

Appendix A.1.37 <hyphenation>

<hyphenation> (hyphenation) summarizes the way in which hyphenation in a source text has been treated in an encoded version of it. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> ... +}

Appendix A.1.37 <hyphenation>

<hyphenation> (hyphenation) summarizes the way in which hyphenation in a source text has been treated in an encoded version of it. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> ... <hyphenation>   <p xml:lang="en">No end-of-line hyphens were present in the source.</p>  </hyphenation> ... -</editorialDecl>
Content model
+</editorialDecl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:hyphenation"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element hyphenation { tei_p+ }

Appendix A.1.38 <idno>

<idno> (identifier) supplies an identifier used to identify some object, such as a person or organisation. If it is a URL, it should have @type="URI". [14.3.1. Basic Principles 2.2.4. Publication, Distribution, Licensing, etc. 2.2.5. The Series Statement 3.12.2.4. Imprint, Size of a Document, and Reprint Information]
Moduleheader — Formal specification
Attributes
typecategorizes the identifier.
StatusRequired
Legal values are:
URI
Uniform Resource Identifier ParlaMint should be a resolvable URL, with the subtype classifying the type of web site.
VIAF
The URL of the Virtual Internet Authority File assigned to link different names in catalogs around the world for the same entity.
subtype
StatusOptional
Legal values are:
handle
The permanent identifier of type handle.
government
A governmental web site.
politicalParty
The web site of a political party.
parliament
A web site of the parliament.
ministry
The web site of a ministry.
personal
The personal web site of a person.
business
A web site belonging to a bussiness.
publicService
The web site of a pubic service.
wikimedia
A web site of Wikimedia, e.g. Wikipedia.
facebook
A Facebook web site.
twitter
A Twitter web site.
tiktok
A TikTok web site.
instagram
An Instagram web site.
Note

this attribute should always be used with type="URI"

Member of
Contained by
core: bibl
namesdates: org person
May containCharacter data only
Note

<idno> should be used for labels which identify an object or concept in a formal cataloguing system such as a database or an RDF store, or in a distributed system such as the World Wide Web. Some suggested values for type on <idno> are ISBN, ISSN, DOI, and URI.

Example
<publicationStmt> ... +
Schema Declaration
+element hyphenation { tei_p+ }

Appendix A.1.38 <idno>

<idno> (identifier) supplies an identifier used to identify some object, such as a person or organisation. If it is a URL, it should have @type="URI". [14.3.1. Basic Principles 2.2.4. Publication, Distribution, Licensing, etc. 2.2.5. The Series Statement 3.12.2.4. Imprint, Size of a Document, and Reprint Information]
Moduleheader — Formal specification
Attributes
typecategorizes the identifier.
StatusRequired
Legal values are:
URI
Uniform Resource Identifier ParlaMint should be a resolvable URL, with the subtype classifying the type of web site.
VIAF
The URL of the Virtual Internet Authority File assigned to link different names in catalogs around the world for the same entity.
subtype
StatusOptional
Legal values are:
handle
The permanent identifier of type handle.
government
A governmental web site.
politicalParty
The web site of a political party.
parliament
A web site of the parliament.
ministry
The web site of a ministry.
personal
The personal web site of a person.
business
A web site belonging to a bussiness.
publicService
The web site of a pubic service.
wikimedia
A web site of Wikimedia, e.g. Wikipedia.
facebook
A Facebook web site.
twitter
A Twitter web site.
tiktok
A TikTok web site.
instagram
An Instagram web site.
Note

this attribute should always be used with type="URI"

Member of
Contained by
core: bibl
namesdates: org person
May containCharacter data only
Note

<idno> should be used for labels which identify an object or concept in a formal cataloguing system such as a database or an RDF store, or in a distributed system such as the World Wide Web. Some suggested values for type on <idno> are ISBN, ISSN, DOI, and URI.

Example
<publicationStmt> ... <idno type="URIsubtype="handle">http://hdl.handle.net/11356/1432</idno> ... -</publicationStmt>
Example
<sourceDesc> +</publicationStmt>
Example
<sourceDesc>  <bibl>   <title type="mainxml:lang="sl">Zapisi sej Državnega zbora Republike Slovenije</title>    ...  <idno type="URI">https://www.dz-rs.si</idno>    ...  </bibl> -</sourceDesc>
Example
<idno type="URIsubtype="wikimedia" +</sourceDesc>
Example
<idno type="URIsubtype="wikimedia"  xml:lang="sl">https://sl.wikipedia.org/wiki/Pozitivna_Slovenija</idno> <idno type="URIsubtype="wikimedia" - xml:lang="en">https://en.wikipedia.org/wiki/Positive_Slovenia</idno>
Content model
xml:lang="en">https://en.wikipedia.org/wiki/Positive_Slovenia</idno>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element idno
 {
    tei_att.global.attribute.xmllang,
@@ -1865,18 +1887,18 @@
     | "instagram"
    }?,
    text
-}

Appendix A.1.39 <incident>

<incident> (incident) marks any phenomenon or occurrence, not necessarily vocalized or communicative, for example incidental noises or other events affecting communication. [8.3.3. Vocal, Kinesic, Incident]
Modulespoken — Formal specification
Attributes
type
StatusRecommended
Legal values are:
action
incident
leaving
entering
break
pause
sound
editorial
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Example
<incident type="action"> +}

Appendix A.1.39 <incident>

<incident> (incident) marks any phenomenon or occurrence, not necessarily vocalized or communicative, for example incidental noises or other events affecting communication. [8.3.3. Vocal, Kinesic, Incident]
Modulespoken — Formal specification
Attributes
type
StatusRecommended
Legal values are:
action
incident
leaving
entering
break
pause
sound
editorial
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Example
<incident type="action">  <desc>He stands and with him the whole Assembly</desc> -</incident>
Example
<incident type="sound"> +</incident>
Example
<incident type="sound">  <desc>The Assembly observed a minute of silence. Applause.</desc> -</incident>
Example
<incident type="entering"> +</incident>
Example
<incident type="entering">  <desc>Arrival of the President of the Republic of Poland</desc> -</incident>
Content model
+</incident>
Content model
 <content>
  <elementRef key="desc" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element incident
 {
    tei_att.global.attribute.xmlid,
@@ -1897,7 +1919,7 @@
     | "editorial"
    }?,
    tei_desc+
-}

Appendix A.1.40 <include>

<include> is an element from the XML namespace of the XML Inclusions (XInclude) W3C recommendation. It is used to include, into a ParlaMint <teiCorpus> root file the elements of the corpus that are stored as separate files. These are the <TEI> corpus components and parts of the corpus root <teiHeader>. Inside <particDesc> these are <listPerson> & <listOrg>, and <taxonomy> inside <classDecl>.
Namespacehttp://www.w3.org/2001/XInclude
Modulederived-module-parlamint
Attributes
href
StatusOptional
Datatypeteidata.pointer
Contained by
core: teiCorpus
corpus: particDesc
header: classDecl
May containEmpty element
ExampleUsing XInclude in ParlaMint to include corpus components into the corpus root:
<teiCorpus xml:lang="en" +}

Appendix A.1.40 <include>

<include> is an element from the XML namespace of the XML Inclusions (XInclude) W3C recommendation. It is used to include, into a ParlaMint <teiCorpus> root file the elements of the corpus that are stored as separate files. These are the <TEI> corpus components and parts of the corpus root <teiHeader>. Inside <particDesc> these are <listPerson> & <listOrg>, and <taxonomy> inside <classDecl>.
Namespacehttp://www.w3.org/2001/XInclude
Modulederived-module-parlamint
Attributes
href
StatusOptional
Datatypeteidata.pointer
Contained by
core: teiCorpus
corpus: particDesc
header: classDecl
May containEmpty element
ExampleUsing XInclude in ParlaMint to include corpus components into the corpus root:
<teiCorpus xml:lang="en"  xml:id="ParlaMint-GB" xmlns="http://www.tei-c.org/ns/1.0">  <teiHeader> ...TEI header of the corpus...  </teiHeader> @@ -1907,18 +1929,18 @@ href="2015/ParlaMint-GB_2015-01-06-commons.xml"/> ... -</teiCorpus>

Appendix A.1.41 <kinesic>

<kinesic> (kinesic) marks any communicative phenomenon, not necessarily vocalized, for example a gesture, frown, etc. [8.3.3. Vocal, Kinesic, Incident]
Modulespoken — Formal specification
Attributes
type
StatusRecommended
Legal values are:
kinesic
applause
ringing
signal
playback
gesture
smiling
laughter
snapping
noise
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Example
<kinesic type="signal"> +</teiCorpus>

Appendix A.1.41 <kinesic>

<kinesic> (kinesic) marks any communicative phenomenon, not necessarily vocalized, for example a gesture, frown, etc. [8.3.3. Vocal, Kinesic, Incident]
Modulespoken — Formal specification
Attributes
type
StatusRecommended
Legal values are:
kinesic
applause
ringing
signal
playback
gesture
smiling
laughter
snapping
noise
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Example
<kinesic type="signal">  <desc>sign for the end of discussion</desc> -</kinesic>
Example
<kinesic type="laughter"> +</kinesic>
Example
<kinesic type="laughter">  <desc xml:lang="hr">smijeh.</desc> -</kinesic>
Example
<kinesic type="applause"> +</kinesic>
Example
<kinesic type="applause">  <desc xml:lang="sl">ploskanje</desc> -</kinesic>
Content model
+</kinesic>
Content model
 <content>
  <elementRef key="desc" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element kinesic
 {
    tei_att.global.attribute.xmlid,
@@ -1941,7 +1963,7 @@
     | "noise"
    }?,
    tei_desc+
-}

Appendix A.1.42 <label>

<label> (label) contains any label or heading used to identify part of a text, typically but not exclusively in a list or glossary. [3.8. Lists]
Modulecore — Formal specification
Attributes
Member of
Contained by
header: application
namesdates: event
May contain
namesdates: orgName
character data
ExampleLabels denote the existence of organisations and connected events:
<org xml:id="DZrole="parliament" +}

Appendix A.1.42 <label>

<label> (label) contains any label or heading used to identify part of a text, typically but not exclusively in a list or glossary. [3.8. Lists]
Modulecore — Formal specification
Attributes
Member of
Contained by
header: application
namesdates: event
May contain
namesdates: orgName
character data
ExampleLabels denote the existence of organisations and connected events:
<org xml:id="DZrole="parliament"  ana="#parla.national #parla.lower">  <orgName xml:lang="slfull="yes">Državni zbor Republike Slovenije</orgName>  <orgName xml:lang="enfull="yes">National Assembly of the Republic of Slovenia</orgName> @@ -1962,11 +1984,11 @@    <label xml:lang="en">Term 8</label>   </event>  </listEvent> -</org>
ExampleLabels may also be used to give a name to the tools used in compiling the corpus:
<application ident="int-tagger" +</org>
ExampleLabels may also be used to give a name to the tools used in compiling the corpus:
<application ident="int-tagger"  version="1.0">  <label>INT Tagger, lemmatizer and Tokenizer</label>  <desc xml:lang="en">INT Tagger, lemmatizer and Tokenizer for modern Dutch, based on old-school machine learning (SVM). It provides the legacy PoS tags (encoded in w/@ana) and the lemmata for Dutch. Not publicly available.</desc> -</application>
ExampleLabels may also be used for other structured list items:
<listEvent> +</application>
ExampleLabels may also be used for other structured list items:
<listEvent>  <head xml:lang="lv">Saeimas sasaukumi</head>  <head xml:lang="en">Legislative period</head>  <event xml:id="PT.12from="2014-11-04" @@ -1978,29 +2000,32 @@   <label xml:lang="lv">13. Saeima</label>   <label xml:lang="en">Term 13</label>  </event> -</listEvent>
Content model
+</listEvent>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <textNode/>
   <elementRef key="orgName"/>
  </alternate>
 </content>
-    
Schema Declaration
-element label { tei_att.global.attribute.xmllang, ( text | tei_orgName ) }

Appendix A.1.43 <langUsage>

<langUsage> (language usage) describes the languages, sublanguages, registers, dialects, etc. represented within a text. [2.4.2. Language Usage 2.4. The Profile Description 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
header: profileDesc
May contain
header: language
Example
<langUsage> +
Schema Declaration
+element label { tei_att.global.attribute.xmllang, ( text | tei_orgName ) }

Appendix A.1.43 <langUsage>

<langUsage> (language usage) describes the languages, sublanguages, registers, dialects, etc. represented within a text. [2.4.2. Language Usage 2.4. The Profile Description 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
header: profileDesc
May contain
header: language
Example
<langUsage>  <language ident="slxml:lang="sl">slovenski</language>  <language ident="enxml:lang="sl">angleški</language>  <language ident="slxml:lang="en">Slovenian</language>  <language ident="enxml:lang="en">English</language> -</langUsage>
Content model
+</langUsage>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:langUsage"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="language" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element langUsage { tei_language+ }

Appendix A.1.44 <language>

<language> (language) characterizes a single language or sublanguage used within a text. [2.4.2. Language Usage]
Moduleheader — Formal specification
Attributes
ident(identifier) Supplies a language code constructed as defined in BCP 47 which is used to identify the language documented by this element, and which may be referenced by the global xml:lang attribute.
StatusRequired
Datatypeteidata.language
usagespecifies the approximate percentage of the text which uses this language.
StatusOptional
DatatypenonNegativeInteger
Contained by
header: langUsage
May containCharacter data only
Note

Particularly for sublanguages, an informal prose characterization should be supplied as content for the element.

Example
<langUsage> +
Schema Declaration
+element langUsage { tei_language+ }

Appendix A.1.44 <language>

<language> (language) characterizes a single language or sublanguage used within a text. [2.4.2. Language Usage]
Moduleheader — Formal specification
Attributes
ident(identifier) Supplies a language code constructed as defined in BCP 47 which is used to identify the language documented by this element, and which may be referenced by the global xml:lang attribute.
StatusRequired
Datatypeteidata.language
usagespecifies the approximate percentage of the text which uses this language.
StatusOptional
DatatypenonNegativeInteger
Contained by
header: langUsage
May containCharacter data only
Note

Particularly for sublanguages, an informal prose characterization should be supplied as content for the element.

Example
<langUsage>  <language ident="esxml:lang="es">Español</language>  <language ident="esxml:lang="en">Spanish</language> -</langUsage>
Example
<langUsage> +</langUsage>
Example
<langUsage>  <language ident="bg-Latnxml:lang="en">Bulgarian in Latin script</language>  <language ident="bgxml:lang="bg">български</language>  <language ident="bgxml:lang="en">Bulgarian</language> @@ -2008,29 +2033,29 @@  <language ident="enxml:lang="en">English</language>  <language ident="frxml:lang="bg">френски</language>  <language ident="frxml:lang="en">French</language> -</langUsage>
Content model
+</langUsage>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element language
 {
    tei_att.global.attribute.xmllang,
    attribute ident { text },
    attribute usage { text }?,
    text
-}

Appendix A.1.45 <licence>

<licence> contains information about a licence or other legal agreement applicable to the text. [2.2.4. Publication, Distribution, Licensing, etc.]
Moduleheader — Formal specification
Contained by
header: availability
May contain
XSD anyURI
Note

A <licence> element should be supplied for each licence agreement applicable to the text in question. The target attribute may be used to reference a full version of the licence. The when, notBefore, notAfter, from or to attributes may be used in combination to indicate the date or dates of applicability of the licence.

ExampleThe <licence> specifies fixed-value CC BY 4.0 URL, and in the following paragraph gives a prose description of the licence:
<licence>http://creativecommons.org/licenses/by/4.0/</licence> +}

Appendix A.1.45 <licence>

<licence> contains information about a licence or other legal agreement applicable to the text. [2.2.4. Publication, Distribution, Licensing, etc.]
Moduleheader — Formal specification
Contained by
header: availability
May contain
XSD anyURI
Note

A <licence> element should be supplied for each licence agreement applicable to the text in question. The target attribute may be used to reference a full version of the licence. The when, notBefore, notAfter, from or to attributes may be used in combination to indicate the date or dates of applicability of the licence.

ExampleThe <licence> specifies fixed-value CC BY 4.0 URL, and in the following paragraph gives a prose description of the licence:
<licence>http://creativecommons.org/licenses/by/4.0/</licence> <p>This work is licensed under the <ref target="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ref> -</p>
ExampleThe textual information on licence can be given in more than one language:
<licence>http://creativecommons.org/licenses/by/4.0/</licence> +</p>
ExampleThe textual information on licence can be given in more than one language:
<licence>http://creativecommons.org/licenses/by/4.0/</licence> <p xml:lang="hr">Ovaj rad je dostupan pod <ref target="http://creativecommons.org/licenses/by/4.0/">međunarodnom licencom Creative Commons Imenovanje 4.0</ref> </p> <p xml:lang="en">This work is licensed under the <ref target="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ref> -</p>
Content model
+</p>
Content model
 <content>
  <dataRef name="anyURI"/>
 </content>
-    
Schema Declaration
-element licence { xsd:anyURI }

Appendix A.1.47 <linkGrp>

<linkGrp> (link group) defines a collection of associations or hypertextual links. [17.1. Links]
Modulelinking — Formal specification
Attributes
targFunc
StatusRequired
Legal values are:
head argument
type
StatusRequired
Legal values are:
UD-SYN
Member of
Contained by
analysis: s
May contain
linking: link
Note

May contain one or more <link> or <ptr> elements.

A web or link group is an administrative convenience, which should be used to collect a set of links together for any purpose, not simply to supply a default value for the type attribute.

ExampleSyntactic analysis is stored in the link group, <linkGrp> element, which is then composed of <link> elements. The example below illustrating this is given, for readability, without the word-level linguistic attributes and with shortened IDs:
<s xml:id="ParlaMint-GB_2021-01-06.seg393.8"> +
Schema Declaration
+element link { attribute ana { text }, attribute target { list { ? } }, empty }

Appendix A.1.47 <linkGrp>

<linkGrp> (link group) defines a collection of associations or hypertextual links. [17.1. Links]
Modulelinking — Formal specification
Attributes
targFunc
StatusRequired
Legal values are:
head argument
type
StatusRequired
Legal values are:
UD-SYN
Member of
Contained by
analysis: s
May contain
linking: link
Note

May contain one or more <link> or <ptr> elements.

A web or link group is an administrative convenience, which should be used to collect a set of links together for any purpose, not simply to supply a default value for the type attribute.

ExampleSyntactic analysis is stored in the link group, <linkGrp> element, which is then composed of <link> elements. The example below illustrating this is given, for readability, without the word-level linguistic attributes and with shortened IDs:
<s xml:id="ParlaMint-GB_2021-01-06.seg393.8">  <w xml:id="ParlaMint-GB_2021-01-06.seg393.8.1">I</w>  <w xml:id="ParlaMint-GB_2021-01-06.seg393.8.2">support</w>  <w xml:id="ParlaMint-GB_2021-01-06.seg393.8.3">the</w> @@ -2079,18 +2104,18 @@   <link ana="ud-syn:punct"    target="#ParlaMint-GB_2021-01-06.seg393.8.2 #ParlaMint-GB_2021-01-06.seg393.8.5"/>  </linkGrp> -</s>
Content model
+</s>
Content model
 <content>
  <elementRef maxOccurs="unbounded"
   key="link"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element linkGrp
 {
    attribute targFunc { "head argument" },
    attribute type { "UD-SYN" },
    tei_link+
-}

Appendix A.1.48 <listEvent>

<listEvent> (list of events) contains a list of descriptions, each of which provides information about an identifiable event. [14.3.1. Basic Principles]
Modulenamesdates — Formal specification
Member of
Contained by
namesdates: org
May contain
core: head
namesdates: event
Example
<listEvent> +}

Appendix A.1.48 <listEvent>

<listEvent> (list of events) contains a list of descriptions, each of which provides information about an identifiable event. [14.3.1. Basic Principles]
Modulenamesdates — Formal specification
Member of
Contained by
namesdates: org
May contain
core: head
namesdates: event
Example
<listEvent>  <event xml:id="GOV.11from="2013-03-20"   to="2014-09-18">   <label xml:lang="sl">11. vlada Republike Slovenije (20. marec 2013 - 18. september 2014)</label> @@ -2101,7 +2126,7 @@   <label xml:lang="sl">14. vlada Republike Slovenije (13. marec 2020 - danes)</label>   <label xml:lang="en">14th Government of the Republic of Slovenia (March 13, 2020 - today)</label>  </event> -</listEvent>
Example
<org ana="#parla.national #parla.upper" +</listEvent>
Example
<org ana="#parla.national #parla.upper"  role="parliamentxml:id="LEG">  <orgName full="yesxml:lang="it">Senato della Repubblica Italiana</orgName>  <orgName full="yesxml:lang="it">Senate of the Republic of Italy</orgName> @@ -2117,7 +2142,10 @@    <label xml:lang="en">XVIII Legislative Term</label>   </event>  </listEvent> -</org>
Content model
+</org>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:listEvent"/> +</sch:pattern>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="head" minOccurs="0"
@@ -2126,8 +2154,8 @@
    maxOccurs="unbounded"/>
  </sequence>
 </content>
-    
Schema Declaration
-element listEvent { tei_head*, tei_event* }

Appendix A.1.49 <listOrg>

<listOrg> (list of organizations) contains a list of elements, each of which provides information about an identifiable organisation. [14.2.2. Organizational Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
corpus: particDesc
May contain
core: head
namesdates: listRelation org
Note

The type attribute may be used to distinguish lists of organizations of a particular type if convenient.

Example
<listOrg> +
Schema Declaration
+element listEvent { tei_head*, tei_event* }

Appendix A.1.49 <listOrg>

<listOrg> (list of organizations) contains a list of elements, each of which provides information about an identifiable organisation. [14.2.2. Organizational Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
corpus: particDesc
May contain
core: head
namesdates: listRelation org
Note

The type attribute may be used to distinguish lists of organizations of a particular type if convenient.

Example
<listOrg>  <org xml:id="government.GB"   role="government"> ...  </org> @@ -2139,7 +2167,10 @@ ... <listRelation> ...  </listRelation> -</listOrg>
Content model
+</listOrg>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:listOrg"/> +</sch:pattern>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="head" minOccurs="0"
@@ -2150,13 +2181,13 @@
    minOccurs="0" maxOccurs="1"/>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element listOrg
 {
    tei_att.global.attribute.xmlid,
    tei_att.global.attribute.xmllang,
    ( tei_head*, tei_org+, tei_listRelation? )
-}

Appendix A.1.50 <listPerson>

<listPerson> (list of persons) contains a list of descriptions, each of which provides information about an identifiable person or a group of people, for example the participants in a language interaction, or the people referred to in a historical source. [14.3.2. The Person Element 16.2. Contextual Information 2.4. The Profile Description 16.3.2. Declarable Elements]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
corpus: particDesc
May contain
core: head
namesdates: person
Note

The type attribute may be used to distinguish lists of people of a particular type if convenient.

Example
<listPerson> +}

Appendix A.1.50 <listPerson>

<listPerson> (list of persons) contains a list of descriptions, each of which provides information about an identifiable person or a group of people, for example the participants in a language interaction, or the people referred to in a historical source. [14.3.2. The Person Element 16.2. Contextual Information 2.4. The Profile Description 16.3.2. Declarable Elements]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
corpus: particDesc
May contain
core: head
namesdates: person
Note

The type attribute may be used to distinguish lists of people of a particular type if convenient.

Example
<listPerson>  <head>List of speakers</head>  <person xml:id="SayeedaWarsi"> ...  </person> @@ -2164,7 +2195,11 @@  </person> ... -</listPerson>
Content model
+</listPerson>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:listPerson"/> +</sch:pattern>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="head" minOccurs="0"
@@ -2173,13 +2208,13 @@
    maxOccurs="unbounded"/>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element listPerson
 {
    tei_att.global.attribute.xmlid,
    tei_att.global.attribute.xmllang,
    ( tei_head*, tei_person+ )
-}

Appendix A.1.51 <listPrefixDef>

<listPrefixDef> (list of prefix definitions) contains a list of definitions of prefixing schemes used in teidata.pointer values, showing how abbreviated URIs using each scheme may be expanded into full URIs. [17.2.3. Using Abbreviated Pointers]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
header: prefixDef
ExampleIn this example, two private URI scheme prefixes are defined and patterns are provided for dereferencing them. Each prefix is also supplied with a human-readable explanation in a <p> element.
<listPrefixDef> +}

Appendix A.1.51 <listPrefixDef>

<listPrefixDef> (list of prefix definitions) contains a list of definitions of prefixing schemes used in teidata.pointer values, showing how abbreviated URIs using each scheme may be expanded into full URIs. [17.2.3. Using Abbreviated Pointers]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
header: prefixDef
ExampleIn this example, two private URI scheme prefixes are defined and patterns are provided for dereferencing them. Each prefix is also supplied with a human-readable explanation in a <p> element.
<listPrefixDef>  <prefixDef ident="ud-syn"   matchPattern="(.+)replacementPattern="#$1">   <p>Private URIs with this prefix point to elements giving their name. In this document they are simply local references into the UD-SYN taxonomy categories in the corpus root TEI header.</p> @@ -2188,13 +2223,13 @@   replacementPattern="#NER.cnec2.0.$1">   <p>Taxonomy for named entities (cnec2.0)</p>  </prefixDef> -</listPrefixDef>
Content model
+</listPrefixDef>
Content model
 <content>
  <elementRef key="prefixDef" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element listPrefixDef { tei_prefixDef+ }

Appendix A.1.52 <listRelation>

<listRelation> provides information about relationships identified amongst people, places, and organisations, either informally as prose or as formally expressed relation links. [14.3.2.3. Personal Relationships]
Modulenamesdates — Formal specification
Member of
Contained by
namesdates: listOrg
May contain
namesdates: relation
Note

May contain a prose description organized as paragraphs, or a sequence of <relation> elements.

Example
<listOrg> +
Schema Declaration
+element listPrefixDef { tei_prefixDef+ }

Appendix A.1.52 <listRelation>

<listRelation> provides information about relationships identified amongst people, places, and organisations, either informally as prose or as formally expressed relation links. [14.3.2.3. Personal Relationships]
Modulenamesdates — Formal specification
Member of
Contained by
namesdates: listOrg
May contain
namesdates: relation
Note

May contain a prose description organized as paragraphs, or a sequence of <relation> elements.

Example
<listOrg>  <org role="parliamentaryGroup"   xml:id="party.LD">   <orgName full="yes">Liberal Democrat</orgName> @@ -2224,20 +2259,20 @@   <relation>...</relation>    ...  </listRelation> -</listOrg>
Content model
+</listOrg>
Content model
 <content>
  <elementRef key="relation" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element listRelation { tei_relation+ }

Appendix A.1.53 <measure>

<measure> (measure) either gives (in teiHeader//extent) the number of occurences of certain items (typicaly elements) in the corpus or corpus component or the score of an annotation. ParlaMint currently uses it for giving the sentiment score of utterrances and sentences. In this case measure should have empty content. [3.6.3. Numbers and Measures]
Modulecore — Formal specification
Attributes
unit
StatusOptional
Legal values are:
speeches
words
tokens
optional value
quantity(quantity) specifies the number of the specified units that comprise the measurement
Derived fromatt.measurement
StatusRequired
Datatypeteidata.numeric
typespecifies the type of measurement in any convenient typology.
Derived fromatt.typed
StatusOptional
Datatypeteidata.enumerated
Member of
Contained by
analysis: s
header: extent
linking: seg
spoken: u
May containCharacter data only
Example
<measure unit="speechesquantity="75122" +
Schema Declaration
+element listRelation { tei_relation+ }

Appendix A.1.53 <measure>

<measure> (measure) either gives (in teiHeader//extent) the number of occurences of certain items (typicaly elements) in the corpus or corpus component or the score of an annotation. ParlaMint currently uses it for giving the sentiment score of utterrances and sentences. In this case measure should have empty content. [3.6.3. Numbers and Measures]
Modulecore — Formal specification
Attributes
unit
StatusOptional
Legal values are:
speeches
words
tokens
optional value
quantity(quantity) specifies the number of the specified units that comprise the measurement
Derived fromatt.measurement
StatusRequired
Datatypeteidata.numeric
typespecifies the type of measurement in any convenient typology.
Derived fromatt.typed
StatusOptional
Datatypeteidata.enumerated
Member of
Contained by
analysis: s
header: extent
linking: seg
spoken: u
May containCharacter data only
Example
<measure unit="speechesquantity="75122"  xml:lang="sl">75.122 govorov</measure> <measure unit="speechesquantity="75122"  xml:lang="en">75,122 speeches</measure> <measure unit="wordsquantity="20190034"  xml:lang="sl">20.190.034 besed</measure> <measure unit="wordsquantity="20190034" - xml:lang="en">20,190,034 words</measure>
ExampleSentiment score of a sentence:
<s xml:id="ParlaMint-SI_2000-10-27-SDZ3-Redna-01.ana.seg1.2">xml:lang="en">20,190,034 words</measure>
ExampleSentiment score of a sentence:
<s xml:id="ParlaMint-SI_2000-10-27-SDZ3-Redna-01.ana.seg1.2">  <measure type="sentimentquantity="4.1"   ana="senti:mixpos"   corresp="#ParlaMint-SI_2000-10-27-SDZ3-Redna-01.ana.seg1.2"/> @@ -2245,11 +2280,11 @@   msd="UPosTag=ADJ|Case=Nom|Degree=Pos|Gender=Fem|Number=Plur|VerbForm=Partana="mte:Appfpnlemma="spoštovan">Spoštovane</w> ... -</s>
Content model
+</s>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element measure
 {
    tei_att.global.attribute.xmllang,
@@ -2259,7 +2294,7 @@
    attribute quantity { text },
    attribute type { text }?,
    text
-}

Appendix A.1.54 <media>

<media> indicates the location of any form of external media such as an audio or video clip etc. [3.10. Graphics and Other Non-textual Components]
Modulecore — Formal specification
Attributes
mimeType(MIME media type) specifies the applicable multimedia internet mail extension (MIME) media type.
Derived fromatt.internetMedia
StatusRequired
Datatype1–∞ occurrences of teidata.word separated by whitespace
Member of
Contained by
spoken: recording
May containEmpty element
Note

The attributes available for this element are not appropriate in all cases. For example, it makes no sense to specify the temporal duration of a graphic. Such errors are not currently detected.

The mimeType attribute must be used to specify the MIME media type of the resource specified by the url attribute.

Example
<recording type="audio"> +}

Appendix A.1.54 <media>

<media> indicates the location of any form of external media such as an audio or video clip etc. [3.10. Graphics and Other Non-textual Components]
Modulecore — Formal specification
Attributes
mimeType(MIME media type) specifies the applicable multimedia internet mail extension (MIME) media type.
Derived fromatt.internetMedia
StatusRequired
Datatype1–∞ occurrences of teidata.word separated by whitespace
Member of
Contained by
spoken: recording
May containEmpty element
Note

The attributes available for this element are not appropriate in all cases. For example, it makes no sense to specify the temporal duration of a graphic. Such errors are not currently detected.

The mimeType attribute must be used to specify the MIME media type of the resource specified by the url attribute.

Example
<recording type="audio">  <media xml:id="ps2013-009-01-001-001.audio1"   mimeType="audio/mp3"   source="https://www.psp.cz/eknih/2013ps/audio/2014/05/07/2014050713581412.mp3" @@ -2274,11 +2309,11 @@   url="2013ps/audio/2014/05/07/2014050714181432.mp3"/> ... -</recording>
Content model
+</recording>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element media
 {
    tei_att.global.attribute.xmlid,
@@ -2286,14 +2321,14 @@
    tei_att.resourced.attributes,
    attribute mimeType { list { + } },
    empty
-}

Appendix A.1.55 <meeting>

<meeting> contains the formalized descriptive title for a meeting or conference, for use in a bibliographic description for an item derived from such a meeting, or as a heading or preamble to publications emanating from it. [3.12.2.2. Titles, Authors, and Editors]
Modulecore — Formal specification
Attributes
Contained by
header: titleStmt
May containCharacter data only
ExampleThe specification of the particular sessions that the corpus or corpus component contains are encoded with <meeting>:
<meeting n="7corresp="#DZ" +}

Appendix A.1.55 <meeting>

<meeting> contains the formalized descriptive title for a meeting or conference, for use in a bibliographic description for an item derived from such a meeting, or as a heading or preamble to publications emanating from it. [3.12.2.2. Titles, Authors, and Editors]
Modulecore — Formal specification
Attributes
Contained by
header: titleStmt
May containCharacter data only
ExampleThe specification of the particular sessions that the corpus or corpus component contains are encoded with <meeting>:
<meeting n="7corresp="#DZ"  ana="#parla.lower #parla.term #DZ.7">7. mandat</meeting> <meeting n="8corresp="#DZ" - ana="#parla.lower #parla.term #DZ.8">8. mandat</meeting>
Content model
ana="#parla.lower #parla.term #DZ.8">8. mandat</meeting>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element meeting
 {
    tei_att.global.attribute.n,
@@ -2301,7 +2336,7 @@
    tei_att.global.analytic.attribute.ana,
    tei_att.global.linking.attribute.corresp,
    text
-}

Appendix A.1.56 <name>

<name> (name, proper noun) contains a proper noun or noun phrase. [3.6.1. Referring Strings]
Modulecore — Formal specification
Attributes
type
StatusOptional
Legal values are:
PER
LOC
ORG
MISC
city
country
address
org
place
Member of
Contained by
analysis: s
core: name unit
corpus: setting
header: change
namesdates: placeName
May contain
analysis: pc w
character data
Note

Proper nouns referring to people, places, and organizations may be tagged instead with <persName>, <placeName>, or <orgName>, when the TEI module for names and dates is included.

ExampleThe element is used to mark up Named Entities in the linguistically analysed corpus, in which case it should have the type attribute with one of the allowed values. It can also have a ref attribute to link it a definition:
... +}

Appendix A.1.56 <name>

<name> (name, proper noun) contains a proper noun or noun phrase. [3.6.1. Referring Strings]
Modulecore — Formal specification
Attributes
type
StatusOptional
Legal values are:
PER
LOC
ORG
MISC
city
country
address
org
place
Member of
Contained by
analysis: s
core: name unit
corpus: setting
header: change
namesdates: placeName
May contain
analysis: pc w
character data
Note

Proper nouns referring to people, places, and organizations may be tagged instead with <persName>, <placeName>, or <orgName>, when the TEI module for names and dates is included.

ExampleThe element is used to mark up Named Entities in the linguistically analysed corpus, in which case it should have the type attribute with one of the allowed values. It can also have a ref attribute to link it a definition:
... <w lemma="andmsd="UPosTag=CCONJ">and</w> <name type="ORG"  ref="https://en.wikipedia.org/wiki/Westminster"> @@ -2310,14 +2345,14 @@ </name> <w lemma=",msd="UPosTag=PUNCT">,</w> ... -
ExampleElement <name> is used in the TEI header to specify the location of the parliament:
<name type="place">Westminster</name> +
ExampleElement <name> is used in the TEI header to specify the location of the parliament:
<name type="place">Westminster</name> <name type="city">London</name> -<name type="countrykey="GB">U.K.</name>
ExampleThe element is used in the TEI header to denote person's responsibility for changes:
<revisionDesc> +<name type="countrykey="GB">U.K.</name>
ExampleThe element is used in the TEI header to denote person's responsibility for changes:
<revisionDesc>  <change when="2021-06-11">   <name>Tomaž Erjavec</name>: Finalized encoding.</change>  <change when="2021-05-28">   <name>Tomaž Erjavec</name>: Built corpus.</change> -</revisionDesc>
Content model
+</revisionDesc>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -2336,7 +2371,7 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element name
 {
    tei_att.global.attribute.xmlid,
@@ -2373,7 +2408,7 @@
     | tei_pb
     | text
    )+
-}

Appendix A.1.58 <namespace>

<namespace> (namespace) supplies the formal name of the namespace to which the elements documented by its children belong. [2.3.4. The Tagging Declaration]
Moduleheader — Formal specification
Attributes
name
StatusRequired
Legal values are:
http://www.tei-c.org/ns/1.0
Contained by
header: tagsDecl
May contain
header: tagUsage
ExampleTo distinguish the TEI elements from the possible use of elements from other namespaces, a <namespace> element giving the TEI namespace is introduced first:
<tagsDecl> +
Schema Declaration
+element nameLink { tei_att.global.attribute.xmllang, text }

Appendix A.1.58 <namespace>

<namespace> (namespace) supplies the formal name of the namespace to which the elements documented by its children belong. [2.3.4. The Tagging Declaration]
Moduleheader — Formal specification
Attributes
name
StatusRequired
Legal values are:
http://www.tei-c.org/ns/1.0
Contained by
header: tagsDecl
May contain
header: tagUsage
ExampleTo distinguish the TEI elements from the possible use of elements from other namespaces, a <namespace> element giving the TEI namespace is introduced first:
<tagsDecl>  <namespace name="http://www.tei-c.org/ns/1.0">   <tagUsage gi="textoccurs="414"/>   <tagUsage gi="bodyoccurs="414"/> @@ -2402,28 +2437,32 @@   <tagUsage gi="kinesicoccurs="560"/>   <tagUsage gi="descoccurs="10234"/>  </namespace> -</tagsDecl>
Content model
+</tagsDecl>
Content model
 <content>
  <elementRef key="tagUsage" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element namespace
 {
    attribute name { "http://www.tei-c.org/ns/1.0" },
    tei_tagUsage+
-}

Appendix A.1.59 <normalization>

<normalization> (normalization) indicates the extent of normalization or regularization of the original source carried out in converting it to electronic form. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> ... +}

Appendix A.1.59 <normalization>

<normalization> (normalization) indicates the extent of normalization or regularization of the original source carried out in converting it to electronic form. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> ... <normalization>   <p xml:lang="en">Text has not been normalised, except for spacing.</p>  </normalization> ... -</editorialDecl>
Content model
+</editorialDecl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:normalization"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element normalization { tei_p+ }

Appendix A.1.60 <note>

<note> (note) contains a note or annotation. [3.9.1. Notes and Simple Annotation 2.2.6. The Notes Statement 3.12.2.8. Notes and Statement of Language 10.3.5.4. Notes within Entries]
Modulecore — Formal specification
Attributes
type
StatusRecommended
Sample values include:
narrative
Description in the third person of events taking place in the meeting, e.g. "Mr X. takes the Chair".
summary
Summaries of speeches that are individually not interesting, e.g. "Question put and agreed to".
speaker
Name, role and possible description of a person doing the speech
vote
Outcome of a vote
location
The location of the speaker, who was not on the podium
date
Date of the session
president
Chairman of a meeting
comment
Comment of parliamentary reporter
time
Date and time of the beginning and end of the debate
quorum
The presence of the members of parliament
debate
Comments on the conduct of debates
Member of
Contained by
analysis: s
core: name unit
linking: seg
namesdates: affiliation state
spoken: u
textstructure: div
May contain
core: pb time
character data
Example<note> element is used to encode transcriber comments such as who spoke, what the time was, interruptions, notes on what is happening in the chamber, results of voting etc.:
<note type="speaker">The president, Dr. Milan Brglez:</note> +
Schema Declaration
+element normalization { tei_p+ }

Appendix A.1.60 <note>

<note> (note) contains a note or annotation. [3.9.1. Notes and Simple Annotation 2.2.6. The Notes Statement 3.12.2.8. Notes and Statement of Language 10.3.5.4. Notes within Entries]
Modulecore — Formal specification
Attributes
type
StatusRecommended
Sample values include:
narrative
Description in the third person of events taking place in the meeting, e.g. "Mr X. takes the Chair".
summary
Summaries of speeches that are individually not interesting, e.g. "Question put and agreed to".
speaker
Name, role and possible description of a person doing the speech
vote
Outcome of a vote
location
The location of the speaker, who was not on the podium
date
Date of the session
president
Chairman of a meeting
comment
Comment of parliamentary reporter
time
Date and time of the beginning and end of the debate
quorum
The presence of the members of parliament
debate
Comments on the conduct of debates
Member of
Contained by
analysis: s
core: name unit
linking: seg
namesdates: affiliation state
spoken: u
textstructure: div
May contain
core: pb time
character data
Example<note> element is used to encode transcriber comments such as who spoke, what the time was, interruptions, notes on what is happening in the chamber, results of voting etc.:
<note type="speaker">The president, Dr. Milan Brglez:</note> ... <note type="time">The session began at 10 o'clock.</note> ... @@ -2431,13 +2470,13 @@ ... <note type="vote-noes">2 voted against the adoption of the measure.</note> ... -
ExampleThe <note> element can be further qualified by the <time> element to specify the date and time recorded in the note; and can also contain a page break, <pb>:
<note type="time">The session began <pb/> at <time when="2016-04-13T010:00:00">10 o'clock</time>.</note>
ExampleThe <note> element may also be used to mark any additional information on debate sections:
<div type="debateSection"> +
ExampleThe <note> element can be further qualified by the <time> element to specify the date and time recorded in the note; and can also contain a page break, <pb>:
<note type="time">The session began <pb/> at <time when="2016-04-13T010:00:00">10 o'clock</time>.</note>
ExampleThe <note> element may also be used to mark any additional information on debate sections:
<div type="debateSection">  <head>Business Before Questions</head>  <note>Death of a Member</note>  <u xml:id="ParlaMint-GB_2019-02-18-commons.u1">...</u> ... <note>End of debateSection.</note> -</div>
Content model
+</div>
Content model
 <content>
  <alternate minOccurs="0"
   maxOccurs="unbounded">
@@ -2446,7 +2485,7 @@
   <elementRef key="time"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element note
 {
    tei_att.global.attribute.xmlid,
@@ -2456,12 +2495,12 @@
    tei_att.typed.attribute.subtype,
    attribute type { text }?,
    ( text | tei_pb | tei_time )*
-}

Appendix A.1.61 <num>

<num> (number) contains a number, written in any form. [3.6.3. Numbers and Measures]
Modulecore — Formal specification
Attributes
typeindicates the type of numeric value.
Derived fromatt.typed
StatusOptional
Datatypeteidata.enumerated
Suggested values include:
cardinal
absolute number, e.g. 21, 21.5
ordinal
ordinal number, e.g. 21st
fraction
fraction, e.g. one half or three-quarters
percentage
a percentage
Note

If a different typology is desired, other values can be used for this attribute.

Member of
Contained by
analysis: s
core: name unit
May contain
analysis: pc w
character data
Note

Detailed analyses of quantities and units of measure in historical documents may also use the feature structure mechanism described in chapter 19. Feature Structures. The <num> element is intended for use in simple applications.

ExampleThe element can be used for fine-grained Named Entities which include numbers:
<num ana="ne:n_" +}

Appendix A.1.61 <num>

<num> (number) contains a number, written in any form. [3.6.3. Numbers and Measures]
Modulecore — Formal specification
Attributes
typeindicates the type of numeric value.
Derived fromatt.typed
StatusOptional
Datatypeteidata.enumerated
Suggested values include:
cardinal
absolute number, e.g. 21, 21.5
ordinal
ordinal number, e.g. 21st
fraction
fraction, e.g. one half or three-quarters
percentage
a percentage
Note

If a different typology is desired, other values can be used for this attribute.

Member of
Contained by
analysis: s
core: name unit
May contain
analysis: pc w
character data
Note

Detailed analyses of quantities and units of measure in historical documents may also use the feature structure mechanism described in chapter 19. Feature Structures. The <num> element is intended for use in simple applications.

ExampleThe element can be used for fine-grained Named Entities which include numbers:
<num ana="ne:n_"  xml:id="ParlaMint-CZ_2018-11-13-ps2017-020-09-004-010.ne138">  <w xml:id="ParlaMint-CZ_2018-11-13-ps2017-020-09-004-010.u6.p17.s3.w12"   lemma="428"   msd="UPosTag=NUM|NumForm=Digit|NumType=Cardjoin="right">428</w> -</num>
Content model
+</num>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -2470,7 +2509,7 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element num
 {
    tei_att.global.attribute.xmlid,
@@ -2479,7 +2518,7 @@
    tei_att.typed.attribute.subtype,
    attribute type { "cardinal" | "ordinal" | "fraction" | "percentage" }?,
    ( tei_w | tei_pc | text )+
-}

Appendix A.1.62 <occupation>

<occupation> (occupation) contains an informal description of a person's trade, profession or occupation. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: person
May containCharacter data only
Note

The content of this element may be used as an alternative to the more formal specification made possible by its attributes; it may also be used to supplement the formal specification with commentary or clarification.

Example
<person n="2678xml:id="SimeonovValeri"> +}

Appendix A.1.62 <occupation>

<occupation> (occupation) contains an informal description of a person's trade, profession or occupation. [16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: person
May containCharacter data only
Note

The content of this element may be used as an alternative to the more formal specification made possible by its attributes; it may also be used to supplement the formal specification with commentary or clarification.

Example
<person n="2678xml:id="SimeonovValeri">  <persName xml:lang="bg">   <forename>Валери</forename>   <surname>Симеонов</surname> @@ -2492,11 +2531,11 @@  <occupation>политик</occupation> ... -</person>
Content model
+</person>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element occupation
 {
    tei_att.global.attribute.xmllang,
@@ -2504,7 +2543,7 @@
    tei_att.datable.w3c.attribute.from,
    tei_att.datable.w3c.attribute.to,
    text
-}

Appendix A.1.63 <org>

<org> (organization) provides information about an identifiable organisation such as the government, political party, ministry etc. [14.3.3. Organizational Data]
Modulenamesdates — Formal specification
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
Derived fromatt.global
StatusRequired
DatatypeID
role
StatusRequired
Legal values are:
country
federatedState
republic
government
ministry
parliament
politicalParty
parliamentaryGroup
conferenceOfChairs
boardOfParliament
ngo
institution
senate
committee
subcommittee
commission
delegation
supervisoryBoard
workingGroup
interparliamentaryFriendshipGroup
nationalCouncil
chamberOfThePeople
chamberOfTheNations
europeanCommission
europeanParliament
europeanInstitution
internationalOrganisation
boardOfDirectors
ethnicCommunity
Contained by
namesdates: listOrg
May contain
core: desc head
header: idno
Example
<org xml:id="government.BE" +}

Appendix A.1.63 <org>

<org> (organization) provides information about an identifiable organisation such as the government, political party, ministry etc. [14.3.3. Organizational Data]
Modulenamesdates — Formal specification
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
Derived fromatt.global
StatusRequired
DatatypeID
role
StatusRequired
Legal values are:
country
federatedState
republic
government
ministry
parliament
politicalParty
parliamentaryGroup
conferenceOfChairs
boardOfParliament
ngo
institution
senate
committee
subcommittee
commission
delegation
supervisoryBoard
workingGroup
interparliamentaryFriendshipGroup
nationalCouncil
chamberOfThePeople
chamberOfTheNations
europeanCommission
europeanParliament
europeanInstitution
internationalOrganisation
boardOfDirectors
ethnicCommunity
Contained by
namesdates: listOrg
May contain
core: desc head
header: idno
Example
<org xml:id="government.BE"  role="government">  <orgName xml:lang="enfull="yes">Federal Government of Belgium</orgName>  <orgName xml:lang="nlfull="yes">Federale regering</orgName> @@ -2519,7 +2558,7 @@  </event> ... -</org>
Example
<org xml:id="party.PS2" +</org>
Example
<org xml:id="party.PS2"  role="parliamentaryGroup">  <orgName full="yesxml:lang="sl">Pozitivna Slovenija</orgName>  <orgName full="yesxml:lang="en">Positive Slovenia</orgName> @@ -2531,7 +2570,7 @@   subtype="wikimedia">https://sl.wikipedia.org/wiki/Pozitivna_Slovenija</idno>  <idno type="URIxml:lang="en"   subtype="wikimedia">https://en.wikipedia.org/wiki/Positive_Slovenia</idno> -</org>
Content model
+</org>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="head" minOccurs="0"
@@ -2550,7 +2589,7 @@
    maxOccurs="unbounded"/>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element org
 {
    tei_att.global.attribute.xmllang,
@@ -2597,19 +2636,19 @@
       tei_listEvent?,
       tei_state*
    )
-}

Appendix A.1.64 <orgName>

<orgName> (organization name) contains an organisational name. [14.2.2. Organizational Names]
Modulenamesdates — Formal specification
Attributes
fromindicates the starting point of the period in standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusOptional
Datatypeteidata.temporal.w3c
Note

Used when "the same" party changes its name

toindicates the ending point of the period in standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusOptional
Datatypeteidata.temporal.w3c
Note

Used when "the same" party changes its name

full
StatusOptional
Legal values are:
yes
abb
Member of
Contained by
header: funder
namesdates: affiliation org
May containCharacter data only
Example
<funder> +}

Appendix A.1.64 <orgName>

<orgName> (organization name) contains an organisational name. [14.2.2. Organizational Names]
Modulenamesdates — Formal specification
Attributes
fromindicates the starting point of the period in standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusOptional
Datatypeteidata.temporal.w3c
Note

Used when "the same" party changes its name

toindicates the ending point of the period in standard form, e.g. yyyy-mm-dd.
Derived fromatt.datable.w3c
StatusOptional
Datatypeteidata.temporal.w3c
Note

Used when "the same" party changes its name

full
StatusOptional
Legal values are:
yes
abb
Member of
Contained by
header: funder
namesdates: affiliation org
May containCharacter data only
Example
<funder>  <orgName xml:lang="en">The CLARIN research infrastructure</orgName>  <orgName xml:lang="sl">Raziskovalna infrastruktura CLARIN</orgName> -</funder>
Example
<org xml:id="party.PS1" +</funder>
Example
<org xml:id="party.PS1"  role="parliamentaryGroup">  <orgName full="yesxml:lang="en">Positive Slovenia</orgName>  <orgName full="yesxml:lang="sl">Pozitivna Slovenija</orgName>  <orgName full="abbxml:lang="sl">PS</orgName> -</org>
Content model
+</org>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element orgName
 {
    tei_att.global.attribute.xmllang,
@@ -2618,11 +2657,11 @@
    attribute to { text }?,
    attribute full { "yes" | "abb" }?,
    text
-}

Appendix A.1.65 <p>

<p> (paragraph) marks paragraphs in prose. [3.1. Paragraphs 7.2.5. Speech Contents]
Modulecore — Formal specification
Attributes
Member of
Contained by
May contain
core: ref
character data
Example
<projectDesc> +}

Appendix A.1.65 <p>

<p> (paragraph) marks paragraphs in prose. [3.1. Paragraphs 7.2.5. Speech Contents]
Modulecore — Formal specification
Attributes
Member of
Contained by
May contain
core: ref
character data
Example
<projectDesc>  <p>   <ref target="https://www.clarin.eu/content/parlamint">ParlaMint</ref>  </p> -</projectDesc>
Example
<availability status="free"> +</projectDesc>
Example
<availability status="free">  <licence>http://creativecommons.org/licenses/by/4.0/</licence>  <p>This work is licensed under the <ref target="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ref>.</p>  <p>This work is also licensed under the <ref target="https://www.parliament.uk/site-information/copyright-parliament/open-parliament-licence/">Open Parliament Licence v3.0</ref>.</p> @@ -2637,7 +2676,7 @@ <sch:rule context="tei:l//tei:p"> <sch:assert test="ancestor::tei:floatingText | parent::tei:figure | parent::tei:note"> Abstract model violation: Metrical lines may not contain higher-level structural elements such as div, p, or ab, unless p is a child of figure or note, or is a descendant of floatingText. </sch:assert> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -2645,12 +2684,16 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
-element p { tei_att.global.attribute.xmllang, ( tei_ref | text )+ }

Appendix A.1.66 <particDesc>

<particDesc> (participation description) describes the identifiable speakers and organisations in a ParlaMint corpus. This informations is given in the corpus root teiHeder. Note that the listPerson and listOrg elements are typically stored in separate files. [16.2. Contextual Information]
Modulecorpus — Formal specification
Contained by
header: profileDesc
May contain
derived-module-parlamint: include
namesdates: listOrg listPerson
Note

May contain a prose description organized as paragraphs, or a structured list of persons and person groups, with an optional formal specification of any relationships amongst them.

Example
<particDesc> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-listOrg.xml"/> +
Schema Declaration
+element p { tei_att.global.attribute.xmllang, ( tei_ref | text )+ }

Appendix A.1.66 <particDesc>

<particDesc> (participation description) describes the identifiable speakers and organisations in a ParlaMint corpus. This informations is given in the corpus root teiHeder. Note that the listPerson and listOrg elements are typically stored in separate files. [16.2. Contextual Information]
Modulecorpus — Formal specification
Contained by
header: profileDesc
May contain
derived-module-parlamint: include
namesdates: listOrg listPerson
Note

May contain a prose description organized as paragraphs, or a structured list of persons and person groups, with an optional formal specification of any relationships amongst them.

Example
<particDesc> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" + href="ParlaMint-SI-listOrg.xml"/> <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" - href="href="ParlaMint-SI-listPerson.xml"/> -</particDesc>
Content model
+ href="ParlaMint-SI-listPerson.xml"/>
+</particDesc>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:particDesc"/> +</sch:pattern>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <alternate minOccurs="1" maxOccurs="1">
@@ -2663,22 +2706,22 @@
   </alternate>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element particDesc
 {
    ( tei_listOrg | tei_include ), ( tei_listPerson | tei_include )
-}

Appendix A.1.67 <pb>

<pb> (page beginning) marks the beginning of a new page in a paginated document. [3.11.3. Milestone Elements]
Modulecore — Formal specification
Attributes
Member of
Contained by
analysis: phr s
core: name note
linking: seg
spoken: u
textstructure: div
May containEmpty element
Note

A <pb> element should appear at the start of the page which it identifies. The global n attribute indicates the number or other value associated with this page. This will normally be the page number or signature printed on it, since the physical sequence number is implicit in the presence of the <pb> element itself.

The type attribute may be used to characterize the page beginning in any respect. The more specialized attributes break, ed, or edRef should be preferred when the intent is to indicate whether or not the page beginning is word-breaking, or to note the source from which it derives.

Example
<body> +}

Appendix A.1.67 <pb>

<pb> (page beginning) marks the beginning of a new page in a paginated document. [3.11.3. Milestone Elements]
Modulecore — Formal specification
Attributes
Member of
Contained by
analysis: phr s
core: name note
linking: seg
spoken: u
textstructure: div
May containEmpty element
Note

A <pb> element should appear at the start of the page which it identifies. The global n attribute indicates the number or other value associated with this page. This will normally be the page number or signature printed on it, since the physical sequence number is implicit in the presence of the <pb> element itself.

The type attribute may be used to characterize the page beginning in any respect. The more specialized attributes break, ed, or edRef should be preferred when the intent is to indicate whether or not the page beginning is word-breaking, or to note the source from which it derives.

Example
<body>  <div type="debateSection">   <pb source="https://www.psp.cz/eknih/2013ps/stenprot/017schuz/s017357.htm"    n="1"    xml:id="ParlaMint-CZ_2014-10-01-ps2013-017-09-003-036.pb1corresp="#ps2013-017-09-003-036.audio1"/>    ...  </div> -</body>
Content model
+</body>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element pb
 {
    tei_att.global.attribute.xmlid,
@@ -2686,7 +2729,7 @@
    tei_att.global.linking.attribute.corresp,
    tei_att.global.source.attribute.source,
    empty
-}

Appendix A.1.68 <pc>

<pc> (punctuation character) contains a character or string of characters regarded as constituting a single punctuation mark. [18.1.2. Below the Word Level 18.4.2. Lightweight Linguistic Annotation]
Moduleanalysis — Formal specification
Attributes
xml:id
StatusRequired
DatatypeID
msd
StatusRequired
Datatypeteidata.text
Member of
Contained by
analysis: phr s
May containCharacter data only
Example
<s> +}

Appendix A.1.68 <pc>

<pc> (punctuation character) contains a character or string of characters regarded as constituting a single punctuation mark. [18.1.2. Below the Word Level 18.4.2. Lightweight Linguistic Annotation]
Moduleanalysis — Formal specification
Attributes
xml:id
StatusRequired
DatatypeID
msd
StatusRequired
Datatypeteidata.text
Member of
Contained by
analysis: phr s
May containCharacter data only
Example
<s>  <w lemma="I"   msd="UPosTag=PRON|Case=Nom|Number=Sing|Person=1|PronType=Prspos="PRP">I</w>  <w lemma="support" @@ -2696,11 +2739,11 @@  <w lemma="amendment"   msd="UPosTag=NOUN|Number=Singpos="NNjoin="right">amendment</w>  <pc msd="UPosTag=PUNCTpos=".">.</pc> -</s>
Content model
+</s>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element pc
 {
    tei_att.global.attribute.xmllang,
@@ -2712,13 +2755,13 @@
    attribute xml:id { text },
    attribute msd { text },
    text
-}

Appendix A.1.69 <persName>

<persName> (personal name) contains a proper noun or proper-noun phrase referring to a person, possibly including one or more of the person's forenames, surnames, honorifics, added names, etc. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
core: respStmt
namesdates: person
May contain
core: term
character data
Note

Special persons (like 'anonymous', 'group' etc.) have their name in <term>.

Example
<persName> +}

Appendix A.1.69 <persName>

<persName> (personal name) contains a proper noun or proper-noun phrase referring to a person, possibly including one or more of the person's forenames, surnames, honorifics, added names, etc. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
core: respStmt
namesdates: person
May contain
core: term
character data
Note

Special persons (like 'anonymous', 'group' etc.) have their name in <term>.

Example
<persName>  <surname>Broekers-Knol</surname>  <forename>Ankie</forename> -</persName>
Example
<respStmt> +</persName>
Example
<respStmt>  <persName>Matthew Coole</persName>  <resp>TEI corpus encoding</resp> -</respStmt>
Content model
+</respStmt>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <alternate minOccurs="1"
@@ -2743,7 +2786,7 @@
   </alternate>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element persName
 {
    tei_att.global.attribute.xmlid,
@@ -2762,7 +2805,7 @@
     | tei_term+
     | ( text )
    )
-}

Appendix A.1.70 <person>

<person> (person) provides information about a speaker in the corpus, at the very least their name and sex. [14.3.2. The Person Element 16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: listPerson
May contain
Note

May contain either a prose description organized as paragraphs, or a sequence of more specific demographic elements drawn from the model.personPart class.

Example
<person xml:id="AliciaKearns"> +}

Appendix A.1.70 <person>

<person> (person) provides information about a speaker in the corpus, at the very least their name and sex. [14.3.2. The Person Element 16.2.2. The Participant Description]
Modulenamesdates — Formal specification
Attributes
Contained by
namesdates: listPerson
May contain
Note

May contain either a prose description organized as paragraphs, or a sequence of more specific demographic elements drawn from the model.personPart class.

Example
<person xml:id="AliciaKearns">  <persName>   <forename>Alicia</forename>   <forename>Alexandra Martha</forename> @@ -2774,7 +2817,7 @@  <affiliation from="2019-12-12"   ref="#party.CONrole="member"/>  <idno subtype="contacttype="URI">https://members.parliament.uk/member/4805/contact</idno> -</person>
Example
<person xml:id="AdamowiczPiotr"> +</person>
Example
<person xml:id="AdamowiczPiotr">  <persName>   <forename>Piotr</forename>   <surname>Adamowicz</surname> @@ -2782,7 +2825,7 @@  <birth when="1961-06-26">26.06.1961</birth>  <sex value="M"/>  <affiliation role="memberref="#party.KO"/> -</person>
Content model
+</person>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="persName" minOccurs="1"
@@ -2808,7 +2851,7 @@
   </alternate>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element person
 {
    tei_att.global.attribute.xmlid,
@@ -2827,7 +2870,7 @@
        | tei_figure*
       )+
    )
-}

Appendix A.1.71 <phr>

<phr> (phrase) contains a semantic multi-word unit. [18.1. Linguistic Segment Categories]
Moduleanalysis — Formal specification
Attributes
ana(analysis) indicates one or more elements containing interpretations of the element on which the ana attribute appears.
Derived fromatt.global.analytic
StatusRequired
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
function(function) characterizes the function of the segment.
Derived fromatt.segLike
StatusRequired
Datatypeteidata.enumerated
type
StatusOptional
Legal values are:
sem
Member of
Contained by
analysis: s
May contain
analysis: pc w
core: pb
character data
Note

The type attribute may be used to indicate the type of phrase, taking values such as noun, verb, preposition, etc. as appropriate.

ExampleThe element is used to mark multi-word units (MWEs) which have a semantic interpretation. The type should be set to sem. The MWE should be marked with the function (all semantic tags) and ana (semantic categories) attributes:
... +}

Appendix A.1.71 <phr>

<phr> (phrase) contains a semantic multi-word unit. [18.1. Linguistic Segment Categories]
Moduleanalysis — Formal specification
Attributes
ana(analysis) indicates one or more elements containing interpretations of the element on which the ana attribute appears.
Derived fromatt.global.analytic
StatusRequired
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
function(function) characterizes the function of the segment.
Derived fromatt.segLike
StatusRequired
Datatypeteidata.enumerated
type
StatusOptional
Legal values are:
sem
Member of
Contained by
analysis: s
May contain
analysis: pc w
core: pb
character data
Note

The type attribute may be used to indicate the type of phrase, taking values such as noun, verb, preposition, etc. as appropriate.

ExampleThe element is used to mark multi-word units (MWEs) which have a semantic interpretation. The type should be set to sem. The MWE should be marked with the function (all semantic tags) and ana (semantic categories) attributes:
... ... <phr type="semfunction="Z4ana="sem:Z4">  <w pos="INmsd="UPosTag=ADPlemma="on" @@ -2840,7 +2883,7 @@   lemma="handfunction="Z4ana="sem:Z4join="right">hand</w> </phr> ... -
Content model
+
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -2850,7 +2893,7 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element phr
 {
    tei_att.global.attribute.xmlid,
@@ -2859,7 +2902,7 @@
    attribute function { text },
    attribute type { "sem" }?,
    ( tei_w | tei_pc | tei_pb | text )+
-}

Appendix A.1.72 <placeName>

<placeName> (place name) contains a place name. [14.2.3. Place Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: birth death
May contain
core: name
character data
Example
<placeName ref="https://www.geonames.org/2523918">Palermo</placeName>
Example
<placeName>Tours-Saint-Symphorien, Indre-et-Loire</placeName>
Content model
+}

Appendix A.1.72 <placeName>

<placeName> (place name) contains a place name. [14.2.3. Place Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: birth death
May contain
core: name
character data
Example
<placeName ref="https://www.geonames.org/2523918">Palermo</placeName>
Example
<placeName>Tours-Saint-Symphorien, Indre-et-Loire</placeName>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <elementRef key="name" minOccurs="0"
@@ -2867,33 +2910,33 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element placeName
 {
    tei_att.global.attribute.xmllang,
    tei_att.canonical.attribute.ref,
    ( tei_name? | text )
-}

Appendix A.1.73 <prefixDef>

<prefixDef> (prefix definition) defines a prefixing scheme used in teidata.pointer values, showing how abbreviated URIs using the scheme may be expanded into full URIs. [17.2.3. Using Abbreviated Pointers]
Moduleheader — Formal specification
Attributes
matchPatternspecifies a regular expression against which the values of other attributes can be matched.
Derived fromatt.patternReplacement
StatusRequired
Datatypeteidata.pattern
replacementPatternspecifies a ‘replacement pattern’, that is, the skeleton of a relative or absolute URI containing references to groups in the matchPattern which, once subpattern substitution has been performed, complete the URI.
Derived fromatt.patternReplacement
StatusRequired
Datatypeteidata.replacement
Note

Using TEI-defined XPointer schemes is not allowed.

identsupplies a name which functions as the prefix for an abbreviated pointing scheme such as a private URI scheme. The prefix constitutes the text preceding the first colon.
StatusRequired
Datatypeteidata.prefix
Note

The value is limited to teidata.prefix so that it may be mapped directly to a URI prefix.

Contained by
May contain
core: p
Note

The abbreviated pointer may be dereferenced to produce either an absolute or a relative URI reference. In the latter case it is combined with the value of xml:base in force at the place where the pointing attribute occurs to form an absolute URI in the usual manner as prescribed by XML Base.

Example
<prefixDef ident="mtematchPattern="(.+)" +}

Appendix A.1.73 <prefixDef>

<prefixDef> (prefix definition) defines a prefixing scheme used in teidata.pointer values, showing how abbreviated URIs using the scheme may be expanded into full URIs. [17.2.3. Using Abbreviated Pointers]
Moduleheader — Formal specification
Attributes
matchPatternspecifies a regular expression against which the values of other attributes can be matched.
Derived fromatt.patternReplacement
StatusRequired
Datatypeteidata.pattern
replacementPatternspecifies a ‘replacement pattern’, that is, the skeleton of a relative or absolute URI containing references to groups in the matchPattern which, once subpattern substitution has been performed, complete the URI.
Derived fromatt.patternReplacement
StatusRequired
Datatypeteidata.replacement
Note

Using TEI-defined XPointer schemes is not allowed.

identsupplies a name which functions as the prefix for an abbreviated pointing scheme such as a private URI scheme. The prefix constitutes the text preceding the first colon.
StatusRequired
Datatypeteidata.prefix
Note

The value is limited to teidata.prefix so that it may be mapped directly to a URI prefix.

Contained by
May contain
core: p
Note

The abbreviated pointer may be dereferenced to produce either an absolute or a relative URI reference. In the latter case it is combined with the value of xml:base in force at the place where the pointing attribute occurs to form an absolute URI in the usual manner as prescribed by XML Base.

Example
<prefixDef ident="mtematchPattern="(.+)"  replacementPattern="http://nl.ijs.si/ME/V6/msd/tables/msd-fslib-hbs.xml#$1">  <p xml:lang="en">Private URIs with this prefix point to feature-structure elements defining the Serbocroatian MULTEXT-East Version 6 MSDs.</p> -</prefixDef>
Content model
+</prefixDef>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element prefixDef
 {
    attribute matchPattern { text },
    attribute replacementPattern { text },
    attribute ident { text },
    tei_p+
-}

Appendix A.1.74 <profileDesc>

<profileDesc> (text-profile description) provides a detailed description of non-bibliographic aspects of a text, specifically the languages and sublanguages used, the situation in which it was produced, the participants and their setting. [2.4. The Profile Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Contained by
header: teiHeader
May contain
Note

Although the content model permits it, it is rarely meaningful to supply multiple occurrences for any of the child elements of <profileDesc> unless these are documenting multiple texts.

ExampleGeneral structure of the element <profileDesc>:
<profileDesc> +}

Appendix A.1.74 <profileDesc>

<profileDesc> (text-profile description) provides a detailed description of non-bibliographic aspects of a text, specifically the languages and sublanguages used, the situation in which it was produced, the participants and their setting. [2.4. The Profile Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Contained by
header: teiHeader
May contain
Note

Although the content model permits it, it is rarely meaningful to supply multiple occurrences for any of the child elements of <profileDesc> unless these are documenting multiple texts.

ExampleGeneral structure of the element <profileDesc>:
<profileDesc>  <settingDesc>...</settingDesc>  <textClass>...</textClass>  <particDesc>...</particDesc>  <langUsage>...</langUsage> -</profileDesc>
ExampleProfile description of a corpus root:
<profileDesc> +</profileDesc>
ExampleProfile description of a corpus root:
<profileDesc>  <settingDesc>   <setting>    <name type="address">Šubičeva ulica 4</name> @@ -2910,9 +2953,9 @@  </textClass>  <particDesc>    <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" -   href="href="ParlaMint-SI-listOrg.xml"/> +   href="ParlaMint-SI-listOrg.xml"/>    <xi:include xmlns:xi="http://www.w3.org/2001/XInclude" -   href="href="ParlaMint-SI-listPerson.xml"/> +   href="ParlaMint-SI-listPerson.xml"/>  </particDesc>  <langUsage>   <langUsage> @@ -2922,7 +2965,7 @@    <language ident="enxml:lang="en">English</language>   </langUsage>  </langUsage> -</profileDesc>
ExampleProfile description for a corpus component. In contrast to the corpus root, only the first, the <settingDesc> is used in corpus components.
<profileDesc> +</profileDesc>
ExampleProfile description for a corpus component. In contrast to the corpus root, only the first, the <settingDesc> is used in corpus components.
<profileDesc>  <settingDesc>   <setting>    <name type="city">Ljubljana</name> @@ -2931,7 +2974,7 @@     ana="#parla.sitting">28.8.2014</date>   </setting>  </settingDesc> -</profileDesc>
Content model
+</profileDesc>
Content model
 <content>
  <elementRef key="settingDesc"/>
  <elementRef key="textClass" minOccurs="0"
@@ -2941,37 +2984,41 @@
  <elementRef key="langUsage" minOccurs="0"
   maxOccurs="1"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element profileDesc
 {
    tei_settingDesc,
    tei_textClass?,
    tei_particDesc?,
    tei_langUsage?
-}

Appendix A.1.75 <projectDesc>

<projectDesc> (project description) describes in detail the aim or purpose for which an electronic file was encoded, together with any other relevant information concerning the process by which it was assembled or collected. [2.3.1. The Project Description 2.3. The Encoding Description 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
core: p
Example
<projectDesc> +}

Appendix A.1.75 <projectDesc>

<projectDesc> (project description) describes in detail the aim or purpose for which an electronic file was encoded, together with any other relevant information concerning the process by which it was assembled or collected. [2.3.1. The Project Description 2.3. The Encoding Description 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
core: p
Example
<projectDesc>  <p xml:lang="sl">Glavni cilji projekta <ref target="https://www.clarin.eu/content/parlamint">ParlaMint</ref> so    (1) izdelati večjezično množico na enak način kodiranih korpusov    zapiskov parlamentarnih sej, ...</p>  <p xml:lang="en">The <ref target="https://www.clarin.eu/content/parlamint">ParlaMint</ref>    project aims to (1) create a multilingual set of uniformly encoded    comparable corpora of parliamentary proceedings, ...</p> -</projectDesc>
Content model
+</projectDesc>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:projectDesc"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element projectDesc { tei_p+ }

Appendix A.1.76 <pubPlace>

<pubPlace> (publication place) contains the name of the place where a bibliographic item was published. [3.12.2.4. Imprint, Size of a Document, and Reprint Information]
Modulecore — Formal specification
Contained by
May contain
core: ref
character data
Example
<pubPlace> +
Schema Declaration
+element projectDesc { tei_p+ }

Appendix A.1.76 <pubPlace>

<pubPlace> (publication place) contains the name of the place where a bibliographic item was published. [3.12.2.4. Imprint, Size of a Document, and Reprint Information]
Modulecore — Formal specification
Contained by
May contain
core: ref
character data
Example
<pubPlace>  <ref target="https://github.com/clarin-eric/ParlaMint">https://github.com/clarin-eric/ParlaMint</ref> -</pubPlace>
Content model
+</pubPlace>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <elementRef key="ref"/>
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
-element pubPlace { tei_ref | text }

Appendix A.1.77 <publicationStmt>

<publicationStmt> (publication statement) groups information concerning the publication or distribution of an electronic or other text. [2.2.4. Publication, Distribution, Licensing, etc. 2.2. The File Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
Note

Where a publication statement contains several members of the model.publicationStmtPart.agency or model.publicationStmtPart.detail classes rather than one or more paragraphs or anonymous blocks, care should be taken to ensure that the repeated elements are presented in a meaningful order. It is a conformance requirement that elements supplying information about publication place, address, identifier, availability, and date be given following the name of the publisher, distributor, or authority concerned, and preferably in that order.

Example
<publicationStmt> +
Schema Declaration
+element pubPlace { tei_ref | text }

Appendix A.1.77 <publicationStmt>

<publicationStmt> (publication statement) groups information concerning the publication or distribution of an electronic or other text. [2.2.4. Publication, Distribution, Licensing, etc. 2.2. The File Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
Note

Where a publication statement contains several members of the model.publicationStmtPart.agency or model.publicationStmtPart.detail classes rather than one or more paragraphs or anonymous blocks, care should be taken to ensure that the repeated elements are presented in a meaningful order. It is a conformance requirement that elements supplying information about publication place, address, identifier, availability, and date be given following the name of the publisher, distributor, or authority concerned, and preferably in that order.

Example
<publicationStmt>  <publisher>   <orgName xml:lang="sl">Raziskovalna infrastrukutra CLARIN</orgName>   <orgName xml:lang="en">CLARIN research infrastructure</orgName> @@ -2984,7 +3031,7 @@   <p xml:lang="en">This work is licensed under the <ref target="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ref>.</p>  </availability>  <date when="2021-06-11">11. 6. 2021</date> -</publicationStmt>
Content model
+</publicationStmt>
Content model
 <content>
  <elementRef key="publisher"/>
  <elementRef key="idno"/>
@@ -2993,7 +3040,7 @@
  <elementRef key="availability"/>
  <elementRef key="date"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element publicationStmt
 {
    tei_publisher,
@@ -3001,10 +3048,10 @@
    tei_pubPlace?,
    tei_availability,
    tei_date
-}

Appendix A.1.78 <publisher>

<publisher> (publisher) provides the name of the organisation responsible for the publication or distribution of a bibliographic item. [3.12.2.4. Imprint, Size of a Document, and Reprint Information 2.2.4. Publication, Distribution, Licensing, etc.]
Modulecore — Formal specification
Attributes
Contained by
core: bibl
May contain
core: ref
namesdates: orgName
character data
Note

Use the full form of the name by which a company is usually referred to, rather than any abbreviation of it which may appear on a title page

Example
<publisher> +}

Appendix A.1.78 <publisher>

<publisher> (publisher) provides the name of the organisation responsible for the publication or distribution of a bibliographic item. [3.12.2.4. Imprint, Size of a Document, and Reprint Information 2.2.4. Publication, Distribution, Licensing, etc.]
Modulecore — Formal specification
Attributes
Contained by
core: bibl
May contain
core: ref
namesdates: orgName
character data
Note

Use the full form of the name by which a company is usually referred to, rather than any abbreviation of it which may appear on a title page

Example
<publisher>  <orgName>CLARIN research infrastructure</orgName>  <ref target="https://www.clarin.eu/">www.clarin.eu</ref> -</publisher>
Content model
+</publisher>
Content model
 <content>
  <alternate minOccurs="1" maxOccurs="1">
   <sequence minOccurs="1" maxOccurs="1">
@@ -3016,37 +3063,43 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element publisher
 {
    tei_att.global.attribute.xmllang,
    ( ( tei_orgName+, tei_ref* ) | text )
-}

Appendix A.1.79 <quotation>

<quotation> (quotation) specifies editorial practice adopted with respect to quotation marks in the original. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> ... +}

Appendix A.1.79 <quotation>

<quotation> (quotation) specifies editorial practice adopted with respect to quotation marks in the original. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> ... <quotation>   <p xml:lang="en">Quotation marks have been left in the text and are not explicitly marked up.</p>  </quotation> </editorialDecl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:quotation"/> +</sch:pattern>
Schematron
<sch:rule context="tei:quotation"> <sch:report test="not( @marks ) and not( tei:p )"> On <sch:name/>, either the @marks attribute should be used, or a paragraph of description provided </sch:report> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element quotation { tei_p+ }

Appendix A.1.80 <recording>

<recording> (recording event) provides details of an audio or video recording event used as the source of a spoken text, either directly or from a public broadcast. [8.2. Documenting the Source of Transcribed Speech 16.3.2. Declarable Elements]
Modulespoken — Formal specification
Attributes
typethe kind of recording.
Derived fromatt.typed
StatusOptional
Datatypeteidata.enumerated
Legal values are:
audio
audio recording[Default]
video
audio and video recording
Contained by
May contain
core: media
Note

The dur attribute is used to indicate the original duration of the recording.

Example
<recording type="audio"> +
Schema Declaration
+element quotation { tei_p+ }

Appendix A.1.80 <recording>

<recording> (recording event) provides details of an audio or video recording event used as the source of a spoken text, either directly or from a public broadcast. [8.2. Documenting the Source of Transcribed Speech 16.3.2. Declarable Elements]
Modulespoken — Formal specification
Attributes
typethe kind of recording.
Derived fromatt.typed
StatusOptional
Datatypeteidata.enumerated
Legal values are:
audio
audio recording[Default]
video
audio and video recording
Contained by
May contain
core: media
Note

The dur attribute is used to indicate the original duration of the recording.

Example
<recording type="audio">  <media xml:id="ps2013-044-02-000-000.audio1"   mimeType="audio/mp3"   source="https://www.psp.cz/eknih/2013ps/audio/2016/04/13/2016041308580912.mp3"   url="2013ps/audio/2016/04/13/2016041308580912.mp3"/> -</recording>
Content model
+</recording>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:recording"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="media" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element recording { attribute type { "audio" | "video" }?, tei_media+ }

Appendix A.1.81 <recordingStmt>

<recordingStmt> (recording statement) describes a set of recordings used as the basis for transcription of a spoken text. [8.2. Documenting the Source of Transcribed Speech 2.2.7. The Source Description]
Modulespoken — Formal specification
Contained by
header: sourceDesc
May contain
spoken: recording
Example
<recordingStmt> +
Schema Declaration
+element recording { attribute type { "audio" | "video" }?, tei_media+ }

Appendix A.1.81 <recordingStmt>

<recordingStmt> (recording statement) describes a set of recordings used as the basis for transcription of a spoken text. [8.2. Documenting the Source of Transcribed Speech 2.2.7. The Source Description]
Modulespoken — Formal specification
Contained by
header: sourceDesc
May contain
spoken: recording
Example
<recordingStmt>  <recording type="audio">   <media xml:id="ps2017-020-09-004-010.audio1"    mimeType="audio/mp3" @@ -3062,13 +3115,13 @@    url="2017ps/audio/2018/11/13/2018111318281842.mp3"/>    ...  </recording> -</recordingStmt>
Content model
+</recordingStmt>
Content model
 <content>
  <elementRef key="recording" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element recordingStmt { tei_recording+ }

Appendix A.1.82 <ref>

<ref> (reference) defines a reference to another location, possibly modified by additional text or comment. [3.7. Simple Links and Cross-References 17.1. Links]
Modulecore — Formal specification
Attributes
targetspecifies the destination of the reference by supplying one or more URI References.
Derived fromatt.pointing
StatusRecommended
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Member of
Contained by
May containCharacter data only
Note

The target and cRef attributes are mutually exclusive.

Example
<projectDesc> +
Schema Declaration
+element recordingStmt { tei_recording+ }

Appendix A.1.82 <ref>

<ref> (reference) defines a reference to another location, possibly modified by additional text or comment. [3.7. Simple Links and Cross-References 17.1. Links]
Modulecore — Formal specification
Attributes
targetspecifies the destination of the reference by supplying one or more URI References.
Derived fromatt.pointing
StatusRecommended
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Member of
Contained by
May containCharacter data only
Note

The target and cRef attributes are mutually exclusive.

Example
<projectDesc>  <p>   <ref target="https://www.clarin.eu/content/parlamint">ParlaMint</ref> is a    project that aims to create a multilingual set of comparable corpora of @@ -3077,22 +3130,22 @@ </projectDesc>
Schematron
<sch:rule context="tei:ref"> <sch:report test="@target and @cRef">Only one of the attributes @target and @cRef may be supplied on <sch:name/>.</sch:report> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element ref
 {
    tei_att.global.attribute.xmllang,
    attribute target { list { + } }?,
    text
-}

Appendix A.1.83 <relation>

<relation> (relationship) describes a relationship between two organisations. [14.3.2.3. Personal Relationships]
Modulenamesdates — Formal specification
Attributes
name
StatusRequired
Legal values are:
coalition
opposition
renaming
successor
representing
activeidentifies the ‘active’ participants in a non-mutual relationship, or all the participants in a mutual one.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
mutualsupplies a list of participants amongst all of whom the relationship holds equally.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
passiveidentifies the ‘passive’ participants in a non-mutual relationship.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Contained by
namesdates: listRelation
May containEmpty element
Note

Only one of the attributes active and mutual may be supplied; the attribute passive may be supplied only if the attribute active is supplied. Not all of these constraints can be enforced in all schema languages.

ExampleSpecification of coalition and opposition political parties (or parliamentary groups) in a given time period and legislative period:
<relation name="coalition" +}

Appendix A.1.83 <relation>

<relation> (relationship) describes a relationship between two organisations. [14.3.2.3. Personal Relationships]
Modulenamesdates — Formal specification
Attributes
name
StatusRequired
Legal values are:
coalition
opposition
renaming
successor
representing
activeidentifies the ‘active’ participants in a non-mutual relationship, or all the participants in a mutual one.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
mutualsupplies a list of participants amongst all of whom the relationship holds equally.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
passiveidentifies the ‘passive’ participants in a non-mutual relationship.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
Contained by
namesdates: listRelation
May containEmpty element
Note

Only one of the attributes active and mutual may be supplied; the attribute passive may be supplied only if the attribute active is supplied. Not all of these constraints can be enforced in all schema languages.

ExampleSpecification of coalition and opposition political parties (or parliamentary groups) in a given time period and legislative period:
<relation name="coalition"  mutual="#MR #OpenVld #N-VA #CD_en_Vfrom="2014-10-11to="2018-12-09"  ana="#period_54"/> <relation name="opposition"  active="#Ecolo #cdH #DéFi #Vuye_Wouters #sp.a #PP #PS #PTB #FDFpassive="#government.BE" - from="2014-10-11to="2018-12-09ana="#period_54"/>
ExampleSpecification of parliamentary group representing political parties in the parliament:
<relation name="representing" + from="2014-10-11to="2018-12-09ana="#period_54"/>
ExampleSpecification of parliamentary group representing political parties in the parliament:
<relation name="representing"  active="#parliamentaryGroup.CSSD.1107"  passive="#politicalParty.CSSD.153 #politicalParty.ENO.1from="2013-10-29to="2017-10-26"/>
Schematron
<sch:rule context="tei:relation"> @@ -3103,11 +3156,11 @@ </sch:rule>
Schematron
<sch:rule context="tei:relation"> <sch:report test="@passive and not(@active)">the attribute @passive may be supplied only if the attribute @active is supplied</sch:report> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element relation
 {
    tei_att.global.analytic.attribute.ana,
@@ -3121,19 +3174,19 @@
    ( attribute active { list { + } }? | attribute mutual { list { + } }? ),
    attribute passive { list { + } }?,
    empty
-}

Appendix A.1.84 <resp>

<resp> (responsibility) contains a phrase describing the nature of a person's intellectual responsibility, or an organisation's role in the production or distribution of a work. [3.12.2.2. Titles, Authors, and Editors 2.2.1. The Title Statement 2.2.2. The Edition Statement 2.2.5. The Series Statement]
Modulecore — Formal specification
Attributes
Contained by
core: respStmt
May containCharacter data only
Note

The attribute ref, inherited from the class att.canonical may be used to indicate the kind of responsibility in a normalized form by referring directly to a standardized list of responsibility types, such as that maintained by a naming authority, for example the list maintained at http://www.loc.gov/marc/relators/relacode.html for bibliographic usage.

Example
<respStmt> +}

Appendix A.1.84 <resp>

<resp> (responsibility) contains a phrase describing the nature of a person's intellectual responsibility, or an organisation's role in the production or distribution of a work. [3.12.2.2. Titles, Authors, and Editors 2.2.1. The Title Statement 2.2.2. The Edition Statement 2.2.5. The Series Statement]
Modulecore — Formal specification
Attributes
Contained by
core: respStmt
May containCharacter data only
Note

The attribute ref, inherited from the class att.canonical may be used to indicate the kind of responsibility in a normalized form by referring directly to a standardized list of responsibility types, such as that maintained by a naming authority, for example the list maintained at http://www.loc.gov/marc/relators/relacode.html for bibliographic usage.

Example
<respStmt>  <persName>Andrej Pančur</persName>  <resp>Kodiranje TEI</resp>  <resp xml:lang="en">TEI corpus encoding</resp> -</respStmt>
Content model
+</respStmt>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
-element resp { tei_att.global.attribute.xmllang, text }

Appendix A.1.85 <respStmt>

<respStmt> (statement of responsibility) supplies a statement of responsibility for the intellectual content of a text, edition, recording, or series, where the specialized elements for authors, editors, etc. do not suffice or do not apply. May also be used to encode information about individuals or organisations which have played a role in the production or distribution of a bibliographic work. [3.12.2.2. Titles, Authors, and Editors 2.2.1. The Title Statement 2.2.2. The Edition Statement 2.2.5. The Series Statement]
Modulecore — Formal specification
Contained by
header: titleStmt
May contain
core: resp
namesdates: persName
Example
<respStmt> +
Schema Declaration
+element resp { tei_att.global.attribute.xmllang, text }

Appendix A.1.85 <respStmt>

<respStmt> (statement of responsibility) supplies a statement of responsibility for the intellectual content of a text, edition, recording, or series, where the specialized elements for authors, editors, etc. do not suffice or do not apply. May also be used to encode information about individuals or organisations which have played a role in the production or distribution of a bibliographic work. [3.12.2.2. Titles, Authors, and Editors 2.2.1. The Title Statement 2.2.2. The Edition Statement 2.2.5. The Series Statement]
Modulecore — Formal specification
Contained by
header: titleStmt
May contain
core: resp
namesdates: persName
Example
<respStmt>  <persName>Matthew Coole</persName>  <resp>Data retrieval, Parla-CLARIN TEI XML corpus encoding and linguistic annotation.</resp> -</respStmt>
Example
<respStmt> +</respStmt>
Example
<respStmt>  <persName ref="https://orcid.org/0000-0003-3063-2239">Tommaso Agnoloni</persName>  <persName ref="https://orcid.org/0000-0002-8126-6294">Francesca Frontini</persName>  <persName ref="https://orcid.org/0000-0002-2953-8619">Simonetta Montemagni</persName> @@ -3157,39 +3210,39 @@  <resp xml:lang="en">Cleaning, normalisation and conversion to ParlaMint TEI XML</resp> </respStmt> ... -
Content model
+
Content model
 <content>
  <elementRef key="persName" minOccurs="1"
   maxOccurs="unbounded"/>
  <elementRef key="resp" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element respStmt { tei_persName+, tei_resp+ }

Appendix A.1.86 <revisionDesc>

<revisionDesc> (revision description) summarizes the revision history for a file [2.6. The Revision Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Attributes
Contained by
header: teiHeader
May contain
header: change
Note

If present on this element, the status attribute should indicate the current status of the document. The same attribute may appear on any <change> to record the status at the time of that change. Conventionally <change> elements should be given in reverse date order, with the most recent change at the start of the list.

Example
<revisionDesc> +
Schema Declaration
+element respStmt { tei_persName+, tei_resp+ }

Appendix A.1.86 <revisionDesc>

<revisionDesc> (revision description) summarizes the revision history for a file [2.6. The Revision Description 2.1.1. The TEI Header and Its Components]
Moduleheader — Formal specification
Attributes
Contained by
header: teiHeader
May contain
header: change
Note

If present on this element, the status attribute should indicate the current status of the document. The same attribute may appear on any <change> to record the status at the time of that change. Conventionally <change> elements should be given in reverse date order, with the most recent change at the start of the list.

Example
<revisionDesc>  <change when="2021-06-11">   <name>Tomaž Erjavec</name>: Finalized encoding.</change>  <change when="2021-05-28">   <name>Tomaž Erjavec</name>: Built corpus.</change> -</revisionDesc>
Content model
+</revisionDesc>
Content model
 <content>
  <elementRef key="change" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element revisionDesc { tei_att.global.attribute.xmllang, tei_change+ }

Appendix A.1.87 <roleName>

<roleName> (role name) contains a name component which indicates that the referent has a particular role or position in society, such as an official title or rank. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: affiliation persName
May containCharacter data only
Note

A <roleName> may be distinguished from an <addName> by virtue of the fact that, like a title, it typically exists independently of its holder.

Example
<persName> +
Schema Declaration
+element revisionDesc { tei_att.global.attribute.xmllang, tei_change+ }

Appendix A.1.87 <roleName>

<roleName> (role name) contains a name component which indicates that the referent has a particular role or position in society, such as an official title or rank. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
Member of
Contained by
namesdates: affiliation persName
May containCharacter data only
Note

A <roleName> may be distinguished from an <addName> by virtue of the fact that, like a title, it typically exists independently of its holder.

Example
<persName>  <surname>Murgel</surname>  <forename>Jasna</forename>  <roleName>dr.</roleName> -</persName>
Example
<affiliation role="ministerref="#GOV" +</persName>
Example
<affiliation role="ministerref="#GOV"  from="2020-08-01">  <roleName xml:lang="sl">Minister za obrambo</roleName>  <roleName xml:lang="en">Minister of Defence</roleName> -</affiliation>
Content model
+</affiliation>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
-element roleName { tei_att.global.attribute.xmllang, text }

Appendix A.1.88 <s>

<s> (s-unit) contains a sentence-like division of a text. [18.1. Linguistic Segment Categories 8.4.1. Segmentation]
Moduleanalysis — Formal specification
Attributes
Member of
Contained by
linking: seg
May contain
analysis: pc phr w
linking: linkGrp
Note

The <s> element may be used to mark orthographic sentences, or any other segmentation of a text, provided that the segmentation is end-to-end, complete, and non-nesting. For segmentation which is partial or recursive, the <seg> should be used instead.

The type attribute may be used to indicate the type of segmentation intended, according to any convenient typology.

Example
<s xml:id="ParlaMint-GB_2017-10-30-lords.seg4.1"> +
Schema Declaration
+element roleName { tei_att.global.attribute.xmllang, text }

Appendix A.1.88 <s>

<s> (s-unit) contains a sentence-like division of a text. [18.1. Linguistic Segment Categories 8.4.1. Segmentation]
Moduleanalysis — Formal specification
Attributes
Member of
Contained by
linking: seg
May contain
analysis: pc phr w
linking: linkGrp
Note

The <s> element may be used to mark orthographic sentences, or any other segmentation of a text, provided that the segmentation is end-to-end, complete, and non-nesting. For segmentation which is partial or recursive, the <seg> should be used instead.

The type attribute may be used to indicate the type of segmentation intended, according to any convenient typology.

Example
<s xml:id="ParlaMint-GB_2017-10-30-lords.seg4.1">  <w lemma="I"   msd="UPosTag=PRON|Case=Nom|Number=Sing|Person=1|PronType=Prspos="PRP">I</w>  <w lemma="support" @@ -3202,7 +3255,7 @@ </s>
Schematron
<sch:rule context="tei:s"> <sch:report test="tei:s">You may not nest one s element within another: use seg instead</sch:report> -</sch:rule>
Content model
+</sch:rule>
Content model
 <content>
  <elementRef key="measure" minOccurs="0"/>
  <alternate minOccurs="1"
@@ -3224,7 +3277,7 @@
  <elementRef key="linkGrp" minOccurs="0"
   maxOccurs="1"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element s
 {
    tei_att.global.attribute.xmlid,
@@ -3249,10 +3302,10 @@
     | tei_pb
    )+,
    tei_linkGrp?
-}

Appendix A.1.89 <seg>

<seg> (arbitrary segment) represents any segmentation of text below the ‘chunk’ level. [17.3. Blocks, Segments, and Anchors 6.2. Components of the Verse Line 7.2.5. Speech Contents]
Modulelinking — Formal specification
Attributes
Member of
Contained by
spoken: u
May contain
analysis: s
character data
Note

The <seg> element may be used at the encoder's discretion to mark any segments of the text of interest for processing. One use of the element is to mark text features for which no appropriate markup is otherwise defined. Another use is to provide an identifier for some segment which is to be pointed at by some other element—i.e. to provide a target, or a part of a target, for a <ptr> or other similar element.

Example
<u who="#DavidPriorana="#regular"> +}

Appendix A.1.89 <seg>

<seg> (arbitrary segment) represents any segmentation of text below the ‘chunk’ level. [17.3. Blocks, Segments, and Anchors 6.2. Components of the Verse Line 7.2.5. Speech Contents]
Modulelinking — Formal specification
Attributes
Member of
Contained by
spoken: u
May contain
analysis: s
character data
Note

The <seg> element may be used at the encoder's discretion to mark any segments of the text of interest for processing. One use of the element is to mark text features for which no appropriate markup is otherwise defined. Another use is to provide an identifier for some segment which is to be pointed at by some other element—i.e. to provide a target, or a part of a target, for a <ptr> or other similar element.

Example
<u who="#DavidPriorana="#regular">  <seg>I ask that the draft Regulations laid before the House on 5 December be approved.</seg>  <seg>The relevant document is the 20th Report from the Legislation Committee.</seg> -</u>
Content model
+</u>
Content model
 <content>
  <elementRef key="measure" minOccurs="0"/>
  <alternate minOccurs="1"
@@ -3270,7 +3323,7 @@
   </alternate>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element seg
 {
    tei_att.global.attribute.xmlid,
@@ -3287,55 +3340,63 @@
     | tei_pb
     | ( text | tei_s )*
    )+
-}

Appendix A.1.90 <segmentation>

<segmentation> (segmentation) describes the principles according to which the text has been segmented, for example into sentences, tone-units, graphemic strata, etc. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl> +}

Appendix A.1.90 <segmentation>

<segmentation> (segmentation) describes the principles according to which the text has been segmented, for example into sentences, tone-units, graphemic strata, etc. [2.3.3. The Editorial Practices Declaration 16.3.2. Declarable Elements]
Moduleheader — Formal specification
Contained by
May contain
core: p
Example
<editorialDecl>  <segmentation>   <p xml:lang="en">The texts are segmented into utterances (speeches) and segments (corresponding to paragraphs in the source transcription).</p>  </segmentation> -</editorialDecl>
Content model
+</editorialDecl>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:segmentation"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="p" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element segmentation { tei_p+ }

Appendix A.1.91 <setting>

<setting> describes one particular setting in which a language interaction takes place. [16.2.3. The Setting Description]
Modulecorpus — Formal specification
Contained by
corpus: settingDesc
May contain
core: date name
Note

If the who attribute is not supplied, the setting is assumed to be that of all participants in the language interaction.

Example
<setting> +
Schema Declaration
+element segmentation { tei_p+ }

Appendix A.1.91 <setting>

<setting> describes one particular setting in which a language interaction takes place. [16.2.3. The Setting Description]
Modulecorpus — Formal specification
Contained by
corpus: settingDesc
May contain
core: date name
Note

If the who attribute is not supplied, the setting is assumed to be that of all participants in the language interaction.

Example
<setting>  <name type="place">Commons Chamber</name>  <name type="place">Westminster</name>  <name type="city">London</name>  <name type="countrykey="GB">U.K.</name>  <date when="2019-02-18">February 18th, 2019</date> -</setting>
Content model
+</setting>
Content model
 <content>
  <elementRef key="name" minOccurs="1"
   maxOccurs="unbounded"/>
  <elementRef key="date"/>
 </content>
-    
Schema Declaration
-element setting { tei_name+, tei_date }

Appendix A.1.92 <settingDesc>

<settingDesc> (setting description) describes the setting or settings within which a language interaction takes place, or other places otherwise referred to in a text, edition, or metadata. [16.2. Contextual Information 2.4. The Profile Description]
Modulecorpus — Formal specification
Contained by
header: profileDesc
May contain
corpus: setting
Note

May contain a prose description organized as paragraphs, or a series of <setting> elements. If used to record not settings of language interactions, but other places mentioned in the text, then <place> optionally grouped by <listPlace> inside <standOff> should be preferred.

Example
<settingDesc> +
Schema Declaration
+element setting { tei_name+, tei_date }

Appendix A.1.92 <settingDesc>

<settingDesc> (setting description) describes the setting or settings within which a language interaction takes place, or other places otherwise referred to in a text, edition, or metadata. [16.2. Contextual Information 2.4. The Profile Description]
Modulecorpus — Formal specification
Contained by
header: profileDesc
May contain
corpus: setting
Note

May contain a prose description organized as paragraphs, or a series of <setting> elements. If used to record not settings of language interactions, but other places mentioned in the text, then <place> optionally grouped by <listPlace> inside <standOff> should be preferred.

Example
<settingDesc>  <setting>   <name type="address">Trg sv. Marka 6</name>   <name type="city">Zagreb</name>   <name type="countrykey="HR">Croatia</name>   <date from="2016-11-15to="2020-05-18">15.11.2016 - 18.5.2020</date>  </setting> -</settingDesc>
Content model
+</settingDesc>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:settingDesc"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="setting" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element settingDesc { tei_setting+ }

Appendix A.1.93 <sex>

<sex> (sex) specifies the sex of a person. [14.3.2.1. Personal Characteristics]
Modulenamesdates — Formal specification
Attributes
value
StatusRequired
Legal values are:
M
F
U
O
N
Contained by
namesdates: person
May containEmpty element
Note

As with other culturally-constructed traits such as age and gender, the way in which this concept is described in different cultural contexts varies. The normalizing attributes are provided only as an optional means of simplifying that variety for purposes of interoperability or project-internal taxonomies for consistency, and should not be used where that is inappropriate or unhelpful. The content of the element may be used to describe the intended concept in more detail.

Example
<sex value="M"/>
Content model
+    
Schema Declaration
+element settingDesc { tei_setting+ }

Appendix A.1.93 <sex>

<sex> (sex) specifies the sex of a person. [14.3.2.1. Personal Characteristics]
Modulenamesdates — Formal specification
Attributes
value
StatusRequired
Legal values are:
M
F
U
O
N
Contained by
namesdates: person
May containEmpty element
Note

As with other culturally-constructed traits such as age and gender, the way in which this concept is described in different cultural contexts varies. The normalizing attributes are provided only as an optional means of simplifying that variety for purposes of interoperability or project-internal taxonomies for consistency, and should not be used where that is inappropriate or unhelpful. The content of the element may be used to describe the intended concept in more detail.

Example
<sex value="M"/>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
-element sex { attribute value { "M" | "F" | "U" | "O" | "N" }, empty }

Appendix A.1.94 <sourceDesc>

<sourceDesc> (source description) describes the source(s) from which an electronic text was derived or generated, typically a bibliographic description in the case of a digitized text, or a phrase such as "born digital" for a text which has no previous existence. [2.2.7. The Source Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
core: bibl
ExampleThe source description <sourceDesc> of the corpus root encodes the original digital source of the ParlaMint corpus:
<sourceDesc> +
Schema Declaration
+element sex { attribute value { "M" | "F" | "U" | "O" | "N" }, empty }

Appendix A.1.94 <sourceDesc>

<sourceDesc> (source description) describes the source(s) from which an electronic text was derived or generated, typically a bibliographic description in the case of a digitized text, or a phrase such as "born digital" for a text which has no previous existence. [2.2.7. The Source Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
core: bibl
ExampleThe source description <sourceDesc> of the corpus root encodes the original digital source of the ParlaMint corpus:
<sourceDesc>  <bibl>   <title type="mainxml:lang="sl">Zapisi sej Državnega zbora Republike Slovenije</title>   <title type="mainxml:lang="en">Minutes of the National Assembly of the Republic of Slovenia</title>   <idno type="URI">https://www.dz-rs.si</idno>   <date from="2014-08-01to="2020-07-16">1.8.2014 - 16.7.2020</date>  </bibl> -</sourceDesc>
ExampleFor corpus components the source description is very similar to the one for the corpus root, except it reflects information of the exact meeting. Furthermore, if the audio or video of the meeting is available, this information can also be given:
<sourceDesc> +</sourceDesc>
ExampleFor corpus components the source description is very similar to the one for the corpus root, except it reflects information of the exact meeting. Furthermore, if the audio or video of the meeting is available, this information can also be given:
<sourceDesc>  <bibl>   <title type="mainxml:lang="cs">Parlament České republiky, Poslanecká sněmovna</title>   <title type="mainxml:lang="en">Parliament of the Czech Republic, Chamber of Deputies</title> @@ -3350,23 +3411,27 @@     url="2013ps/audio/2016/04/13/2016041308580912.mp3"/>   </recording>  </recordingStmt> -</sourceDesc>
Content model
+</sourceDesc>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" + value="tei:sourceDesc"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="bibl" minOccurs="1"
   maxOccurs="unbounded"/>
  <elementRef key="recordingStmt"
   minOccurs="0" maxOccurs="1"/>
 </content>
-    
Schema Declaration
-element sourceDesc { tei_bibl+, tei_recordingStmt? }

Appendix A.1.95 <state>

<state> (state) defines additional metadata on a political party or parliamentary group, e.g. its political orientation. [14.3.1. Basic Principles 14.3.2.1. Personal Characteristics]
Modulenamesdates — Formal specification
Attributes
ana(analysis) indicates one or more elements containing interpretations of the element on which the ana attribute appears.
Derived fromatt.global.analytic
StatusOptional
Datatypeteidata.pointer
type
StatusRequired
Legal values are:
politicalOrientation
encoder
Wikipedia
CHES
variable
value
Member of
Contained by
namesdates: org state
May contain
core: note
namesdates: state
Note

Where there is confusion between <trait> and <state> the more general purpose element <state> should be used even for unchanging characteristics. If you wish to distinguish between characteristics that are generally perceived to be time-bound states and those assumed to be fixed traits, then <trait> is available for the more static of these. The <state> element encodes characteristics which are sometimes assumed to change, often at specific times or over a date range, whereas the <trait> elements are used to record characteristics, such as eye-colour, which are less subject to change. Traits are typically, but not necessarily, independent of the volition or action of the holder.

ExampleEncoding political orientation as entered by an encoder:
<state type="politicalOrientation"> +
Schema Declaration
+element sourceDesc { tei_bibl+, tei_recordingStmt? }

Appendix A.1.95 <state>

<state> (state) defines additional metadata on a political party or parliamentary group, e.g. its political orientation. [14.3.1. Basic Principles 14.3.2.1. Personal Characteristics]
Modulenamesdates — Formal specification
Attributes
ana(analysis) indicates one or more elements containing interpretations of the element on which the ana attribute appears.
Derived fromatt.global.analytic
StatusOptional
Datatypeteidata.pointer
type
StatusRequired
Legal values are:
politicalOrientation
encoder
Wikipedia
CHES
variable
value
Member of
Contained by
namesdates: org state
May contain
core: note
namesdates: state
Note

Where there is confusion between <trait> and <state> the more general purpose element <state> should be used even for unchanging characteristics. If you wish to distinguish between characteristics that are generally perceived to be time-bound states and those assumed to be fixed traits, then <trait> is available for the more static of these. The <state> element encodes characteristics which are sometimes assumed to change, often at specific times or over a date range, whereas the <trait> elements are used to record characteristics, such as eye-colour, which are less subject to change. Traits are typically, but not necessarily, independent of the volition or action of the holder.

ExampleEncoding political orientation as entered by an encoder:
<state type="politicalOrientation">  <state type="encoder"   source="#GrietDepoorterana="#orientation.CRR">   <note xml:lang="en">Orientation determined by encoder, using own knowledge of the parliamentary group.</note>  </state> -</state>
ExampleEncoding Wikipedia-sourced political orientation:
<state type="politicalOrientation"> +</state>
ExampleEncoding Wikipedia-sourced political orientation:
<state type="politicalOrientation">  <state type="Wikipedia"   source="https://lv.wikipedia.org/wiki/Attīstībai/Par!ana="#orientation.C"/> -</state>
ExampleEncoding CHES-sourced variables and their values, with @key containing the CHES name for the political party:
<state type="CHESkey="AP!from="2019" +</state>
ExampleEncoding CHES-sourced variables and their values, with @key containing the CHES name for the political party:
<state type="CHESkey="AP!from="2019"  to="2019"  source="https://www.chesdata.eu/s/1999-2019_CHES_dataset_meansv3.csv">  <state type="variableana="#ches.lrgen"> @@ -3377,7 +3442,7 @@   <state type="valuefrom="2019to="2019"    n="5.90"/>  </state> -</state>
Content model
+</state>
Content model
 <content>
  <sequence minOccurs="1" maxOccurs="1">
   <elementRef key="note" minOccurs="0"
@@ -3386,7 +3451,7 @@
    maxOccurs="unbounded"/>
  </sequence>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element state
 {
    tei_att.global.attribute.n,
@@ -3405,10 +3470,10 @@
     | "value"
    },
    ( tei_note*, tei_state* )
-}

Appendix A.1.96 <surname>

<surname> (surname) contains a family (inherited) name, as opposed to a given, baptismal, or nick name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
type
StatusOptional
Legal values are:
birth
patronym
married
Member of
Contained by
namesdates: persName
May containCharacter data only
Example
<persName> +}

Appendix A.1.96 <surname>

<surname> (surname) contains a family (inherited) name, as opposed to a given, baptismal, or nick name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Attributes
type
StatusOptional
Legal values are:
birth
patronym
married
Member of
Contained by
namesdates: persName
May containCharacter data only
Example
<persName>  <surname>Accetto</surname>  <forename>Matej</forename> -</persName>
Example
<persName> +</persName>
Example
<persName>  <forename>Ірина</forename>  <surname type="patronym">Борисівна</surname>  <surname>Щеняєва</surname> @@ -3417,17 +3482,17 @@  <forename>Iryna</forename>  <surname type="patronym">Borysivna</surname>  <surname>Ščenjajeva</surname> -</persName>
Content model
+</persName>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element surname
 {
    tei_att.global.attribute.xmllang,
    attribute type { "birth" | "patronym" | "married" }?,
    text
-}

Appendix A.1.97 <tagUsage>

<tagUsage> (element usage) documents the usage of a specific element within a specified document. [2.3.4. The Tagging Declaration]
Moduleheader — Formal specification
Attributes
gi(generic identifier) specifies the name (generic identifier) of the element indicated by the tag, within the namespace indicated by the parent <namespace> element. All descendats of <text> element and <text> element counts have to be included.
StatusRequired
Datatypeteidata.name
occursspecifies the number of occurrences of this element within the text.
StatusRequired
Datatypeteidata.count
Contained by
header: namespace
May containEmpty element
Example
<tagsDecl> +}

Appendix A.1.97 <tagUsage>

<tagUsage> (element usage) documents the usage of a specific element within a specified document. [2.3.4. The Tagging Declaration]
Moduleheader — Formal specification
Attributes
gi(generic identifier) specifies the name (generic identifier) of the element indicated by the tag, within the namespace indicated by the parent <namespace> element. All descendats of <text> element and <text> element counts have to be included.
StatusRequired
Datatypeteidata.name
occursspecifies the number of occurrences of this element within the text.
StatusRequired
Datatypeteidata.count
Contained by
header: namespace
May containEmpty element
Example
<tagsDecl>  <namespace name="http://www.tei-c.org/ns/1.0">   <tagUsage gi="textoccurs="414"/>   <tagUsage gi="bodyoccurs="414"/> @@ -3442,12 +3507,12 @@   <tagUsage gi="kinesicoccurs="560"/>   <tagUsage gi="descoccurs="10234"/>  </namespace> -</tagsDecl>
Content model
+</tagsDecl>
Content model
 <content>
  <empty/>
 </content>
-    
Schema Declaration
-element tagUsage { attribute gi { text }, attribute occurs { text }, empty }

Appendix A.1.98 <tagsDecl>

<tagsDecl> (tagging declaration) provides detailed information about the tagging applied to a document. [2.3.4. The Tagging Declaration 2.3. The Encoding Description]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
header: namespace
ExampleThe tags declaration, <tagsDecl> of the corpus root gives the count of all the XML tags used in the data part (so, not in the TEI header) of the corpus (for the corpus root) or in an individual component of the corpus.
<encodingDesc> ... +
Schema Declaration
+element tagUsage { attribute gi { text }, attribute occurs { text }, empty }

Appendix A.1.98 <tagsDecl>

<tagsDecl> (tagging declaration) provides detailed information about the tagging applied to a document. [2.3.4. The Tagging Declaration 2.3. The Encoding Description]
Moduleheader — Formal specification
Contained by
header: encodingDesc
May contain
header: namespace
ExampleThe tags declaration, <tagsDecl> of the corpus root gives the count of all the XML tags used in the data part (so, not in the TEI header) of the corpus (for the corpus root) or in an individual component of the corpus.
<encodingDesc> ... <tagsDecl>   <namespace name="http://www.tei-c.org/ns/1.0">    <tagUsage gi="textoccurs="414"/> @@ -3456,12 +3521,12 @@      ...   </namespace>  </tagsDecl> -</encodingDesc>
Content model
+</encodingDesc>
Content model
 <content>
  <elementRef key="namespace"/>
 </content>
-    
Schema Declaration
-element tagsDecl { tei_namespace }

Appendix A.1.99 <taxonomy>

<taxonomy> (taxonomy) defines a typology explicitly by a structured taxonomy. [2.3.7. The Classification Declaration]
Moduleheader — Formal specification
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
Derived fromatt.global
StatusRequired
DatatypeID
Contained by
header: classDecl
May contain
core: desc
header: category
Note

Nested taxonomies are common in many fields, so the <taxonomy> element can be nested.

Example
<taxonomy xml:id="subcorpus"> +
Schema Declaration
+element tagsDecl { tei_namespace }

Appendix A.1.99 <taxonomy>

<taxonomy> (taxonomy) defines a typology explicitly by a structured taxonomy. [2.3.7. The Classification Declaration]
Moduleheader — Formal specification
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
Derived fromatt.global
StatusRequired
DatatypeID
Contained by
header: classDecl
May contain
core: desc
header: category
Note

Nested taxonomies are common in many fields, so the <taxonomy> element can be nested.

Example
<taxonomy xml:id="subcorpus">  <desc xml:lang="sl">   <term>Podkorpusi</term>  </desc> @@ -3480,7 +3545,7 @@   <catDesc xml:lang="en">    <term>COVID</term>: COVID subcorpus, from 2020-01-31 onwards</catDesc>  </category> -</taxonomy>
Example
<taxonomy xml:id="parla.legislature"> +</taxonomy>
Example
<taxonomy xml:id="parla.legislature">  <desc xml:lang="it">   <term>Legislatura</term>  </desc> @@ -3519,21 +3584,21 @@  role="parliamentxml:id="LEG">  <orgName full="yesxml:lang="it">Senato della Repubblica Italiana</orgName>  <orgName full="yesxml:lang="it">Senate of the Republic of Italy</orgName> -</org>
Content model
+</org>
Content model
 <content>
  <elementRef key="desc" minOccurs="1"
   maxOccurs="unbounded"/>
  <elementRef key="category" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element taxonomy
 {
    tei_att.global.attribute.xmllang,
    attribute xml:id { text },
    tei_desc+,
    tei_category+
-}

Appendix A.1.100 <teiCorpus>

<teiCorpus> (TEI corpus) contains one whole corpus, stored in the corpus root file comprising the corpus header and XInclude references to corpus component files, each containing a <TEI> element. [4. Default Text Structure 16.1. Varieties of Composite Text]
Modulecore — Formal specification
Attributes
xml:id
StatusRequired
DatatypeID
xml:lang
StatusRequired
Datatypeteidata.language
Contained by
May contain
derived-module-parlamint: include
header: teiHeader
textstructure: TEI
Note

Should contain one <teiHeader> for the corpus, and a series of <TEI> elements, one for each text.

As with all elements in the TEI scheme (except <egXML>) this element is in the TEI namespace (see 5.7.2. Namespaces). Thus, when it is used as the outermost element of a TEI document, it is necessary to specify the TEI namespace on it. This is customarily achieved by including http://www.tei-c.org/ns/1.0 as the value of the XML namespace declaration (xmlns), without indicating a prefix, and then not using a prefix on TEI elements in the rest of the document. For example: <teiCorpus version="4.8.1" xml:lang="en" xmlns="http://www.tei-c.org/ns/1.0">.

ExampleGeneral structure of a ParlaMint corpus root:
<teiCorpus xml:lang="en" +}

Appendix A.1.100 <teiCorpus>

<teiCorpus> (TEI corpus) contains one whole corpus, stored in the corpus root file comprising the corpus header and XInclude references to corpus component files, each containing a <TEI> element. [4. Default Text Structure 16.1. Varieties of Composite Text]
Modulecore — Formal specification
Attributes
xml:id
StatusRequired
DatatypeID
xml:lang
StatusRequired
Datatypeteidata.language
Contained by
May contain
derived-module-parlamint: include
header: teiHeader
textstructure: TEI
Note

Should contain one <teiHeader> for the corpus, and a series of <TEI> elements, one for each text.

As with all elements in the TEI scheme (except <egXML>) this element is in the TEI namespace (see 5.7.2. Namespaces). Thus, when it is used as the outermost element of a TEI document, it is necessary to specify the TEI namespace on it. This is customarily achieved by including http://www.tei-c.org/ns/1.0 as the value of the XML namespace declaration (xmlns), without indicating a prefix, and then not using a prefix on TEI elements in the rest of the document. For example: <teiCorpus version="4.8.1" xml:lang="en" xmlns="http://www.tei-c.org/ns/1.0">.

ExampleGeneral structure of a ParlaMint corpus root:
<teiCorpus xml:lang="en"  xml:id="ParlaMint-GB" xmlns="http://www.tei-c.org/ns/1.0">  <teiHeader> ...TEI header of the corpus...  </teiHeader> @@ -3543,7 +3608,7 @@ href="2015/ParlaMint-GB_2015-01-06-commons.xml"/> ... -</teiCorpus>
Content model
+</teiCorpus>
Content model
 <content>
  <elementRef key="teiHeader"/>
  <alternate minOccurs="1"
@@ -3552,7 +3617,7 @@
   <elementRef key="include"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element teiCorpus
 {
    tei_att.global.linking.attribute.corresp,
@@ -3560,12 +3625,12 @@
    attribute xml:lang { text },
    tei_teiHeader,
    ( tei_TEI | tei_include )+
-}

Appendix A.1.101 <teiHeader>

<teiHeader> (TEI header) supplies descriptive and declarative metadata associated with a digital resource or set of resources. [2.1.1. The TEI Header and Its Components 16.1. Varieties of Composite Text]
Moduleheader — Formal specification
Attributes
Contained by
core: teiCorpus
textstructure: TEI
May contain
Note

One of the few elements unconditionally required in any TEI document.

ExampleBasic structure of the <teiHeader>:
<teiHeader> +}

Appendix A.1.101 <teiHeader>

<teiHeader> (TEI header) supplies descriptive and declarative metadata associated with a digital resource or set of resources. [2.1.1. The TEI Header and Its Components 16.1. Varieties of Composite Text]
Moduleheader — Formal specification
Attributes
Contained by
core: teiCorpus
textstructure: TEI
May contain
Note

One of the few elements unconditionally required in any TEI document.

ExampleBasic structure of the <teiHeader>:
<teiHeader>  <fileDesc>...</fileDesc>  <encodingDesc>...</encodingDesc>  <profileDesc>...</profileDesc>  <revisionDesc>...</revisionDesc> -</teiHeader>
ExampleExample of a ParlaMint corpus component <teiHeader>:
<teiHeader> +</teiHeader>
ExampleExample of a ParlaMint corpus component <teiHeader>:
<teiHeader>  <fileDesc>   <titleStmt>    <title type="mainxml:lang="lv">Latvijas parlamenta corpus ParlaMint-LV, 12. Saeima, 2014-11-04 [ParlaMint]</title> @@ -3633,7 +3698,7 @@    </setting>   </settingDesc>  </profileDesc> -</teiHeader>
Content model
+</teiHeader>
Content model
 <content>
  <elementRef key="fileDesc"/>
  <elementRef key="encodingDesc"/>
@@ -3641,7 +3706,7 @@
  <elementRef key="revisionDesc"
   minOccurs="0" maxOccurs="1"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element teiHeader
 {
    tei_att.global.attribute.xmllang,
@@ -3649,7 +3714,7 @@
    tei_encodingDesc,
    tei_profileDesc,
    tei_revisionDesc?
-}

Appendix A.1.102 <term>

<term> (term) contains a single-word, multi-word, or symbolic designation which is regarded as a technical term. [3.4.1. Terms and Glosses]
Modulecore — Formal specification
Member of
Contained by
core: desc
header: catDesc
namesdates: persName
May containCharacter data only
Note

When this element appears within an <index> element, it is understood to supply the form under which an index entry is to be made for that location. Elsewhere, it is understood simply to indicate that its content is to be regarded as a technical or specialised term. It may be associated with a <gloss> element by means of its ref attribute; alternatively a <gloss> element may point to a <term> element by means of its target attribute.

In formal terminological work, there is frequently discussion over whether terms must be atomic or may include multi-word lexical items, symbolic designations, or phraseological units. The <term> element may be used to mark any of these. No position is taken on the philosophical issue of what a term can be; the looser definition simply allows the <term> element to be used by practitioners of any persuasion.

As with other members of the att.canonical class, instances of this element occuring in a text may be associated with a canonical definition, either by means of a URI (using the ref attribute), or by means of some system-specific code value (using the key attribute). Because the mutually exclusive target and cRef attributes overlap with the function of the ref attribute, they are deprecated and may be removed at a subsequent release.

Example<term> is used inside taxonomies to name the taxonomy and its categories:
<taxonomy xml:id="subcorpus"> +}

Appendix A.1.102 <term>

<term> (term) contains a single-word, multi-word, or symbolic designation which is regarded as a technical term. [3.4.1. Terms and Glosses]
Modulecore — Formal specification
Member of
Contained by
core: desc
header: catDesc
namesdates: persName
May containCharacter data only
Note

When this element appears within an <index> element, it is understood to supply the form under which an index entry is to be made for that location. Elsewhere, it is understood simply to indicate that its content is to be regarded as a technical or specialised term. It may be associated with a <gloss> element by means of its ref attribute; alternatively a <gloss> element may point to a <term> element by means of its target attribute.

In formal terminological work, there is frequently discussion over whether terms must be atomic or may include multi-word lexical items, symbolic designations, or phraseological units. The <term> element may be used to mark any of these. No position is taken on the philosophical issue of what a term can be; the looser definition simply allows the <term> element to be used by practitioners of any persuasion.

As with other members of the att.canonical class, instances of this element occuring in a text may be associated with a canonical definition, either by means of a URI (using the ref attribute), or by means of some system-specific code value (using the key attribute). Because the mutually exclusive target and cRef attributes overlap with the function of the ref attribute, they are deprecated and may be removed at a subsequent release.

Example<term> is used inside taxonomies to name the taxonomy and its categories:
<taxonomy xml:id="subcorpus">  <desc xml:lang="sl">   <term>Podkorpusi</term>  </desc> @@ -3664,7 +3729,7 @@  </category> ... -</taxonomy>
Example
<catDesc xml:lang="en"> +</taxonomy>
Example
<catDesc xml:lang="en">  <term>acl</term>: Clausal modifier of noun (adjectival clause) </catDesc> <catDesc xml:lang="en"> @@ -3672,22 +3737,22 @@ </catDesc> <catDesc xml:lang="en">  <term>punct</term>: Punctuation -</catDesc>
Content model
+</catDesc>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
-element term { text }

Appendix A.1.103 <text>

<text> (text) contains a single text of any kind, whether unitary or composite, for example a poem or drama, a collection of essays, a novel, a dictionary, or a corpus sample. [4. Default Text Structure 16.1. Varieties of Composite Text]
Moduletextstructure — Formal specification
Attributes
Contained by
textstructure: TEI
May contain
textstructure: body
Note

This element should not be used to represent a text which is inserted at an arbitrary point within the structure of another, for example as in an embedded or quoted narrative; the <floatingText> is provided for this purpose.

Example
<text ana="#reference"> +
Schema Declaration
+element term { text }

Appendix A.1.103 <text>

<text> (text) contains a single text of any kind, whether unitary or composite, for example a poem or drama, a collection of essays, a novel, a dictionary, or a corpus sample. [4. Default Text Structure 16.1. Varieties of Composite Text]
Moduletextstructure — Formal specification
Attributes
Contained by
textstructure: TEI
May contain
textstructure: body
Note

This element should not be used to represent a text which is inserted at an arbitrary point within the structure of another, for example as in an embedded or quoted narrative; the <floatingText> is provided for this purpose.

Example
<text ana="#reference">  <body>   <div type="debateSection">...</div>   <div type="debateSection">...</div>    ...  </body> -</text>
Content model
+</text>
Content model
 <content>
  <elementRef key="body"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element text
 {
    tei_att.global.attribute.xmlid,
@@ -3696,17 +3761,20 @@
    tei_att.global.analytic.attribute.ana,
    tei_att.global.source.attribute.source,
    tei_body
-}

Appendix A.1.104 <textClass>

<textClass> (text classification) groups information which describes the nature or topic of a text in terms of a standard classification scheme, thesaurus, etc. [2.4.3. The Text Classification]
Moduleheader — Formal specification
Contained by
header: profileDesc
May contain
header: catRef
Example
<textClass> +}

Appendix A.1.104 <textClass>

<textClass> (text classification) groups information which describes the nature or topic of a text in terms of a standard classification scheme, thesaurus, etc. [2.4.3. The Text Classification]
Moduleheader — Formal specification
Contained by
header: profileDesc
May contain
header: catRef
Example
<textClass>  <catRef scheme="#parla.legislature"   target="#parla.bi #parla.lower #parla.upper"/> -</textClass>
Content model
+</textClass>
Schematron
+<sch:pattern is-a="declarable"> +<sch:param name="tde" value="tei:textClass"/> +</sch:pattern>
Content model
 <content>
  <elementRef key="catRef"/>
 </content>
-    
Schema Declaration
-element textClass { tei_catRef }

Appendix A.1.105 <time>

<time> (time) contains a phrase defining a time of day in any format. [3.6.4. Dates and Times]
Modulecore — Formal specification
Attributes
Member of
Contained by
analysis: s
core: name note unit
May contain
analysis: pc w
character data
ExampleA note giving the time when e.g. the session started:
<note type="time"> +
Schema Declaration
+element textClass { tei_catRef }

Appendix A.1.105 <time>

<time> (time) contains a phrase defining a time of day in any format. [3.6.4. Dates and Times]
Modulecore — Formal specification
Attributes
Member of
Contained by
analysis: s
core: name note unit
May contain
analysis: pc w
character data
ExampleA note giving the time when e.g. the session started:
<note type="time">  <time when="2016-04-13T09:10:00">(9.10 hodin)</time> -</note>
Content model
+</note>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -3715,7 +3783,7 @@
   <textNode/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element time
 {
    tei_att.global.attribute.xmlid,
@@ -3726,23 +3794,23 @@
    tei_att.datable.w3c.attribute.to,
    tei_att.typed.attributes,
    ( tei_w | tei_pc | text )+
-}

Appendix A.1.106 <title>

<title> (title) contains a title for any kind of work. [3.12.2.2. Titles, Authors, and Editors 2.2.1. The Title Statement 2.2.5. The Series Statement]
Modulecore — Formal specification
Attributes
type
StatusRecommended
Legal values are:
main
sub
Note

Attribute is required in <titleStmt> context.

Member of
Contained by
core: bibl
header: titleStmt
May containCharacter data only
Note

The attributes key and ref, inherited from the class att.canonical may be used to indicate the canonical form for the title; the former, by supplying (for example) the identifier of a record in some external library system; the latter by pointing to an XML element somewhere containing the canonical form of the title.

ExampleThe <title> element as used in the <titleStmt> of the corpus root <teiHeader>:
<title type="mainxml:lang="cs">Český parlamentní korpus ParlaMint-CZ [ParlaMint]</title> +}

Appendix A.1.106 <title>

<title> (title) contains a title for any kind of work. [3.12.2.2. Titles, Authors, and Editors 2.2.1. The Title Statement 2.2.5. The Series Statement]
Modulecore — Formal specification
Attributes
type
StatusRecommended
Legal values are:
main
sub
Note

Attribute is required in <titleStmt> context.

Member of
Contained by
core: bibl
header: titleStmt
May containCharacter data only
Note

The attributes key and ref, inherited from the class att.canonical may be used to indicate the canonical form for the title; the former, by supplying (for example) the identifier of a record in some external library system; the latter by pointing to an XML element somewhere containing the canonical form of the title.

ExampleThe <title> element as used in the <titleStmt> of the corpus root <teiHeader>:
<title type="mainxml:lang="cs">Český parlamentní korpus ParlaMint-CZ [ParlaMint]</title> <title type="mainxml:lang="en">Czech parliamentary corpus ParlaMint-CZ [ParlaMint]</title> <title type="subxml:lang="cs">Parlament České republiky, Poslanecká sněmovna</title> -<title type="subxml:lang="en">Parliament of the Czech Republic, Chamber of Deputies</title>
ExampleThe <title> element as used in the <titleStmt> of the corpus component <teiHeader>:
<title type="mainxml:lang="cs">Český parlamentní korpus ParlaMint-CZ, 2013-11-25 ps2013-001-01-000-000 [ParlaMint]</title> +<title type="subxml:lang="en">Parliament of the Czech Republic, Chamber of Deputies</title>
ExampleThe <title> element as used in the <titleStmt> of the corpus component <teiHeader>:
<title type="mainxml:lang="cs">Český parlamentní korpus ParlaMint-CZ, 2013-11-25 ps2013-001-01-000-000 [ParlaMint]</title> <title type="mainxml:lang="en">Czech parliamentary corpus ParlaMint-CZ, 2013-11-25 ps2013-001-01-000-000 [ParlaMint]</title> <title type="subxml:lang="cs">Parlament České republiky, Poslanecká sněmovna, 2013-11-25, Začátek schůze Poslanecké sněmovny 25. listopadu 2013 ve 14.05 hodin Přítomno: 199 poslanců</title> -<title type="subxml:lang="en">Parliament of the Czech Republic, Chamber of Deputies, 2013-11-25</title>
Content model
+<title type="subxml:lang="en">Parliament of the Czech Republic, Chamber of Deputies, 2013-11-25</title>
Content model
 <content>
  <textNode/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element title
 {
    tei_att.global.attribute.xmllang,
    attribute type { "main" | "sub" }?,
    text
-}

Appendix A.1.107 <titleStmt>

<titleStmt> (title statement) groups information about the title of a work and those responsible for its content. [2.2.1. The Title Statement 2.2. The File Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
ExampleThe <titleStmt> element gives the title of the corpus root or component, along with the specification of the particular session(s) of the parliament contained, the persons responsible for compiling the corpus and the funder(s) of the project:
<titleStmt> +}

Appendix A.1.107 <titleStmt>

<titleStmt> (title statement) groups information about the title of a work and those responsible for its content. [2.2.1. The Title Statement 2.2. The File Description]
Moduleheader — Formal specification
Contained by
header: fileDesc
May contain
ExampleThe <titleStmt> element gives the title of the corpus root or component, along with the specification of the particular session(s) of the parliament contained, the persons responsible for compiling the corpus and the funder(s) of the project:
<titleStmt>  <title type="main">Slovenski parlamentarni korpus ParlaMint-SI [ParlaMint]</title>  <title type="mainxml:lang="en">Slovenian parliamentary corpus ParlaMint-SI [ParlaMint]</title>  <title type="sub">Zapisi sej Državnega zbora Republike Slovenije, 7. in 8. mandat (2014 - 2020)</title> @@ -3765,7 +3833,7 @@   <orgName>Slovenska raziskovalna infrastruktura CLARIN.SI</orgName>   <orgName xml:lang="en">The Slovenian research infrastructure CLARIN.SI</orgName>  </funder> -</titleStmt>
Content model
+</titleStmt>
Content model
 <content>
  <elementRef key="title" minOccurs="1"
   maxOccurs="unbounded"/>
@@ -3776,11 +3844,11 @@
  <elementRef key="funder" minOccurs="0"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
-element titleStmt { tei_title+, tei_meeting+, tei_respStmt*, tei_funder* }

Appendix A.1.108 <u>

<u> (utterance) contains a stretch of speech usually preceded and followed by silence or by a change of speaker. [8.3.1. Utterances]
Modulespoken — Formal specification
Attributes
Member of
Contained by
textstructure: div
May contain
linking: seg
Note

Prose and a mixture of speech elements

Although individual transcriptions may consistently use <u> elements for turns or other units, and although in most cases a <u> will be delimited by pause or change of speaker, <u> is not required to represent a turn or any communicative event, nor to be bounded by pauses or change of speaker. At a minimum, a <u> is some phonetic production by a given speaker.

ExampleThe element <u> marks up a speech, as illustrated below:
<u who="#DavidPriorana="#regular"> +
Schema Declaration
+element titleStmt { tei_title+, tei_meeting+, tei_respStmt*, tei_funder* }

Appendix A.1.108 <u>

<u> (utterance) contains a stretch of speech usually preceded and followed by silence or by a change of speaker. [8.3.1. Utterances]
Modulespoken — Formal specification
Attributes
Member of
Contained by
textstructure: div
May contain
linking: seg
Note

Prose and a mixture of speech elements

Although individual transcriptions may consistently use <u> elements for turns or other units, and although in most cases a <u> will be delimited by pause or change of speaker, <u> is not required to represent a turn or any communicative event, nor to be bounded by pauses or change of speaker. At a minimum, a <u> is some phonetic production by a given speaker.

ExampleThe element <u> marks up a speech, as illustrated below:
<u who="#DavidPriorana="#regular">  <seg>I ask that the draft Regulations laid before the House on 5 December be approved.</seg>  <seg>The relevant document is the 20th Report from the Legislation Committee.</seg> -</u>
Content model
+</u>
Content model
 <content>
  <elementRef key="measure" minOccurs="0"/>
  <alternate minOccurs="1"
@@ -3794,7 +3862,7 @@
   <elementRef key="seg"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element u
 {
    tei_att.global.attribute.xmlid,
@@ -3816,7 +3884,7 @@
     | tei_pb
     | tei_seg
    )+
-}

Appendix A.1.109 <unit>

<unit> contains a symbol, a word or a phrase referring to a unit of measurement in any kind of formal or informal system. [3.6.3. Numbers and Measures]
Modulecore — Formal specification
Attributes
Member of
Contained by
core: unit
May contain
ExampleThe element can be used for fine-grained Named Entities which include units:
<num ana="ne:nc" +}

Appendix A.1.109 <unit>

<unit> contains a symbol, a word or a phrase referring to a unit of measurement in any kind of formal or informal system. [3.6.3. Numbers and Measures]
Modulecore — Formal specification
Attributes
Member of
Contained by
core: unit
May contain
ExampleThe element can be used for fine-grained Named Entities which include units:
<num ana="ne:nc"  xml:id="ParlaMint-CZ_2013-12-06-ps2013-003-01-001-001.ne53">  <w xml:id="ParlaMint-CZ_2013-12-06-ps2013-003-01-001-001.u2.p10.s1.w9"   lemma="3" @@ -3830,7 +3898,7 @@  <w xml:id="ParlaMint-CZ_2013-12-06-ps2013-003-01-001-001.u2.p10.s1.w11"   lemma=""   msd="UPosTag=NOUN|Gender=Fem|Polarity=Posjoin="right"></w> -</unit>
Content model
+</unit>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -3850,7 +3918,7 @@
   <elementRef key="vocal"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element unit
 {
    tei_att.global.attribute.xmlid,
@@ -3873,14 +3941,14 @@
     | tei_incident
     | tei_vocal
    )+
-}

Appendix A.1.110 <vocal>

<vocal> (vocal) marks any vocalized but not necessarily lexical phenomenon, for example voiced pauses, non-lexical backchannels, etc. [8.3.3. Vocal, Kinesic, Incident]
Modulespoken — Formal specification
Attributes
type
StatusRecommended
Legal values are:
greeting
question
clarification
speaking
interruption
exclamat
laughter
shouting
murmuring
noise
signal
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Example
<vocal type="interruption"> +}

Appendix A.1.110 <vocal>

<vocal> (vocal) marks any vocalized but not necessarily lexical phenomenon, for example voiced pauses, non-lexical backchannels, etc. [8.3.3. Vocal, Kinesic, Incident]
Modulespoken — Formal specification
Attributes
type
StatusRecommended
Legal values are:
greeting
question
clarification
speaking
interruption
exclamat
laughter
shouting
murmuring
noise
signal
Member of
Contained by
analysis: s
core: name unit
linking: seg
spoken: u
textstructure: div
May contain
core: desc
Example
<vocal type="interruption">  <desc>Interruption from the chair: Your time is up.</desc> -</vocal>
Content model
+</vocal>
Content model
 <content>
  <elementRef key="desc" minOccurs="1"
   maxOccurs="unbounded"/>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element vocal
 {
    tei_att.global.attribute.xmlid,
@@ -3904,7 +3972,7 @@
     | "signal"
    }?,
    tei_desc+
-}

Appendix A.1.111 <w>

<w> (word) represents a grammatical (not necessarily orthographic) word. [18.1. Linguistic Segment Categories 18.4.2. Lightweight Linguistic Annotation]
Moduleanalysis — Formal specification
Attributes
Member of
Contained by
analysis: phr s w
May contain
analysis: w
character data
Example
<s xml:id="ParlaMint-GB_2017-10-30-lords.seg4.1"> +}

Appendix A.1.111 <w>

<w> (word) represents a grammatical (not necessarily orthographic) word. [18.1. Linguistic Segment Categories 18.4.2. Lightweight Linguistic Annotation]
Moduleanalysis — Formal specification
Attributes
Member of
Contained by
analysis: phr s w
May contain
analysis: w
character data
Example
<s xml:id="ParlaMint-GB_2017-10-30-lords.seg4.1">  <w lemma="I"   msd="UPosTag=PRON|Case=Nom|Number=Sing|Person=1|PronType=Prspos="PRP">I</w>  <w lemma="support" @@ -3914,12 +3982,12 @@  <w lemma="amendment"   msd="UPosTag=NOUN|Number=Singpos="NNjoin="right">amendment</w>  <pc msd="UPosTag=PUNCTpos=".">.</pc> -</s>
ExampleCertain frameworks, in particular the Universal Dependencies, allow for tokens to be decomposed into several words, and it is these syntactic words, and not tokens, that are further annotated. For example, Czech has the word ‘abyste’ which is in UD decomposed into two syntactic words, ‘aby’ and ‘byste’, which can be encoded in the <w> element:
<w>abyste +</s>
ExampleCertain frameworks, in particular the Universal Dependencies, allow for tokens to be decomposed into several words, and it is these syntactic words, and not tokens, that are further annotated. For example, Czech has the word ‘abyste’ which is in UD decomposed into two syntactic words, ‘aby’ and ‘byste’, which can be encoded in the <w> element:
<w>abyste <w norm="abylemma="aby"   msd="UPosTag=SCONJ"/>  <w norm="bystelemma="být"   msd="UPosTag=AUX|Mood=Cnd|Number=Plur|Person=2|VerbForm=Fin"/> -</w>
Content model
+</w>
Content model
 <content>
  <alternate minOccurs="1"
   maxOccurs="unbounded">
@@ -3927,7 +3995,7 @@
   <elementRef key="w"/>
  </alternate>
 </content>
-    
Schema Declaration
+    
Schema Declaration
 element w
 {
    tei_att.global.attribute.xmlid,
@@ -3936,7 +4004,7 @@
    tei_att.linguistic.attributes,
    tei_att.segLike.attribute.function,
    ( text | tei_w )+
-}

Appendix A.2 Model classes

Appendix A.2.1 model.addressLike

model.addressLike groups elements used to represent a postal or email address. [1. The TEI Infrastructure]
Moduletei — Formal specification
Used by
Membersaffiliation email

Appendix A.2.2 model.attributable

model.attributable groups elements that contain a word or phrase that can be attributed to a source. [3.3.3. Quotation 4.3.2. Floating Texts]
Moduletei — Formal specification
Used by
Membersmodel.quoteLike

Appendix A.2.3 model.biblLike

model.biblLike groups elements containing a bibliographic description. [3.12. Bibliographic Citations and References]
Moduletei — Formal specification
Used by
Membersbibl

Appendix A.2.4 model.dateLike

model.dateLike groups elements containing temporal expressions. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Used by
Membersdate time

Appendix A.2.5 model.divPart

model.divPart groups paragraph-level elements appearing directly within divisions. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.divPart.spoken[u] model.lLike model.pLike[p]
Note

Note that this element class does not include members of the model.inter class, which can appear either within or between paragraph-level items.

Appendix A.2.6 model.divPart.spoken

model.divPart.spoken groups elements structurally analogous to paragraphs within spoken texts. [8.1. General Considerations and Overview]
Modulespoken — Formal specification
Used by
Membersu
Note

Spoken texts may be structured in many ways; elements in this class are typically larger units such as turns or utterances.

Appendix A.2.7 model.emphLike

model.emphLike groups phrase-level elements which are typographically distinct and to which a specific function can be attributed. [3.3. Highlighting and Quotation]
Moduletei — Formal specification
Used by
Membersterm title

Appendix A.2.8 model.global

Appendix A.2.9 model.global.edit

model.global.edit groups globally available elements which perform a specifically editorial function. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersgap

Appendix A.2.10 model.global.meta

model.global.meta groups globally available elements which describe the status of other elements. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Memberslink linkGrp
Note

Elements in this class are typically used to hold groups of links or of abstract interpretations, or by provide indications of certainty etc. It may find be convenient to localize all metadata elements, for example to contain them within the same divison as the elements that they relate to; or to locate them all to a division of their own. They may however appear at any point in a TEI text.

Appendix A.2.11 model.global.spoken

model.global.spoken groups elements which may appear globally within spoken texts. [8.1. General Considerations and Overview]
Modulespoken — Formal specification
Used by
Membersincident kinesic vocal
Note

This class groups elements which can appear anywhere within transcribed speech.

Appendix A.2.12 model.graphicLike

model.graphicLike groups elements containing images, formulae, and similar objects. [3.10. Graphics and Other Non-textual Components]
Moduletei — Formal specification
Used by
Membersgraphic media

Appendix A.2.13 model.highlighted

model.highlighted groups phrase-level elements which are typographically distinct. [3.3. Highlighting and Quotation]
Moduletei — Formal specification
Used by
Membersmodel.emphLike[term title] model.hiLike

Appendix A.2.14 model.inter

model.inter groups elements which can appear either within or between paragraph-like elements. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.attributable[model.quoteLike] model.biblLike[bibl] model.egLike model.labelLike[desc label] model.listLike[listEvent listOrg listPerson listRelation] model.oddDecl model.stageLike

Appendix A.2.15 model.labelLike

model.labelLike groups elements used to gloss or explain other parts of a document.
Moduletei — Formal specification
Used by
Membersdesc label

Appendix A.2.16 model.limitedPhrase

model.limitedPhrase groups phrase-level elements excluding those elements primarily intended for transcription of existing sources. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.emphLike[term title] model.hiLike model.pPart.data[model.addressLike[affiliation email] model.dateLike[date time] model.measureLike[measure num unit] model.nameLike[model.nameLike.agent[name orgName persName] model.offsetLike model.persNamePart[addName forename nameLink roleName surname] model.placeStateLike[model.placeNamePart[placeName] state] idno]] model.pPart.editorial model.pPart.msdesc model.phrase.xml model.ptrLike[ref]

Appendix A.2.17 model.listLike

model.listLike groups list-like elements. [3.8. Lists]
Moduletei — Formal specification
Used by
MemberslistEvent listOrg listPerson listRelation

Appendix A.2.18 model.measureLike

model.measureLike groups elements which denote a number, a quantity, a measurement, or similar piece of text that conveys some numerical meaning. [3.6.3. Numbers and Measures]
Moduletei — Formal specification
Used by
Membersmeasure num unit

Appendix A.2.19 model.milestoneLike

model.milestoneLike groups milestone-style elements used to represent reference systems. [1.3. The TEI Class System 3.11.3. Milestone Elements]
Moduletei — Formal specification
Used by
Memberspb

Appendix A.2.20 model.nameLike

model.nameLike groups elements which name or refer to a person, place, or organization.
Moduletei — Formal specification
Used by
Membersmodel.nameLike.agent[name orgName persName] model.offsetLike model.persNamePart[addName forename nameLink roleName surname] model.placeStateLike[model.placeNamePart[placeName] state] idno
Note

A superset of the naming elements that may appear in datelines, addresses, statements of responsibility, etc.

Appendix A.2.21 model.nameLike.agent

model.nameLike.agent groups elements which contain names of individuals or corporate bodies. [3.6. Names, Numbers, Dates, Abbreviations, and Addresses]
Moduletei — Formal specification
Used by
Membersname orgName persName
Note

This class is used in the content model of elements which reference names of people or organizations.

Appendix A.2.22 model.noteLike

model.noteLike groups globally-available note-like elements. [3.9. Notes, Annotation, and Indexing]
Moduletei — Formal specification
Used by
Membersnote

Appendix A.2.23 model.pLike

model.pLike groups paragraph-like elements.
Moduletei — Formal specification
Used by
Membersp

Appendix A.2.25 model.pPart.edit

model.pPart.edit groups phrase-level elements for simple editorial correction and transcription. [3.5. Simple Editorial Changes]
Moduletei — Formal specification
Used by
Membersmodel.pPart.editorial model.pPart.transcriptional

Appendix A.2.26 model.paraPart

Appendix A.2.27 model.persNamePart

model.persNamePart groups elements which form part of a personal name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Used by
MembersaddName forename nameLink roleName surname

Appendix A.2.28 model.phrase

model.phrase groups elements which can occur at the level of individual words or phrases. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.graphicLike[graphic media] model.highlighted[model.emphLike[term title] model.hiLike] model.lPart model.pPart.data[model.addressLike[affiliation email] model.dateLike[date time] model.measureLike[measure num unit] model.nameLike[model.nameLike.agent[name orgName persName] model.offsetLike model.persNamePart[addName forename nameLink roleName surname] model.placeStateLike[model.placeNamePart[placeName] state] idno]] model.pPart.edit[model.pPart.editorial model.pPart.transcriptional] model.pPart.msdesc model.phrase.xml model.ptrLike[ref] model.segLike[pc phr s seg w] model.specDescLike
Note

This class of elements can occur within paragraphs, list items, lines of verse, etc.

Appendix A.2.29 model.placeNamePart

model.placeNamePart groups elements which form part of a place name. [14.2.3. Place Names]
Moduletei — Formal specification
Used by
MembersplaceName

Appendix A.2.30 model.placeStateLike

model.placeStateLike groups elements which describe changing states of a place.
Moduletei — Formal specification
Used by
Membersmodel.placeNamePart[placeName] state

Appendix A.2.31 model.ptrLike

model.ptrLike groups elements used for purposes of location and reference. [3.7. Simple Links and Cross-References]
Moduletei — Formal specification
Used by
Membersref

Appendix A.2.32 model.segLike

model.segLike groups elements used for arbitrary segmentation. [17.3. Blocks, Segments, and Anchors 18.1. Linguistic Segment Categories]
Moduletei — Formal specification
Used by
Memberspc phr s seg w
Note

The principles on which segmentation is carried out, and any special codes or attribute values used, should be defined explicitly in the <segmentation> element of the <encodingDesc> within the associated TEI header.

Appendix A.3 Attribute classes

Appendix A.3.1 att.ascribed

att.ascribed provides attributes for elements representing speech or action that can be ascribed to a specific individual. [3.3.3. Quotation 8.3. Elements Unique to Spoken Texts]
Moduletei — Formal specification
Membersatt.ascribed.directed[kinesic u vocal] change incident setting
Attributes
whoindicates the person, or group of people, to whom the element content is ascribed.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
In the following example from Hamlet, speeches (<sp>) in the body of the play are linked to <role> elements in the <castList> using the who attribute.
<castItem type="role"> +}

Appendix A.2 Model classes

Appendix A.2.1 model.addressLike

model.addressLike groups elements used to represent a postal or email address. [1. The TEI Infrastructure]
Moduletei — Formal specification
Used by
Membersaffiliation email

Appendix A.2.2 model.attributable

model.attributable groups elements that contain a word or phrase that can be attributed to a source. [3.3.3. Quotation 4.3.2. Floating Texts]
Moduletei — Formal specification
Used by
Membersmodel.quoteLike

Appendix A.2.3 model.biblLike

model.biblLike groups elements containing a bibliographic description. [3.12. Bibliographic Citations and References]
Moduletei — Formal specification
Used by
Membersbibl

Appendix A.2.4 model.dateLike

model.dateLike groups elements containing temporal expressions. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Used by
Membersdate time

Appendix A.2.5 model.divPart

model.divPart groups paragraph-level elements appearing directly within divisions. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.divPart.spoken[u] model.lLike model.pLike[p]
Note

Note that this element class does not include members of the model.inter class, which can appear either within or between paragraph-level items.

Appendix A.2.6 model.divPart.spoken

model.divPart.spoken groups elements structurally analogous to paragraphs within spoken texts. [8.1. General Considerations and Overview]
Modulespoken — Formal specification
Used by
Membersu
Note

Spoken texts may be structured in many ways; elements in this class are typically larger units such as turns or utterances.

Appendix A.2.7 model.emphLike

model.emphLike groups phrase-level elements which are typographically distinct and to which a specific function can be attributed. [3.3. Highlighting and Quotation]
Moduletei — Formal specification
Used by
Membersterm title

Appendix A.2.8 model.global

Appendix A.2.9 model.global.edit

model.global.edit groups globally available elements which perform a specifically editorial function. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersgap

Appendix A.2.10 model.global.meta

model.global.meta groups globally available elements which describe the status of other elements. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Memberslink linkGrp
Note

Elements in this class are typically used to hold groups of links or of abstract interpretations, or by provide indications of certainty etc. It may find be convenient to localize all metadata elements, for example to contain them within the same divison as the elements that they relate to; or to locate them all to a division of their own. They may however appear at any point in a TEI text.

Appendix A.2.11 model.global.spoken

model.global.spoken groups elements which may appear globally within spoken texts. [8.1. General Considerations and Overview]
Modulespoken — Formal specification
Used by
Membersincident kinesic vocal
Note

This class groups elements which can appear anywhere within transcribed speech.

Appendix A.2.12 model.graphicLike

model.graphicLike groups elements containing images, formulae, and similar objects. [3.10. Graphics and Other Non-textual Components]
Moduletei — Formal specification
Used by
Membersgraphic media

Appendix A.2.13 model.highlighted

model.highlighted groups phrase-level elements which are typographically distinct. [3.3. Highlighting and Quotation]
Moduletei — Formal specification
Used by
Membersmodel.emphLike[term title] model.hiLike

Appendix A.2.14 model.inter

model.inter groups elements which can appear either within or between paragraph-like elements. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.attributable[model.quoteLike] model.biblLike[bibl] model.egLike model.labelLike[desc label] model.listLike[listEvent listOrg listPerson listRelation] model.oddDecl model.stageLike

Appendix A.2.15 model.labelLike

model.labelLike groups elements used to gloss or explain other parts of a document.
Moduletei — Formal specification
Used by
Membersdesc label

Appendix A.2.16 model.limitedPhrase

model.limitedPhrase groups phrase-level elements excluding those elements primarily intended for transcription of existing sources. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.emphLike[term title] model.hiLike model.pPart.data[model.addressLike[affiliation email] model.dateLike[date time] model.measureLike[measure num unit] model.nameLike[model.nameLike.agent[name orgName persName] model.offsetLike model.persNamePart[addName forename nameLink roleName surname] model.placeStateLike[model.placeNamePart[placeName] state] idno]] model.pPart.editorial model.pPart.msdesc model.phrase.xml model.ptrLike[ref]

Appendix A.2.17 model.listLike

model.listLike groups list-like elements. [3.8. Lists]
Moduletei — Formal specification
Used by
MemberslistEvent listOrg listPerson listRelation

Appendix A.2.18 model.measureLike

model.measureLike groups elements which denote a number, a quantity, a measurement, or similar piece of text that conveys some numerical meaning. [3.6.3. Numbers and Measures]
Moduletei — Formal specification
Used by
Membersmeasure num unit

Appendix A.2.19 model.milestoneLike

model.milestoneLike groups milestone-style elements used to represent reference systems. [1.3. The TEI Class System 3.11.3. Milestone Elements]
Moduletei — Formal specification
Used by
Memberspb

Appendix A.2.20 model.nameLike

model.nameLike groups elements which name or refer to a person, place, or organization.
Moduletei — Formal specification
Used by
Membersmodel.nameLike.agent[name orgName persName] model.offsetLike model.persNamePart[addName forename nameLink roleName surname] model.placeStateLike[model.placeNamePart[placeName] state] idno
Note

A superset of the naming elements that may appear in datelines, addresses, statements of responsibility, etc.

Appendix A.2.21 model.nameLike.agent

model.nameLike.agent groups elements which contain names of individuals or corporate bodies. [3.6. Names, Numbers, Dates, Abbreviations, and Addresses]
Moduletei — Formal specification
Used by
Membersname orgName persName
Note

This class is used in the content model of elements which reference names of people or organizations.

Appendix A.2.22 model.noteLike

model.noteLike groups globally-available note-like elements. [3.9. Notes, Annotation, and Indexing]
Moduletei — Formal specification
Used by
Membersnote

Appendix A.2.23 model.pLike

model.pLike groups paragraph-like elements.
Moduletei — Formal specification
Used by
Membersp

Appendix A.2.25 model.pPart.edit

model.pPart.edit groups phrase-level elements for simple editorial correction and transcription. [3.5. Simple Editorial Changes]
Moduletei — Formal specification
Used by
Membersmodel.pPart.editorial model.pPart.transcriptional

Appendix A.2.26 model.paraPart

Appendix A.2.27 model.persNamePart

model.persNamePart groups elements which form part of a personal name. [14.2.1. Personal Names]
Modulenamesdates — Formal specification
Used by
MembersaddName forename nameLink roleName surname

Appendix A.2.28 model.phrase

model.phrase groups elements which can occur at the level of individual words or phrases. [1.3. The TEI Class System]
Moduletei — Formal specification
Used by
Membersmodel.graphicLike[graphic media] model.highlighted[model.emphLike[term title] model.hiLike] model.lPart model.pPart.data[model.addressLike[affiliation email] model.dateLike[date time] model.measureLike[measure num unit] model.nameLike[model.nameLike.agent[name orgName persName] model.offsetLike model.persNamePart[addName forename nameLink roleName surname] model.placeStateLike[model.placeNamePart[placeName] state] idno]] model.pPart.edit[model.pPart.editorial model.pPart.transcriptional] model.pPart.msdesc model.phrase.xml model.ptrLike[ref] model.segLike[pc phr s seg w] model.specDescLike
Note

This class of elements can occur within paragraphs, list items, lines of verse, etc.

Appendix A.2.29 model.placeNamePart

model.placeNamePart groups elements which form part of a place name. [14.2.3. Place Names]
Moduletei — Formal specification
Used by
MembersplaceName

Appendix A.2.30 model.placeStateLike

model.placeStateLike groups elements which describe changing states of a place.
Moduletei — Formal specification
Used by
Membersmodel.placeNamePart[placeName] state

Appendix A.2.31 model.ptrLike

model.ptrLike groups elements used for purposes of location and reference. [3.7. Simple Links and Cross-References]
Moduletei — Formal specification
Used by
Membersref

Appendix A.2.32 model.segLike

model.segLike groups elements used for arbitrary segmentation. [17.3. Blocks, Segments, and Anchors 18.1. Linguistic Segment Categories]
Moduletei — Formal specification
Used by
Memberspc phr s seg w
Note

The principles on which segmentation is carried out, and any special codes or attribute values used, should be defined explicitly in the <segmentation> element of the <encodingDesc> within the associated TEI header.

Appendix A.3 Attribute classes

Appendix A.3.1 att.ascribed

att.ascribed provides attributes for elements representing speech or action that can be ascribed to a specific individual. [3.3.3. Quotation 8.3. Elements Unique to Spoken Texts]
Moduletei — Formal specification
Membersatt.ascribed.directed[kinesic u vocal] change incident setting
Attributes
whoindicates the person, or group of people, to whom the element content is ascribed.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
In the following example from Hamlet, speeches (<sp>) in the body of the play are linked to <role> elements in the <castList> using the who attribute.
<castItem type="role">  <role xml:id="Barnardo">Bernardo</role> </castItem> <castItem type="role"> @@ -4009,7 +4077,7 @@ <time when-iso="14">around two</time> <time when-iso="15,5">half past three</time>
All of the examples of the when attribute in the att.datable.w3c class are also valid with respect to this attribute.
He likes to be punctual. I said <q>  <time when-iso="12">around noon</time> -</q>, and he showed up at <time when-iso="12:00:00">12 O'clock</time> on the dot.
The second occurence of <time> could have been encoded with the when attribute, as 12:00:00 is a valid time with respect to the W3C XML Schema Part 2: Datatypes Second Edition specification. The first occurence could not.
notBefore-isospecifies the earliest possible date for the event in standard form, e.g. yyyy-mm-dd.
StatusOptional
Datatypeteidata.temporal.iso
notAfter-isospecifies the latest possible date for the event in standard form, e.g. yyyy-mm-dd.
StatusOptional
Datatypeteidata.temporal.iso
from-isoindicates the starting point of the period in standard form.
StatusOptional
Datatypeteidata.temporal.iso
to-isoindicates the ending point of the period in standard form.
StatusOptional
Datatypeteidata.temporal.iso
Note

The value of these attributes should be a normalized representation of the date, time, or combined date & time intended, in any of the standard formats specified by ISO 8601:2004, using the Gregorian calendar.

If both when-iso and dur-iso are specified, the values should be interpreted as indicating a span of time by its starting time (or date) and duration. That is,
<date when-iso="2007-06-01dur-iso="P8D"/>
indicates the same time period as
<date when-iso="2007-06-01/P8D"/>

In providing a ‘regularized’ form, no claim is made that the form in the source text is incorrect; the regularized form is simply that chosen as the main form for purposes of unifying variant forms under a single heading.

Appendix A.3.5 att.datable.w3c

att.datable.w3c provides attributes for normalization of elements that contain datable events conforming to the W3C XML Schema Part 2: Datatypes Second Edition. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Membersatt.datable[affiliation application birth change date death education event funder idno licence meeting name occupation orgName persName placeName relation resp sex state time title]
Attributes
whensupplies the value of the date or time in a standard form, e.g. yyyy-mm-dd.
StatusOptional
Datatypeteidata.temporal.w3c
Examples of W3C date, time, and date & time formats.
<p> +</q>, and he showed up at <time when-iso="12:00:00">12 O'clock</time> on the dot.
The second occurrence of <time> could have been encoded with the when attribute, as 12:00:00 is a valid time with respect to the W3C XML Schema Part 2: Datatypes Second Edition specification. The first occurrence could not.
notBefore-isospecifies the earliest possible date for the event in standard form, e.g. yyyy-mm-dd.
StatusOptional
Datatypeteidata.temporal.iso
notAfter-isospecifies the latest possible date for the event in standard form, e.g. yyyy-mm-dd.
StatusOptional
Datatypeteidata.temporal.iso
from-isoindicates the starting point of the period in standard form.
StatusOptional
Datatypeteidata.temporal.iso
to-isoindicates the ending point of the period in standard form.
StatusOptional
Datatypeteidata.temporal.iso
Note

The value of these attributes should be a normalized representation of the date, time, or combined date & time intended, in any of the standard formats specified by ISO 8601:2004, using the Gregorian calendar.

If both when-iso and dur-iso are specified, the values should be interpreted as indicating a span of time by its starting time (or date) and duration. That is,
<date when-iso="2007-06-01dur-iso="P8D"/>
indicates the same time period as
<date when-iso="2007-06-01/P8D"/>

In providing a ‘regularized’ form, no claim is made that the form in the source text is incorrect; the regularized form is simply that chosen as the main form for purposes of unifying variant forms under a single heading.

Appendix A.3.5 att.datable.w3c

att.datable.w3c provides attributes for normalization of elements that contain datable events conforming to the W3C XML Schema Part 2: Datatypes Second Edition. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Membersatt.datable[affiliation application birth change date death education event funder idno licence meeting name occupation orgName persName placeName relation resp sex state time title]
Attributes
whensupplies the value of the date or time in a standard form, e.g. yyyy-mm-dd.
StatusOptional
Datatypeteidata.temporal.w3c
Examples of W3C date, time, and date & time formats.
<p>  <date when="1945-10-24">24 Oct 45</date>  <date when="1996-09-24T07:25:00Z">September 24th, 1996 at 3:25 in the morning</date>  <time when="1999-01-04T20:42:00-05:00">Jan 4 1999 at 8 pm</time> @@ -4159,7 +4227,17 @@ <!-- inside an <entry> element: --> <usg type="domain" - valueDatcat="#domain.medical_and_health_sciences.medicine">Med.</usg>
In the Morais dictionary, the relevant domain labels are in the header, getting referenced inside the dictionary, from <usg> elements. The vocabulary used for dictionary-internal labelling is in turn anchored in the MorDigital controlled vocabulary service of the NOVA University of Lisbon – School of Social Sciences and Humanities (NOVA FCSH).
Note

The TEI Abstract Model can be expressed as a hierarchy of attribute-value matrices (AVMs) of various types and of various levels of complexity, nested or grouped in various ways. At the most abstract level, an AVM consists of an information container and the value (contents) of that container.

A simple example of an XML serialization of such structures is, on the one hand, the opening and closing tags that delimit and name the container, and, on the other, the content enclosed by the two tags that constitues the value. An analogous example is an attribute name and the value of that attribute.

In a TEI XML example of two equivalent serializations expressing the name-value pair <part-of-speech,common-noun>, namely <pos>commonNoun</pos> and pos="common-noun", one would classify the element <pos> and the attribute pos as containers (mapping onto the first member of the relevant name-value pair), while the character data content of <pos> or the value of pos would be seen as mapping onto the second member of the pair.

The att.datcat class provides means of addressing the containers and their values, while at the same time providing a way to interpret them in the context of external taxonomies or ontologies. Aligning e.g. both the <pos> element and the pos attribute with the same value of an external reference point (i.e., an entry in an agreed taxonomy) affirms the identity of the concept serialised by both the element container and the attribute container, and optionally provides a definition of that concept (in the case at hand, the concept part of speech).

The value of the att.datcat attributes should be a PID (persistent identifier) that points to a specific — and, ideally, shared — taxonomy or ontology. Among the resources that can, to a lesser or greater extent, be used as inventories of (more or less) standardized linguistic categories are the GOLD ontology, CLARIN CCR, OLiA, or TermWeb's DatCatInfo, and also the Universal Dependencies inventory, on the assumption that its URIs are going to persist. It is imaginable that a project may choose to address a local taxonomy store instead, but this risks losing the advantage of interchangeability with other projects.

Historically, datcat and valueDatcat originate from the (now obsolete) ISO 12620:2009 standard, describing the data model and procedures for a Data Category Registry (DCR). The current version of that standard, ISO 12620-1, does not standardize the serialization of pointers, merely mentioning the TEI att.datcat as an example.

Note that no constraint prevents the occurrence of a combination of att.datcat attributes: the <fDecl> element, which is a natural bearer of the targetDatcat attribute, is an instance of a specific modeling element, and, in principle, could be semantically fixed by an appropriate reference taxonomy of modeling devices.

Appendix A.3.7 att.declarable

att.declarable provides attributes for those elements in the TEI header which may be independently selected by means of the special purpose decls attribute. [16.3. Associating Contextual Information with a Text]
Moduletei — Formal specification
Membersavailability bibl correction editorialDecl equipment equipment hyphenation langUsage listEvent listOrg listPerson normalization particDesc projectDesc quotation recording segmentation settingDesc sourceDesc textClass
Attributes
defaultindicates whether or not this element is selected by default when its parent is selected.
StatusOptional
Datatypeteidata.truthValue
Legal values are:
true
This element is selected if its parent is selected
false
This element can only be selected explicitly, unless it is the only one of its kind, in which case it is selected if its parent is selected.[Default]
Note

The rules governing the association of declarable elements with individual parts of a TEI text are fully defined in chapter 16.3. Associating Contextual Information with a Text. Only one element of a particular type may have a default attribute with a value of true.

Appendix A.3.8 att.duration

att.duration provides attributes for normalization of elements that contain datable events.
Modulespoken — Formal specification
Membersatt.timed[gap incident kinesic media u vocal] date recording time
Attributes
Note

This ‘superclass’ provides attributes that can be used to provide normalized values of temporal information. By default, the attributes from the att.duration.w3c class are provided. If the module for names & dates is loaded, this class also provides attributes from the att.duration.iso class. In general, the possible values of attributes restricted to the W3C datatypes form a subset of those values available via the ISO 8601 standard. However, the greater expressiveness of the ISO datatypes is rarely needed, and there exists much greater software support for the W3C datatypes.

Appendix A.3.9 att.duration.iso

att.duration.iso provides attributes for recording normalized temporal durations. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Membersatt.duration[att.timed[gap incident kinesic media u vocal] date recording time]
Attributes
dur-iso(duration) indicates the length of this element in time.
StatusOptional
Datatypeteidata.duration.iso
Note

If both when and dur or dur-iso are specified, the values should be interpreted as indicating a span of time by its starting time (or date) and duration. In order to represent a time range by a duration and its ending time the when-iso attribute must be used.

In providing a ‘regularized’ form, no claim is made that the form in the source text is incorrect; the regularized form is simply that chosen as the main form for purposes of unifying variant forms under a single heading.

Appendix A.3.10 att.duration.w3c

att.duration.w3c provides attributes for recording normalized temporal durations. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Membersatt.duration[att.timed[gap incident kinesic media u vocal] date recording time]
Attributes
dur(duration) indicates the length of this element in time.
StatusOptional
Datatypeteidata.duration.w3c
Note

If both when and dur are specified, the values should be interpreted as indicating a span of time by its starting time (or date) and duration. In order to represent a time range by a duration and its ending time the when-iso attribute must be used.

In providing a ‘regularized’ form, no claim is made that the form in the source text is incorrect; the regularized form is simply that chosen as the main form for purposes of unifying variant forms under a single heading.

Appendix A.3.11 att.fragmentable

att.fragmentable provides attributes for representing fragmentation of a structural element, typically as a consequence of some overlapping hierarchy.
Moduletei — Formal specification
Membersatt.divLike[div] att.segLike[pc phr s seg w] p
Attributes
partspecifies whether or not its parent element is fragmented in some way, typically by some other overlapping structure: for example a speech which is divided between two or more verse stanzas, a paragraph which is split across a page division, a verse line which is divided between two speakers.
StatusOptional
Datatypeteidata.enumerated
Legal values are:
Y
(yes) the element is fragmented in some (unspecified) respect
N
(no) the element is not fragmented, or no claim is made as to its completeness[Default]
I
(initial) this is the initial part of a fragmented element
M
(medial) this is a medial part of a fragmented element
F
(final) this is the final part of a fragmented element
Note

The values I, M, or F should be used only where it is clear how the element may be reconstituted.

Appendix A.3.12 att.global

att.global provides attributes common to all elements in the TEI encoding scheme. [1.3.1.1. Global Attributes]
Moduletei — Formal specification
MembersTEI addName affiliation appInfo application availability bibl birth body catDesc catRef category change classDecl correction date death desc div edition editionStmt editorialDecl education email encodingDesc equipment equipment event extent figure fileDesc forename funder gap graphic head hyphenation idno incident kinesic label langUsage language licence link linkGrp listEvent listOrg listPerson listPrefixDef listRelation measure media meeting name nameLink namespace normalization note num occupation org orgName p particDesc pb pc persName person phr placeName prefixDef profileDesc projectDesc pubPlace publicationStmt publisher quotation recording recordingStmt ref relation resp respStmt revisionDesc roleName s seg segmentation setting settingDesc sex sourceDesc state surname tagUsage tagsDecl taxonomy teiCorpus teiHeader term text textClass time title titleStmt u unit vocal w
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
StatusOptional
DatatypeID
Note

The xml:id attribute may be used to specify a canonical reference for an element; see section 3.11. Reference Systems.

n(number) gives a number (or other label) for an element, which is not necessarily unique within the document.
StatusOptional
Datatypeteidata.text
Note

The value of this attribute is always understood to be a single token, even if it contains space or other punctuation characters, and need not be composed of numbers only. It is typically used to specify the numbering of chapters, sections, list items, etc.; it may also be used in the specification of a standard reference system for the text.

xml:lang(language) indicates the language of the element content using a ‘tag’ generated according to BCP 47.
StatusOptional
Datatypeteidata.language
<p> … The consequences of + valueDatcat="#domain.medical_and_health_sciences.medicine">Med.</usg>
In the Morais dictionary, the relevant domain labels are in the header, getting referenced inside the dictionary, from <usg> elements. The vocabulary used for dictionary-internal labelling is in turn anchored in the MorDigital controlled vocabulary service of the NOVA University of Lisbon – School of Social Sciences and Humanities (NOVA FCSH).
Note

The TEI Abstract Model can be expressed as a hierarchy of attribute-value matrices (AVMs) of various types and of various levels of complexity, nested or grouped in various ways. At the most abstract level, an AVM consists of an information container and the value (contents) of that container.

A simple example of an XML serialization of such structures is, on the one hand, the opening and closing tags that delimit and name the container, and, on the other, the content enclosed by the two tags that constitues the value. An analogous example is an attribute name and the value of that attribute.

In a TEI XML example of two equivalent serializations expressing the name-value pair <part-of-speech,common-noun>, namely <pos>commonNoun</pos> and pos="common-noun", one would classify the element <pos> and the attribute pos as containers (mapping onto the first member of the relevant name-value pair), while the character data content of <pos> or the value of pos would be seen as mapping onto the second member of the pair.

The att.datcat class provides means of addressing the containers and their values, while at the same time providing a way to interpret them in the context of external taxonomies or ontologies. Aligning e.g. both the <pos> element and the pos attribute with the same value of an external reference point (i.e., an entry in an agreed taxonomy) affirms the identity of the concept serialised by both the element container and the attribute container, and optionally provides a definition of that concept (in the case at hand, the concept part of speech).

The value of the att.datcat attributes should be a PID (persistent identifier) that points to a specific — and, ideally, shared — taxonomy or ontology. Among the resources that can, to a lesser or greater extent, be used as inventories of (more or less) standardized linguistic categories are the GOLD ontology, CLARIN CCR, OLiA, or TermWeb's DatCatInfo, and also the Universal Dependencies inventory, on the assumption that its URIs are going to persist. It is imaginable that a project may choose to address a local taxonomy store instead, but this risks losing the advantage of interchangeability with other projects.

Historically, datcat and valueDatcat originate from the (now obsolete) ISO 12620:2009 standard, describing the data model and procedures for a Data Category Registry (DCR). The current version of that standard, ISO 12620-1, does not standardize the serialization of pointers, merely mentioning the TEI att.datcat as an example.

Note that no constraint prevents the occurrence of a combination of att.datcat attributes: the <fDecl> element, which is a natural bearer of the targetDatcat attribute, is an instance of a specific modeling element, and, in principle, could be semantically fixed by an appropriate reference taxonomy of modeling devices.

Appendix A.3.7 att.declarable

att.declarable provides attributes for those elements in the TEI header which may be independently selected by means of the special purpose decls attribute. [16.3. Associating Contextual Information with a Text]
Moduletei — Formal specification
Membersavailability bibl correction editorialDecl equipment equipment hyphenation langUsage listEvent listOrg listPerson normalization particDesc projectDesc quotation recording segmentation settingDesc sourceDesc textClass
Attributes
defaultindicates whether or not this element is selected by default when its parent is selected.
StatusOptional
Datatypeteidata.truthValue
Legal values are:
true
This element is selected if its parent is selected
false
This element can only be selected explicitly, unless it is the only one of its kind, in which case it is selected if its parent is selected.[Default]
Schematron
+<sch:pattern id="declarable" abstract="true"> +<sch:rule context="$tde[ ancestor::tei:teiHeader and following-sibling::$tde and not( + preceding-sibling::$tde ) ]"> + <sch:report test="../child::$tde[ not( @xml:id ) ]"> When there is more than one <sch:name/>, each must have an @xml:id + </sch:report> + <sch:assert test="count( ../child::$tde[ normalize-space( @default ) = ('1','true') + ] ) eq 1"> When there is more than one <sch:name/>, one and only one must have a @default of 'true'. + </sch:assert> +</sch:rule> +</sch:pattern>
Note

The rules governing the association of declarable elements with individual parts of a TEI text are fully defined in chapter 16.3. Associating Contextual Information with a Text. Only one element of a particular type may have a default attribute with a value of true.

Appendix A.3.8 att.duration

att.duration provides attributes for normalization of elements that contain datable events.
Modulespoken — Formal specification
Membersatt.timed[gap incident kinesic media u vocal] date recording time
Attributes
Note

This ‘superclass’ provides attributes that can be used to provide normalized values of temporal information. By default, the attributes from the att.duration.w3c class are provided. If the module for names & dates is loaded, this class also provides attributes from the att.duration.iso class. In general, the possible values of attributes restricted to the W3C datatypes form a subset of those values available via the ISO 8601 standard. However, the greater expressiveness of the ISO datatypes is rarely needed, and there exists much greater software support for the W3C datatypes.

Appendix A.3.9 att.duration.iso

att.duration.iso provides attributes for recording normalized temporal durations. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Membersatt.duration[att.timed[gap incident kinesic media u vocal] date recording time]
Attributes
dur-iso(duration) indicates the length of this element in time.
StatusOptional
Datatypeteidata.duration.iso
Note

If both when and dur or dur-iso are specified, the values should be interpreted as indicating a span of time by its starting time (or date) and duration. In order to represent a time range by a duration and its ending time the when-iso attribute must be used.

In providing a ‘regularized’ form, no claim is made that the form in the source text is incorrect; the regularized form is simply that chosen as the main form for purposes of unifying variant forms under a single heading.

Appendix A.3.10 att.duration.w3c

att.duration.w3c provides attributes for recording normalized temporal durations. [3.6.4. Dates and Times 14.4. Dates]
Moduletei — Formal specification
Membersatt.duration[att.timed[gap incident kinesic media u vocal] date recording time]
Attributes
dur(duration) indicates the length of this element in time.
StatusOptional
Datatypeteidata.duration.w3c
Note

If both when and dur are specified, the values should be interpreted as indicating a span of time by its starting time (or date) and duration. In order to represent a time range by a duration and its ending time the when-iso attribute must be used.

In providing a ‘regularized’ form, no claim is made that the form in the source text is incorrect; the regularized form is simply that chosen as the main form for purposes of unifying variant forms under a single heading.

Appendix A.3.11 att.fragmentable

att.fragmentable provides attributes for representing fragmentation of a structural element, typically as a consequence of some overlapping hierarchy.
Moduletei — Formal specification
Membersatt.divLike[div] att.segLike[pc phr s seg w] p
Attributes
partspecifies whether or not its parent element is fragmented in some way, typically by some other overlapping structure: for example a speech which is divided between two or more verse stanzas, a paragraph which is split across a page division, a verse line which is divided between two speakers.
StatusOptional
Datatypeteidata.enumerated
Legal values are:
Y
(yes) the element is fragmented in some (unspecified) respect
N
(no) the element is not fragmented, or no claim is made as to its completeness[Default]
I
(initial) this is the initial part of a fragmented element
M
(medial) this is a medial part of a fragmented element
F
(final) this is the final part of a fragmented element
Note

The values I, M, or F should be used only where it is clear how the element may be reconstituted.

Appendix A.3.12 att.global

att.global provides attributes common to all elements in the TEI encoding scheme. [1.3.1.1. Global Attributes]
Moduletei — Formal specification
MembersTEI addName affiliation appInfo application availability bibl birth body catDesc catRef category change classDecl correction date death desc div edition editionStmt editorialDecl education email encodingDesc equipment equipment event extent figure fileDesc forename funder gap graphic head hyphenation idno incident kinesic label langUsage language licence link linkGrp listEvent listOrg listPerson listPrefixDef listRelation measure media meeting name nameLink namespace normalization note num occupation org orgName p particDesc pb pc persName person phr placeName prefixDef profileDesc projectDesc pubPlace publicationStmt publisher quotation recording recordingStmt ref relation resp respStmt revisionDesc roleName s seg segmentation setting settingDesc sex sourceDesc state surname tagUsage tagsDecl taxonomy teiCorpus teiHeader term text textClass time title titleStmt u unit vocal w
Attributes
xml:id(identifier) provides a unique identifier for the element bearing the attribute.
StatusOptional
DatatypeID
Note

The xml:id attribute may be used to specify a canonical reference for an element; see section 3.11. Reference Systems.

n(number) gives a number (or other label) for an element, which is not necessarily unique within the document.
StatusOptional
Datatypeteidata.text
Note

The value of this attribute is always understood to be a single token, even if it contains space or other punctuation characters, and need not be composed of numbers only. It is typically used to specify the numbering of chapters, sections, list items, etc.; it may also be used in the specification of a standard reference system for the text.

xml:lang(language) indicates the language of the element content using a ‘tag’ generated according to BCP 47.
StatusOptional
Datatypeteidata.language
<p> … The consequences of this rapid depopulation were the loss of the last <foreign xml:lang="rap">ariki</foreign> or chief (Routledge 1920:205,210) and their connections to @@ -4226,7 +4304,7 @@      allegorical character in mayoral shows.   </p>  </note> -</person>
In this example, a <place> element containing information about the city of London is linked with two <person> elements in a literary personography. This correspondence represents a slightly looser relationship than the one in the preceding example; there is no sense in which an allegorical character could be substituted for the physical city, or vice versa, but there is obviously a correspondence between them.
synch(synchronous) points to elements that are synchronous with the current element.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
nextpoints to the next element of a virtual aggregate of which the current element is part.
StatusOptional
Datatypeteidata.pointer
Note

It is recommended that the element indicated be of the same type as the element bearing this attribute.

prev(previous) points to the previous element of a virtual aggregate of which the current element is part.
StatusOptional
Datatypeteidata.pointer
Note

It is recommended that the element indicated be of the same type as the element bearing this attribute.

Appendix A.3.15 att.global.rendition

att.global.rendition provides rendering attributes common to all elements in the TEI encoding scheme. [1.3.1.1.3. Rendition Indicators]
Moduletei — Formal specification
Membersatt.global[TEI addName affiliation appInfo application availability bibl birth body catDesc catRef category change classDecl correction date death desc div edition editionStmt editorialDecl education email encodingDesc equipment equipment event extent figure fileDesc forename funder gap graphic head hyphenation idno incident kinesic label langUsage language licence link linkGrp listEvent listOrg listPerson listPrefixDef listRelation measure media meeting name nameLink namespace normalization note num occupation org orgName p particDesc pb pc persName person phr placeName prefixDef profileDesc projectDesc pubPlace publicationStmt publisher quotation recording recordingStmt ref relation resp respStmt revisionDesc roleName s seg segmentation setting settingDesc sex sourceDesc state surname tagUsage tagsDecl taxonomy teiCorpus teiHeader term text textClass time title titleStmt u unit vocal w]
Attributes
rend(rendition) indicates how the element in question was rendered or presented in the source text.
StatusOptional
Datatype1–∞ occurrences of teidata.word separated by whitespace
<head rend="align(center) case(allcaps)"> +</person>
In this example, a <place> element containing information about the city of London is linked with two <person> elements in a literary personography. This correspondence represents a slightly looser relationship than the one in the preceding example; there is no sense in which an allegorical character could be substituted for the physical city, or vice versa, but there is obviously a correspondence between them.
synch(synchronous) points to elements that are synchronous with the current element.
StatusOptional
Datatype1–∞ occurrences of teidata.pointer separated by whitespace
next(next) points to the next element of a virtual aggregate of which the current element is part.
StatusOptional
Datatypeteidata.pointer
Note

It is recommended that the element indicated be of the same type as the element bearing this attribute.

prev(previous) points to the previous element of a virtual aggregate of which the current element is part.
StatusOptional
Datatypeteidata.pointer
Note

It is recommended that the element indicated be of the same type as the element bearing this attribute.

Appendix A.3.15 att.global.rendition

att.global.rendition provides rendering attributes common to all elements in the TEI encoding scheme. [1.3.1.1.3. Rendition Indicators]
Moduletei — Formal specification
Membersatt.global[TEI addName affiliation appInfo application availability bibl birth body catDesc catRef category change classDecl correction date death desc div edition editionStmt editorialDecl education email encodingDesc equipment equipment event extent figure fileDesc forename funder gap graphic head hyphenation idno incident kinesic label langUsage language licence link linkGrp listEvent listOrg listPerson listPrefixDef listRelation measure media meeting name nameLink namespace normalization note num occupation org orgName p particDesc pb pc persName person phr placeName prefixDef profileDesc projectDesc pubPlace publicationStmt publisher quotation recording recordingStmt ref relation resp respStmt revisionDesc roleName s seg segmentation setting settingDesc sex sourceDesc state surname tagUsage tagsDecl taxonomy teiCorpus teiHeader term text textClass time title titleStmt u unit vocal w]
Attributes
rend(rendition) indicates how the element in question was rendered or presented in the source text.
StatusOptional
Datatype1–∞ occurrences of teidata.word separated by whitespace
<head rend="align(center) case(allcaps)">  <lb/>To The <lb/>Duchesse <lb/>of <lb/>Newcastle, <lb/>On Her <lb/>  <hi rend="case(mixed)">New Blazing-World</hi>. @@ -4384,7 +4462,7 @@ </div>
Note

The type attribute is present on a number of elements, not all of which are members of att.typed, usually because these elements restrict the possible values for the attribute in a specific way.

subtype(subtype) provides a sub-categorization of the element, if needed.
StatusOptional
Datatypeteidata.enumerated
Note

The subtype attribute may be used to provide any sub-classification for the element additional to that provided by its type attribute.

Schematron
<sch:rule context="tei:*[@subtype]"> <sch:assert test="@type">The <sch:name/> element should not be categorized in detail with @subtype unless also categorized in general with @type</sch:assert> -</sch:rule>
Note

When appropriate, values from an established typology should be used. Alternatively a typology may be defined in the associated TEI header. If values are to be taken from a project-specific list, this should be defined using the <valList> element in the project-specific schema description, as described in 24.3.1.3. Modification of Attribute and Attribute Value Lists .

Appendix A.4 Datatypes

Appendix A.4.1 teidata.certainty

teidata.certainty defines the range of attribute values expressing a degree of certainty.
Moduletei — Formal specification
Used by
Content model
+</sch:rule>
Note

When appropriate, values from an established typology should be used. Alternatively a typology may be defined in the associated TEI header. If values are to be taken from a project-specific list, this should be defined using the <valList> element in the project-specific schema description, as described in 24.3.1.3. Modification of Attribute and Attribute Value Lists .

Appendix A.4 Datatypes

Appendix A.4.1 teidata.certainty

teidata.certainty defines the range of attribute values expressing a degree of certainty.
Moduletei — Formal specification
Used by
Content model
 <content>
  <valList type="closed">
   <valItem ident="high"/>
@@ -4393,29 +4471,29 @@
   <valItem ident="unknown"/>
  </valList>
 </content>
-    
Declaration
-tei_teidata.certainty = "high" | "medium" | "low" | "unknown"
Note

Certainty may be expressed by one of the predefined symbolic values high, medium, or low. The value unknown should be used in cases where the encoder does not wish to assert an opinion about the matter.

Appendix A.4.2 teidata.count

teidata.count defines the range of attribute values used for a non-negative integer value used as a count.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.certainty = "high" | "medium" | "low" | "unknown"
Note

Certainty may be expressed by one of the predefined symbolic values high, medium, or low. The value unknown should be used in cases where the encoder does not wish to assert an opinion about the matter.

Appendix A.4.2 teidata.count

teidata.count defines the range of attribute values used for a non-negative integer value used as a count.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef name="nonNegativeInteger"/>
 </content>
-    
Declaration
-tei_teidata.count = xsd:nonNegativeInteger
Note

Any positive integer value or zero is permitted

Appendix A.4.3 teidata.duration.iso

teidata.duration.iso defines the range of attribute values available for representation of a duration in time using ISO 8601 standard formats.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.count = xsd:nonNegativeInteger
Note

Any positive integer value or zero is permitted

Appendix A.4.3 teidata.duration.iso

teidata.duration.iso defines the range of attribute values available for representation of a duration in time using ISO 8601 standard formats.
Moduletei — Formal specification
Used by
Content model
 <content>
  <dataRef name="token"
   restriction="[0-9.,DHMPRSTWYZ/:+\-]+"/>
 </content>
-    
Declaration
-tei_teidata.duration.iso = token { pattern = "[0-9.,DHMPRSTWYZ/:+\-]+" }
Example
<time dur-iso="PT0,75H">three-quarters of an hour</time>
Example
<date dur-iso="P1,5D">a day and a half</date>
Example
<date dur-iso="P14D">a fortnight</date>
Example
<time dur-iso="PT0.02S">20 ms</time>
Note

A duration is expressed as a sequence of number-letter pairs, preceded by the letter P; the letter gives the unit and may be Y (year), M (month), D (day), H (hour), M (minute), or S (second), in that order. The numbers are all unsigned integers, except for the last, which may have a decimal component (using either . or , as the decimal point; the latter is preferred). If any number is 0, then that number-letter pair may be omitted. If any of the H (hour), M (minute), or S (second) number-letter pairs are present, then the separator T must precede the first ‘time’ number-letter pair.

For complete details, see ISO 8601 Data elements and interchange formats — Information interchange — Representation of dates and times.

Appendix A.4.4 teidata.duration.w3c

teidata.duration.w3c defines the range of attribute values available for representation of a duration in time using W3C datatypes.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.duration.iso = token { pattern = "[0-9.,DHMPRSTWYZ/:+\-]+" }
Example
<time dur-iso="PT0,75H">three-quarters of an hour</time>
Example
<date dur-iso="P1,5D">a day and a half</date>
Example
<date dur-iso="P14D">a fortnight</date>
Example
<time dur-iso="PT0.02S">20 ms</time>
Note

A duration is expressed as a sequence of number-letter pairs, preceded by the letter P; the letter gives the unit and may be Y (year), M (month), D (day), H (hour), M (minute), or S (second), in that order. The numbers are all unsigned integers, except for the last, which may have a decimal component (using either . or , as the decimal point; the latter is preferred). If any number is 0, then that number-letter pair may be omitted. If any of the H (hour), M (minute), or S (second) number-letter pairs are present, then the separator T must precede the first ‘time’ number-letter pair.

For complete details, see ISO 8601 Data elements and interchange formats — Information interchange — Representation of dates and times.

Appendix A.4.4 teidata.duration.w3c

teidata.duration.w3c defines the range of attribute values available for representation of a duration in time using W3C datatypes.
Moduletei — Formal specification
Used by
Content model
 <content>
  <dataRef name="duration"/>
 </content>
-    
Declaration
-tei_teidata.duration.w3c = xsd:duration
Example
<time dur="PT45M">forty-five minutes</time>
Example
<date dur="P1DT12H">a day and a half</date>
Example
<date dur="P7D">a week</date>
Example
<time dur="PT0.02S">20 ms</time>
Note

A duration is expressed as a sequence of number-letter pairs, preceded by the letter P; the letter gives the unit and may be Y (year), M (month), D (day), H (hour), M (minute), or S (second), in that order. The numbers are all unsigned integers, except for the S number, which may have a decimal component (using . as the decimal point). If any number is 0, then that number-letter pair may be omitted. If any of the H (hour), M (minute), or S (second) number-letter pairs are present, then the separator T must precede the first ‘time’ number-letter pair.

For complete details, see the W3C specification.

Appendix A.4.5 teidata.enumerated

teidata.enumerated defines the range of attribute values expressed as a single XML name taken from a list of documented possibilities.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.duration.w3c = xsd:duration
Example
<time dur="PT45M">forty-five minutes</time>
Example
<date dur="P1DT12H">a day and a half</date>
Example
<date dur="P7D">a week</date>
Example
<time dur="PT0.02S">20 ms</time>
Note

A duration is expressed as a sequence of number-letter pairs, preceded by the letter P; the letter gives the unit and may be Y (year), M (month), D (day), H (hour), M (minute), or S (second), in that order. The numbers are all unsigned integers, except for the S number, which may have a decimal component (using . as the decimal point). If any number is 0, then that number-letter pair may be omitted. If any of the H (hour), M (minute), or S (second) number-letter pairs are present, then the separator T must precede the first ‘time’ number-letter pair.

For complete details, see the W3C specification.

Appendix A.4.5 teidata.enumerated

teidata.enumerated defines the range of attribute values expressed as a single XML name taken from a list of documented possibilities.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef key="teidata.word"/>
 </content>
-    
Declaration
-tei_teidata.enumerated = teidata.word
Note

Attributes using this datatype must contain a single ‘word’ which contains only letters, digits, punctuation characters, or symbols: thus it cannot include whitespace.

Typically, the list of documented possibilities will be provided (or exemplified) by a value list in the associated attribute specification, expressed with a <valList> element.

Appendix A.4.6 teidata.language

teidata.language defines the range of attribute values used to identify a particular combination of human language and writing system. [6.1. Language Identification]
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.enumerated = teidata.word
Note

Attributes using this datatype must contain a single ‘word’ which contains only letters, digits, punctuation characters, or symbols: thus it cannot include whitespace.

Typically, the list of documented possibilities will be provided (or exemplified) by a value list in the associated attribute specification, expressed with a <valList> element.

Appendix A.4.6 teidata.language

teidata.language defines the range of attribute values used to identify a particular combination of human language and writing system. [6.1. Language Identification]
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <alternate>
   <dataRef name="language"/>
@@ -4424,13 +4502,13 @@
   </valList>
  </alternate>
 </content>
-    
Declaration
-tei_teidata.language = xsd:language | ( "" )
Note

The values for this attribute are language ‘tags’ as defined in BCP 47. Currently BCP 47 comprises RFC 5646 and RFC 4647; over time, other IETF documents may succeed these as the best current practice.

A ‘language tag’, per BCP 47, is assembled from a sequence of components or subtags separated by the hyphen character (-, U+002D). The tag is made of the following subtags, in the following order. Every subtag except the first is optional. If present, each occurs only once, except the fourth and fifth components (variant and extension), which are repeatable.

language
The IANA-registered code for the language. This is almost always the same as the ISO 639 2-letter language code if there is one. The list of available registered language subtags can be found at https://www.iana.org/assignments/language-subtag-registry. It is recommended that this code be written in lower case.
script
The ISO 15924 code for the script. These codes consist of 4 letters, and it is recommended they be written with an initial capital, the other three letters in lower case. The canonical list of codes is maintained by the Unicode Consortium, and is available at https://unicode.org/iso15924/iso15924-codes.html. The IETF recommends this code be omitted unless it is necessary to make a distinction you need.
region
Either an ISO 3166 country code or a UN M.49 region code that is registered with IANA (not all such codes are registered, e.g. UN codes for economic groupings or codes for countries for which there is already an ISO 3166 2-letter code are not registered). The former consist of 2 letters, and it is recommended they be written in upper case; the list of codes can be searched or browsed at https://www.iso.org/obp/ui/#search/code/. The latter consist of 3 digits; the list of codes can be found at http://unstats.un.org/unsd/methods/m49/m49.htm.
variant
An IANA-registered variation. These codes ‘are used to indicate additional, well-recognized variations that define a language or its dialects that are not covered by other available subtags’.
extension
An extension has the format of a single letter followed by a hyphen followed by additional subtags. There are currently only two extensions in use. Extension T indicates that the content was transformed. For example en-t-it could be used for content in English that was translated from Italian. Extension T is described in the informational RFC 6497. Extension U can be used to embed a variety of locale attributes. It is described in the informational RFC 6067.
private use
An extension that uses the initial subtag of the single letter x (i.e., starts with x-) has no meaning except as negotiated among the parties involved. These should be used with great care, since they interfere with the interoperability that use of RFC 4646 is intended to promote. In order for a document that makes use of these subtags to be TEI-conformant, a corresponding <language> element must be present in the TEI header.

There are two exceptions to the above format. First, there are language tags in the IANA registry that do not match the above syntax, but are present because they have been ‘grandfathered’ from previous specifications.

Second, an entire language tag can consist of only a private use subtag. These tags start with x-, and do not need to follow any further rules established by the IETF and endorsed by these Guidelines. Like all language tags that make use of private use subtags, the language in question must be documented in a corresponding <language> element in the TEI header.

Examples include

sn
Shona
zh-TW
Taiwanese
zh-Hant-HK
Chinese written in traditional script as used in Hong Kong
en-SL
English as spoken in Sierra Leone
pl
Polish
es-MX
Spanish as spoken in Mexico
es-419
Spanish as spoken in Latin America

The W3C Internationalization Activity has published a useful introduction to BCP 47, Language tags in HTML and XML.

Appendix A.4.7 teidata.name

teidata.name defines the range of attribute values expressed as an XML Name.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.language = xsd:language | ( "" )
Note

The values for this attribute are language ‘tags’ as defined in BCP 47. Currently BCP 47 comprises RFC 5646 and RFC 4647; over time, other IETF documents may succeed these as the best current practice.

A ‘language tag’, per BCP 47, is assembled from a sequence of components or subtags separated by the hyphen character (-, U+002D). The tag is made of the following subtags, in the following order. Every subtag except the first is optional. If present, each occurs only once, except the fourth and fifth components (variant and extension), which are repeatable.

language
The IANA-registered code for the language. This is almost always the same as the ISO 639 2-letter language code if there is one. The list of available registered language subtags can be found at https://www.iana.org/assignments/language-subtag-registry. It is recommended that this code be written in lower case.
script
The ISO 15924 code for the script. These codes consist of 4 letters, and it is recommended they be written with an initial capital, the other three letters in lower case. The canonical list of codes is maintained by the Unicode Consortium, and is available at https://unicode.org/iso15924/iso15924-codes.html. The IETF recommends this code be omitted unless it is necessary to make a distinction you need.
region
Either an ISO 3166 country code or a UN M.49 region code that is registered with IANA (not all such codes are registered, e.g. UN codes for economic groupings or codes for countries for which there is already an ISO 3166 2-letter code are not registered). The former consist of 2 letters, and it is recommended they be written in upper case; the list of codes can be searched or browsed at https://www.iso.org/obp/ui/#search/code/. The latter consist of 3 digits; the list of codes can be found at http://unstats.un.org/unsd/methods/m49/m49.htm.
variant
An IANA-registered variation. These codes ‘are used to indicate additional, well-recognized variations that define a language or its dialects that are not covered by other available subtags’.
extension
An extension has the format of a single letter followed by a hyphen followed by additional subtags. There are currently only two extensions in use. Extension T indicates that the content was transformed. For example en-t-it could be used for content in English that was translated from Italian. Extension T is described in the informational RFC 6497. Extension U can be used to embed a variety of locale attributes. It is described in the informational RFC 6067.
private use
An extension that uses the initial subtag of the single letter x (i.e., starts with x-) has no meaning except as negotiated among the parties involved. These should be used with great care, since they interfere with the interoperability that use of RFC 4646 is intended to promote. In order for a document that makes use of these subtags to be TEI-conformant, a corresponding <language> element must be present in the TEI header.

There are two exceptions to the above format. First, there are language tags in the IANA registry that do not match the above syntax, but are present because they have been ‘grandfathered’ from previous specifications.

Second, an entire language tag can consist of only a private use subtag. These tags start with x-, and do not need to follow any further rules established by the IETF and endorsed by these Guidelines. Like all language tags that make use of private use subtags, the language in question must be documented in a corresponding <language> element in the TEI header.

Examples include

sn
Shona
zh-TW
Taiwanese
zh-Hant-HK
Chinese written in traditional script as used in Hong Kong
en-SL
English as spoken in Sierra Leone
pl
Polish
es-MX
Spanish as spoken in Mexico
es-419
Spanish as spoken in Latin America

The W3C Internationalization Activity has published a useful introduction to BCP 47, Language tags in HTML and XML.

Appendix A.4.7 teidata.name

teidata.name defines the range of attribute values expressed as an XML Name.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef name="Name"/>
 </content>
-    
Declaration
-tei_teidata.name = xsd:Name
Note

Attributes using this datatype must contain a single word which follows the rules defining a legal XML name (see https://www.w3.org/TR/REC-xml/#dt-name): for example they cannot include whitespace or begin with digits.

Appendix A.4.8 teidata.numeric

teidata.numeric defines the range of attribute values used for numeric values.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.name = xsd:Name
Note

Attributes using this datatype must contain a single word which follows the rules defining a legal XML name (see https://www.w3.org/TR/REC-xml/#dt-name): for example they cannot include whitespace or begin with digits.

Appendix A.4.8 teidata.numeric

teidata.numeric defines the range of attribute values used for numeric values.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <alternate>
   <dataRef name="double"/>
@@ -4439,63 +4517,63 @@
   <dataRef name="decimal"/>
  </alternate>
 </content>
-    
Declaration
+    
Declaration
 tei_teidata.numeric =
-   xsd:double | token { pattern = "(\-?[\d]+/\-?[\d]+)" } | xsd:decimal
Note

Any numeric value, represented as a decimal number, in floating point format, or as a ratio.

To represent a floating point number, expressed in scientific notation, ‘E notation’, a variant of ‘exponential notation’, may be used. In this format, the value is expressed as two numbers separated by the letter E. The first number, the significand (sometimes called the mantissa) is given in decimal format, while the second is an integer. The value is obtained by multiplying the mantissa by 10 the number of times indicated by the integer. Thus the value represented in decimal notation as 1000.0 might be represented in scientific notation as 10E3.

A value expressed as a ratio is represented by two integer values separated by a solidus (/) character. Thus, the value represented in decimal notation as 0.5 might be represented as a ratio by the string 1/2.

Appendix A.4.9 teidata.outputMeasurement

teidata.outputMeasurement defines a range of values for use in specifying the size of an object that is intended for display.
Moduletei — Formal specification
Used by
Content model
+   xsd:double | token { pattern = "(\-?[\d]+/\-?[\d]+)" } | xsd:decimal
Note

Any numeric value, represented as a decimal number, in floating point format, or as a ratio.

To represent a floating point number, expressed in scientific notation, ‘E notation’, a variant of ‘exponential notation’, may be used. In this format, the value is expressed as two numbers separated by the letter E. The first number, the significand (sometimes called the mantissa) is given in decimal format, while the second is an integer. The value is obtained by multiplying the mantissa by 10 the number of times indicated by the integer. Thus the value represented in decimal notation as 1000.0 might be represented in scientific notation as 10E3.

A value expressed as a ratio is represented by two integer values separated by a solidus (/) character. Thus, the value represented in decimal notation as 0.5 might be represented as a ratio by the string 1/2.

Appendix A.4.9 teidata.outputMeasurement

teidata.outputMeasurement defines a range of values for use in specifying the size of an object that is intended for display.
Moduletei — Formal specification
Used by
Content model
 <content>
  <dataRef name="token"
   restriction="[\-+]?\d+(\.\d+)?(%|cm|mm|in|pt|pc|px|em|ex|ch|rem|vw|vh|vmin|vmax)"/>
 </content>
-    
Declaration
+    
Declaration
 tei_teidata.outputMeasurement =
    token
    {
       pattern = "[\-+]?\d+(\.\d+)?(%|cm|mm|in|pt|pc|px|em|ex|ch|rem|vw|vh|vmin|vmax)"
-   }
Example
<figure> + }
Example
<figure>  <head>The TEI Logo</head>  <figDesc>Stylized yellow angle brackets with the letters <mentioned>TEI</mentioned> in    between and <mentioned>text encoding initiative</mentioned> underneath, all on a white    background.</figDesc>  <graphic height="600pxwidth="600px"   url="http://www.tei-c.org/logos/TEI-600.jpg"/> -</figure>
Note

These values map directly onto the values used by XSL-FO and CSS. For definitions of the units see those specifications; at the time of this writing the most complete list is in the CSS3 working draft.

Appendix A.4.10 teidata.pattern

teidata.pattern defines attribute values which are expressed as a regular expression.
Moduletei — Formal specification
Used by
Element:
Content model
+</figure>
Note

These values map directly onto the values used by XSL-FO and CSS. For definitions of the units see those specifications; at the time of this writing the most complete list is in the CSS3 working draft.

Appendix A.4.10 teidata.pattern

teidata.pattern defines attribute values which are expressed as a regular expression.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef name="token"/>
 </content>
-    
Declaration
-tei_teidata.pattern = token
Note
A regular expression, often called a pattern, is an expression that describes a set of strings. They are usually used to give a concise description of a set, without having to list all elements. For example, the set containing the three strings Handel, Händel, and Haendel can be described by the pattern H(ä|ae?)ndel (or alternatively, it is said that the pattern H(ä|ae?)ndel matches each of the three strings)
Wikipedia

This TEI datatype is mapped to the XSD token datatype, and may therefore contain any string of characters. However, it is recommended that the value used conform to the particular flavour of regular expression syntax supported by XSD Schema.

Appendix A.4.11 teidata.pointer

teidata.pointer defines the range of attribute values used to provide a single URI, absolute or relative, pointing to some other resource, either within the current document or elsewhere.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.pattern = token
Note
A regular expression, often called a pattern, is an expression that describes a set of strings. They are usually used to give a concise description of a set, without having to list all elements. For example, the set containing the three strings Handel, Händel, and Haendel can be described by the pattern H(ä|ae?)ndel (or alternatively, it is said that the pattern H(ä|ae?)ndel matches each of the three strings)
Wikipedia

This TEI datatype is mapped to the XSD token datatype, and may therefore contain any string of characters. However, it is recommended that the value used conform to the particular flavour of regular expression syntax supported by XSD Schema.

Appendix A.4.11 teidata.pointer

teidata.pointer defines the range of attribute values used to provide a single URI, absolute or relative, pointing to some other resource, either within the current document or elsewhere.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef restriction="\S+" name="anyURI"/>
 </content>
-    
Declaration
-tei_teidata.pointer = xsd:anyURI { pattern = "\S+" }
Note

The range of syntactically valid values is defined by RFC 3986 Uniform Resource Identifier (URI): Generic Syntax. Note that the values themselves are encoded using RFC 3987 Internationalized Resource Identifiers (IRIs) mapping to URIs. For example, https://secure.wikimedia.org/wikipedia/en/wiki/% is encoded as https://secure.wikimedia.org/wikipedia/en/wiki/%25 while http://موقع.وزارة-الاتصالات.مصر/ is encoded as http://xn--4gbrim.xn----rmckbbajlc6dj7bxne2c.xn--wgbh1c/

Appendix A.4.12 teidata.prefix

teidata.prefix defines a range of values that may function as a URI scheme name.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.pointer = xsd:anyURI { pattern = "\S+" }
Note

The range of syntactically valid values is defined by RFC 3986 Uniform Resource Identifier (URI): Generic Syntax. Note that the values themselves are encoded using RFC 3987 Internationalized Resource Identifiers (IRIs) mapping to URIs. For example, https://secure.wikimedia.org/wikipedia/en/wiki/% is encoded as https://secure.wikimedia.org/wikipedia/en/wiki/%25 while http://موقع.وزارة-الاتصالات.مصر/ is encoded as http://xn--4gbrim.xn----rmckbbajlc6dj7bxne2c.xn--wgbh1c/

Appendix A.4.12 teidata.prefix

teidata.prefix defines a range of values that may function as a URI scheme name.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef name="token"
   restriction="[a-z][a-z0-9\+\.\-]*"/>
 </content>
-    
Declaration
-tei_teidata.prefix = token { pattern = "[a-z][a-z0-9\+\.\-]*" }
Note

This datatype is used to constrain a string of characters to one that can be used as a URI scheme name according to RFC 3986, section 3.1. Thus only the 26 lowercase letters a–z, the 10 digits 0–9, the plus sign, the period, and the hyphen are permitted, and the value must start with a letter.

Appendix A.4.13 teidata.probCert

teidata.probCert defines a range of attribute values which can be expressed either as a numeric probability or as a coded certainty value.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.prefix = token { pattern = "[a-z][a-z0-9\+\.\-]*" }
Note

This datatype is used to constrain a string of characters to one that can be used as a URI scheme name according to RFC 3986, section 3.1. Thus only the 26 lowercase letters a–z, the 10 digits 0–9, the plus sign, the period, and the hyphen are permitted, and the value must start with a letter.

Appendix A.4.13 teidata.probCert

teidata.probCert defines a range of attribute values which can be expressed either as a numeric probability or as a coded certainty value.
Moduletei — Formal specification
Used by
Content model
 <content>
  <alternate>
   <dataRef key="teidata.probability"/>
   <dataRef key="teidata.certainty"/>
  </alternate>
 </content>
-    
Declaration
-tei_teidata.probCert = teidata.probability | teidata.certainty

Appendix A.4.14 teidata.probability

teidata.probability defines the range of attribute values expressing a probability.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.probCert = teidata.probability | teidata.certainty

Appendix A.4.14 teidata.probability

teidata.probability defines the range of attribute values expressing a probability.
Moduletei — Formal specification
Used by
Content model
 <content>
  <dataRef name="double">
   <dataFacet name="minInclusive" value="0"/>
   <dataFacet name="maxInclusive" value="1"/>
  </dataRef>
 </content>
-    
Declaration
-tei_teidata.probability = xsd:double
Note

Probability is expressed as a real number between 0 and 1; 0 representing certainly false and 1 representing certainly true.

Appendix A.4.15 teidata.replacement

teidata.replacement defines attribute values which contain a replacement template.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.probability = xsd:double
Note

Probability is expressed as a real number between 0 and 1; 0 representing certainly false and 1 representing certainly true.

Appendix A.4.15 teidata.replacement

teidata.replacement defines attribute values which contain a replacement template.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <textNode/>
 </content>
-    
Declaration
-tei_teidata.replacement = text

Appendix A.4.16 teidata.temporal.iso

teidata.temporal.iso defines the range of attribute values expressing a temporal expression such as a date, a time, or a combination of them, that conform to the international standard Data elements and interchange formats – Information interchange – Representation of dates and times.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.replacement = text

Appendix A.4.16 teidata.temporal.iso

teidata.temporal.iso defines the range of attribute values expressing a temporal expression such as a date, a time, or a combination of them, that conform to the international standard Data elements and interchange formats – Information interchange – Representation of dates and times.
Moduletei — Formal specification
Used by
Content model
 <content>
  <alternate>
   <dataRef name="date"/>
@@ -4510,7 +4588,7 @@
    restriction="[0-9.,DHMPRSTWYZ/:+\-]+"/>
  </alternate>
 </content>
-    
Declaration
+    
Declaration
 tei_teidata.temporal.iso =
    xsd:date
  | xsd:gYear
@@ -4520,7 +4598,7 @@
  | xsd:gMonthDay
  | xsd:time
  | xsd:dateTime
- | token { pattern = "[0-9.,DHMPRSTWYZ/:+\-]+" }
Note

If it is likely that the value used is to be compared with another, then a time zone indicator should always be included, and only the dateTime representation should be used.

For all representations for which ISO 8601:2004 describes both a basic and an extended format, these Guidelines recommend use of the extended format.

Appendix A.4.17 teidata.temporal.w3c

teidata.temporal.w3c defines the range of attribute values expressing a temporal expression such as a date, a time, or a combination of them, that conform to the W3C XML Schema Part 2: Datatypes Second Edition specification.
Moduletei — Formal specification
Used by
Element:
Content model
+ | token { pattern = "[0-9.,DHMPRSTWYZ/:+\-]+" }
Note

If it is likely that the value used is to be compared with another, then a time zone indicator should always be included, and only the dateTime representation should be used.

For all representations for which ISO 8601:2004 describes both a basic and an extended format, these Guidelines recommend use of the extended format.

Appendix A.4.17 teidata.temporal.w3c

teidata.temporal.w3c defines the range of attribute values expressing a temporal expression such as a date, a time, or a combination of them, that conform to the W3C XML Schema Part 2: Datatypes Second Edition specification.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <alternate>
   <dataRef name="date"/>
@@ -4533,7 +4611,7 @@
   <dataRef name="dateTime"/>
  </alternate>
 </content>
-    
Declaration
+    
Declaration
 tei_teidata.temporal.w3c =
    xsd:date
  | xsd:gYear
@@ -4542,30 +4620,30 @@
  | xsd:gYearMonth
  | xsd:gMonthDay
  | xsd:time
- | xsd:dateTime
Note

If it is likely that the value used is to be compared with another, then a time zone indicator should always be included, and only the dateTime representation should be used.

Appendix A.4.18 teidata.text

teidata.text defines the range of attribute values used to express some kind of identifying string as a single sequence of Unicode characters possibly including whitespace.
Moduletei — Formal specification
Used by
Element:
Content model
+ | xsd:dateTime
Note

If it is likely that the value used is to be compared with another, then a time zone indicator should always be included, and only the dateTime representation should be used.

Appendix A.4.18 teidata.text

teidata.text defines the range of attribute values used to express some kind of identifying string as a single sequence of Unicode characters possibly including whitespace.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef name="string"/>
 </content>
-    
Declaration
-tei_teidata.text = string
Note

Attributes using this datatype must contain a single ‘token’ in which whitespace and other punctuation characters are permitted.

Appendix A.4.19 teidata.truthValue

teidata.truthValue defines the range of attribute values used to express a truth value.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.text = string
Note

Attributes using this datatype must contain a single ‘token’ in which whitespace and other punctuation characters are permitted.

Appendix A.4.19 teidata.truthValue

teidata.truthValue defines the range of attribute values used to express a truth value.
Moduletei — Formal specification
Used by
Content model
 <content>
  <dataRef name="boolean"/>
 </content>
-    
Declaration
-tei_teidata.truthValue = xsd:boolean
Note

The possible values of this datatype are 1 or true, or 0 or false.

This datatype applies only for cases where uncertainty is inappropriate; if the attribute concerned may have a value other than true or false, e.g. unknown, or inapplicable, it should have the extended version of this datatype: teidata.xTruthValue.

Appendix A.4.20 teidata.versionNumber

teidata.versionNumber defines the range of attribute values used for version numbers.
Moduletei — Formal specification
Used by
Element:
Content model
+    
Declaration
+tei_teidata.truthValue = xsd:boolean
Note

The possible values of this datatype are 1 or true, or 0 or false.

This datatype applies only for cases where uncertainty is inappropriate; if the attribute concerned may have a value other than true or false, e.g. unknown, or inapplicable, it should have the extended version of this datatype: teidata.xTruthValue.

Appendix A.4.20 teidata.versionNumber

teidata.versionNumber defines the range of attribute values used for version numbers.
Moduletei — Formal specification
Used by
Element:
Content model
 <content>
  <dataRef name="token"
   restriction="[\d]+[a-z]*[\d]*(\.[\d]+[a-z]*[\d]*){0,3}"/>
 </content>
-    
Declaration
+    
Declaration
 tei_teidata.versionNumber =
-   token { pattern = "[\d]+[a-z]*[\d]*(\.[\d]+[a-z]*[\d]*){0,3}" }

Appendix A.4.21 teidata.word

teidata.word defines the range of attribute values expressed as a single word or token.
Moduletei — Formal specification
Used by
teidata.enumeratedElement:
Content model
+   token { pattern = "[\d]+[a-z]*[\d]*(\.[\d]+[a-z]*[\d]*){0,3}" }

Appendix A.4.21 teidata.word

teidata.word defines the range of attribute values expressed as a single word or token.
Moduletei — Formal specification
Used by
teidata.enumeratedElement:
Content model
 <content>
  <dataRef name="token"
   restriction="[^\p{C}\p{Z}]+"/>
 </content>
-    
Declaration
-tei_teidata.word = token { pattern = "[^\p{C}\p{Z}]+" }
Note

Attributes using this datatype must contain a single ‘word’ which contains only letters, digits, punctuation characters, or symbols: thus it cannot include whitespace.

Appendix A.4.22 teidata.xTruthValue

teidata.xTruthValue (extended truth value) defines the range of attribute values used to express a truth value which may be unknown.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.word = token { pattern = "[^\p{C}\p{Z}]+" }
Note

Attributes using this datatype must contain a single ‘word’ which contains only letters, digits, punctuation characters, or symbols: thus it cannot include whitespace.

Appendix A.4.22 teidata.xTruthValue

teidata.xTruthValue (extended truth value) defines the range of attribute values used to express a truth value which may be unknown.
Moduletei — Formal specification
Used by
Content model
 <content>
  <alternate>
   <dataRef name="boolean"/>
@@ -4575,15 +4653,15 @@
   </valList>
  </alternate>
 </content>
-    
Declaration
-tei_teidata.xTruthValue = xsd:boolean | ( "unknown" | "inapplicable" )
Note

In cases where where uncertainty is inappropriate, use the datatype teidata.TruthValue.

Appendix A.4.23 teidata.xpath

teidata.xpath defines attribute values which contain an XPath expression.
Moduletei — Formal specification
Used by
Content model
+    
Declaration
+tei_teidata.xTruthValue = xsd:boolean | ( "unknown" | "inapplicable" )
Note

In cases where where uncertainty is inappropriate, use the datatype teidata.TruthValue.

Appendix A.4.23 teidata.xpath

teidata.xpath defines attribute values which contain an XPath expression.
Moduletei — Formal specification
Used by
Content model
 <content>
  <textNode/>
 </content>
-    
Declaration
-tei_teidata.xpath = text
Note

Any XPath expression using the syntax defined in 6.2..

When writing programs that evaluate XPath expressions, programmers should be mindful of the possibility of malicious code injection attacks. For further information about XPath injection attacks, see the article at OWASP.

Notes
1
Note that this is a illustrative example, i.e. a valid ParlaMint corpus would also need certain attributes to be defined on the illustrated elements. This holds for all the examples in this section.
2
Note that parliaments also have unaffiliated (or independent) MPs, that can either belong to a special ‘unaffiliated’ parliamentary group or don't belong to any parliamentary group. For the former, they are simply not affiliated to any parliamentary group. For the latter, an ‘unaffiliated’ parlimentaryGroup organisation must be created, and such MPs are affiliated with it as members.
3
The ideal situation is that the organisation somebody is affiliated with is specificed as a organisation, using the <org> element (cf. the Section on Organisations) but if this is not the case, using <orgName> directly in the <affiliation> is an alternative encoding.
4
Note that, in general, the utterance can also be split in the middle of a sentence, which brings with it problems for automatic linguistic processing, as, ideally, the parts should be first joined, and only then processed.
5
These are typically tagset developed and used for specific languages and can be found in the XPOS column of CoNLL-U files, which is the native format for UD treebanks.
6
Note that the example is rendered in three lines, however, the correct encoding in the corpus is actually in a single line, without any spaces between the elements, as otherwise the new line and indenting spaces are actually a part of the word ‘abyste’.
7
Because <name> and <phr> can give conflicting markup (i.e. crossing tags) the current script annotates phrases only where they are not related to names, i.e. not only conflicting markup, but also nestings of phr/name and name/phr are forbidden and such MWEs are not retained in the XML. Furthermore, due to a bug in the script, phrases adjecent to names are also not retained. We hope to introduce a better script and encoding in the future.
Tomaž Erjavec, tomaz.erjavec@ijs.si, Matyáš Kopp, kopp@ufal.mff.cuni.cz and Andrej Pančur, andrej.pancur@inz.si. Date: 2025-06-13
Notes
1
Note that this is a illustrative example, i.e. a valid ParlaMint corpus would also need certain attributes to be defined on the illustrated elements. This holds for all the examples in this chapter.
2
Note that parliaments also have unaffiliated (or independent) MPs, that can either belong to a special ‘unaffiliated’ parliamentary group or don't belong to any parliamentary group. For the former, they are simply not affiliated to any parliamentary group. For the latter, an ‘unaffiliated’ parlimentaryGroup organisation must be created, and such MPs are affiliated with it as members.
3
The ideal situation is that the organisation somebody is affiliated with is specificed as a organisation, using the <org> element (cf. the Section on Organisations) but if this is not the case, using <orgName> directly in the <affiliation> is an alternative encoding.
4
Note that, in general, the utterance can also be split in the middle of a sentence, which brings with it problems for automatic linguistic processing, as, ideally, the parts should be first joined, and only then processed.
5
These are typically tagset developed and used for specific languages and can be found in the XPOS column of CoNLL-U files, which is the native format for UD treebanks.
6
Note that the example is rendered in three lines, however, the correct encoding in the corpus is actually in a single line, without any spaces between the elements, as otherwise the new line and indenting spaces are actually a part of the word ‘abyste’.
7
Because <name> and <phr> can give conflicting markup (i.e. crossing tags) the current script annotates phrases only where they are not related to names, i.e. not only conflicting markup, but also nestings of phr/name and name/phr are forbidden and such MWEs are not retained in the XML. Furthermore, due to a bug in the script, phrases adjecent to names are also not retained. We hope to introduce a better script and encoding in the future.
Tomaž Erjavec, tomaz.erjavec@ijs.si, Matyáš Kopp, kopp@ufal.mff.cuni.cz and Andrej Pančur, andrej.pancur@inz.si. Date: 2025-09-30
\ No newline at end of file