Corpus of Ukrainian Narratives with Audio data (CUNA) / Корпус українських наративів з аудіоданими (КУНА) The Corpus of Ukrainian Narratives with Audio data (CUNA) comprises 40 in-depth semi-structured oral interviews with speakers from Ukraine documenting their discursive construction of Europe, Europeanness and Ukraine’s Europeanisation in the context of identity formation amid significant events in Ukraine’s recent history. These include Ukraine’s declaration of independence from the Soviet Union, civil resistance movements, Russia’s aggression beginning with the 2014 annexation of Crimea and armed combat in Donbas and continuing with the full-scale invasion launched in 2022 as well as Ukraine’s ongoing efforts toward accession to the European Union. In addition to political and historical themes, the interviews convey biographical, social and cultural memories related to professional life, volunteering, displacement, migration, language practices, specific geographical places in Ukraine and abroad, and a sense of both regional and national belonging. The interviews were conducted between June and August 2024 with women aged 40 and above. At the time of the interviews, 20 participants resided in Slovenia and 20 in Ukraine. The language of the interviews was predominantly Ukrainian, although the participants episodically code-switched to other languages including Russian, Slovenian, English and Crimean Tatar. All participants have consented in writing to the storage and publication of the recordings and their transcripts under a CC BY 4.0 License. The recordings were automatically transcribed using the open-source speech recognition system Whisper. The resulting transcripts were manually aligned with audio segments in the open-source editor Audacity, then manually edited to ensure the accuracy of the speech recognition and conformity with the writing norms of Ukrainian and the other languages present in the data. Each transcript was annotated for both speaker-related and speech-related data. Speaker metadata was annotated manually and includes: speaker type (guest, host), birth cohort (in 5-year intervals), sex (female), education level (higher, vocational, complete secondary), country of residence at the time of the interview (Slovenia, Ukraine), place of birth (Ukrainian SSR or other), region of longest-term residence in Ukraine (Centre, East, West, South including Crimea), and self-identified mother tongue and ethnicity. Speech data includes automatic linguistic annotation with UDPipe2 (tokenisation, PoS tagging, morphological annotation, lemmatisation) as well as manual annotation for word-level code-switching and sentence-level topics based on questionnaire prompts. Personal names of private individuals mentioned in the transcripts were manually anonymised.
DARIAH-SI/CUNA
Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|