German-English code-switching
We describe a corpus of German-English code-switching in social media interactions. We focus on some challenges in annotating CS, especially due to words whose language ID cannot be easily determined. We introduce a novel schema for such word-level annotation, with which we manually annotated a subset of the corpus. We then trained classifiers to predict and identify switches, and applied them to the remainder of the corpus. Thereby, we created a large-scale corpus of German-English mixed utterances with precise indications of CS points.
If you use this dataset, please cite: @inproceedings{osmelak-wintner-2023-denglisch, title = "The Denglisch Corpus of {G}erman-{E}nglish Code-Switching", author = "Osmelak, Doreen and Wintner, Shuly", booktitle = "Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP", month = may, year = "2023", address = "Dubrovnik, Croatia", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.sigtyp-1.5", pages = "42--51", }