This collection of files contains datasets for 7 languages (Russian, Basque, Turkish, Spanish, Czech, English, and Polish) corresponding to 3 different SES formats, namely, ses-udpipe, ses-ixapipes and ses-morpheus. These files are ready to use and are suitable to fine-tune such models as XLM-RoBERTa-large in a token classification task in to perform lemmatization.
For more information and if you use our data please refer to the paper:
@inproceedings{toporkov-agerri-2024-evaluating,
title = "Evaluating Shortest Edit Script Methods for Contextual Lemmatization",
author = "Toporkov, Olia and
Agerri, Rodrigo",
editor = "Calzolari, Nicoletta and
Kan, Min-Yen and
Hoste, Veronique and
Lenci, Alessandro and
Sakti, Sakriani and
Xue, Nianwen",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "<https://aclanthology.org/2024.lrec-main.572/>",
pages = "6451--6463",
}