Skip to content

Latest commit

 

History

37 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ses-lemma

This collection of files contains datasets for 7 languages (Russian, Basque, Turkish, Spanish, Czech, English, and Polish) corresponding to 3 different SES formats, namely, ses-udpipe, ses-ixapipes and ses-morpheus. These files are ready to use and are suitable to fine-tune such models as XLM-RoBERTa-large in a token classification task in to perform lemmatization.

For more information and if you use our data please refer to the paper:

@inproceedings{toporkov-agerri-2024-evaluating,
    title = "Evaluating Shortest Edit Script Methods for Contextual Lemmatization",
    author = "Toporkov, Olia  and
      Agerri, Rodrigo",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "<https://aclanthology.org/2024.lrec-main.572/>",
    pages = "6451--6463",
}

About

Evaluating Shortest Edit Script Methods for Contextual Lemmatization

Resources

Stars

0 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors

Languages