Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

9 Commits
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ANETAC: Arabic Named Entity Transliteration and Classification Dataset

Description

This is a dataset containing 79,924 named entites in Arabic and their transliteration in English. The dataset was published as part of the paper Arabic Machine Transliteration using an Attention-based Encoder-decoder Model

The dataset is split into train, test, and development sets:

Arabic-English named entites transliteration Dataset

Examples

Examples of the instances present in the dataset are provided in the below Table:

Arabic-English named entites transliteration examples

Statistics

The statistics the dataset are provided in the below Table:

Arabic-English named entites transliteration statistics

Citations

If you want to use the dataset please cite the following arXiv paper:

@article{ameur2019anetac,
  title={Anetac: Arabic named entity transliteration and classification dataset},
  author={Ameur, Mohamed Seghir Hadj and Meziane, Farid and Guessoum, Ahmed},
  journal={arXiv preprint arXiv:1907.03110},
  year={2019}
}

Contacts:

For all questions please contact mohamedhadjameur@gmail.com