This is a dataset containing 79,924 named entites in Arabic and their transliteration in English. The dataset was published as part of the paper Arabic Machine Transliteration using an Attention-based Encoder-decoder Model
The dataset is split into train, test, and development sets:
Examples of the instances present in the dataset are provided in the below Table:
The statistics the dataset are provided in the below Table:
If you want to use the dataset please cite the following arXiv paper:
@article{ameur2019anetac,
title={Anetac: Arabic named entity transliteration and classification dataset},
author={Ameur, Mohamed Seghir Hadj and Meziane, Farid and Guessoum, Ahmed},
journal={arXiv preprint arXiv:1907.03110},
year={2019}
}
For all questions please contact mohamedhadjameur@gmail.com