Zomi language dataset for NLP and AI training — clean corpus, golden textbooks, normalizer
-
Updated
Jun 16, 2026 - Python
Zomi language dataset for NLP and AI training — clean corpus, golden textbooks, normalizer
A script for preprocessing a frequency list from the Norwegian Web as Corpus (NoWaC) in R
Add a description, image, and links to the language-corpus topic page so that developers can more easily learn about it.
To associate your repository with the language-corpus topic, visit your repo's landing page and select "manage topics."