Breaking the Language Barrier for the Next Billion Users.
The Multilingual Nigeria NLU Engine is an AI project dedicated to making advanced Large Language Models (LLMs) accessible to native speakers of Nigeria's major indigenous languages. Current AI models often struggle with the nuances, dialects, and low-resource nature of West African languages. This project aims to solve that.
Nigeria has over 500 indigenous languages, yet most AI tools are English-centric. This creates a digital divide for millions of speakers of Hausa, Yoruba, Igbo, and Fulfulde.
- Scrape and clean localized datasets from Wikipedia, Common Crawl, and local news outlets.
- Develop a "Native Speaker" verification pipeline for manual annotation.
- Curate a specialized corpus for Nigerian Pidgin football and market slang.
- Fine-tune Afro-XLM-R or Llama-3 variants on curated Nigerian corpora.
- Implement a hybrid CNN-Transformer architecture for robust classification in low-resource settings.
- Launch a lightweight API for integration into healthcare (MedRoute) and education tools.
- Build a "Trans-Linguist" chatbot demo that switches between languages fluently.
- Frameworks: PyTorch, Hugging Face Transformers.
- Base Models: Afro-XLM-R, mBART.
- Deployment: FastAPI, Docker.
Umar Yusuf Lamiya - umaryusuf9191@gmail.com