Skip to content

Latest commit

Β 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 

Repository files navigation

Multilingual Nigeria NLU Engine 🌍

Breaking the Language Barrier for the Next Billion Users.

The Multilingual Nigeria NLU Engine is an AI project dedicated to making advanced Large Language Models (LLMs) accessible to native speakers of Nigeria's major indigenous languages. Current AI models often struggle with the nuances, dialects, and low-resource nature of West African languages. This project aims to solve that.

🎯 The Problem

Nigeria has over 500 indigenous languages, yet most AI tools are English-centric. This creates a digital divide for millions of speakers of Hausa, Yoruba, Igbo, and Fulfulde.

πŸ› οΈ Technical Roadmap

Phase 1: Data Collection & Curation (Current)

  • Scrape and clean localized datasets from Wikipedia, Common Crawl, and local news outlets.
  • Develop a "Native Speaker" verification pipeline for manual annotation.
  • Curate a specialized corpus for Nigerian Pidgin football and market slang.

Phase 2: Model Architecture & Fine-tuning

  • Fine-tune Afro-XLM-R or Llama-3 variants on curated Nigerian corpora.
  • Implement a hybrid CNN-Transformer architecture for robust classification in low-resource settings.

Phase 3: Deployment & API

  • Launch a lightweight API for integration into healthcare (MedRoute) and education tools.
  • Build a "Trans-Linguist" chatbot demo that switches between languages fluently.

πŸ§ͺ Tech Stack

  • Frameworks: PyTorch, Hugging Face Transformers.
  • Base Models: Afro-XLM-R, mBART.
  • Deployment: FastAPI, Docker.

πŸ“§ Contact

Umar Yusuf Lamiya - umaryusuf9191@gmail.com

About

A specialized NLU engine designed to bridge the AI language gap for Hausa, Yoruba, Igbo, Fulfulde, and Nigerian Pidgin.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages