Natural Language Processing

Development and Evaluation of Large Language Models for Under-resourced Ethiopian Languages

This research proposes developing and evaluating a multilingual Large Language Model (LLM) for under-resourced Ethiopian languages. It focuses on building quality multilingual datasets, improving model training through techniques such as transfer learning and fine-tuning, and evaluating performance across tasks including translation, question answering, summarization, text generation, and sentiment analysis. The study also addresses fairness, hallucination, cultural accuracy, code-switching, dialect variation, and limited data. The expected result is an open, reproducible framework with datasets, benchmarks, methodologies, and a prototype Ethiopian-language LLM that can support inclusive AI, education, digital language preservation, and public services.

100
Progress
15
Researchers

Roles needed

  • Data Annotation
  • Data Preprocessing

Milestones

  1. Initial Manuscript Draft Date to be set · pending

Newsletter

Keep up with the group.

Calls for collaborators, papers and events — a short note when there is something worth saying.

The address is not shared, and one click leaves the list.