Natural Language Processing
Development and Evaluation of Large Language Models for Under-resourced Ethiopian Languages
This research proposes developing and evaluating a multilingual Large Language Model (LLM) for under-resourced Ethiopian languages. It focuses on building quality multilingual datasets, improving model training through techniques such as transfer learning and fine-tuning, and evaluating performance across tasks including translation, question answering, summarization, text generation, and sentiment analysis. The study also addresses fairness, hallucination, cultural accuracy, code-switching, dialect variation, and limited data. The expected result is an open, reproducible framework with datasets, benchmarks, methodologies, and a prototype Ethiopian-language LLM that can support inclusive AI, education, digital language preservation, and public services.
- 100
- Progress
- 15
- Researchers
Roles needed
- Data Annotation
- Data Preprocessing
Milestones
- Initial Manuscript Draft Date to be set · pending