Related Experiment Video
Updated: Aug 10, 2026

A Study on an Intelligent Diagnosis and Treatment Assistant System for Acupuncture in Diminished Ovarian Reserve Based on a Knowledge Graph
Published on: May 29, 2026
A parallel English-Akan maternal health dataset to support machine translation and text-to-speech systems
Isaac Wiafe1, Akon Obu Ekpezu1,2, Fiifi Baffoe Payin Winful1
1Department of Computer Science, University of Ghana, Ghana.
Abstract:
Audio and text datasets are essential for developing machine translation, text-to-speech systems, and speech-enabled systems. However, domain-specific datasets for low-resource African languages remain limited, particularly in the healthcare domain. This study addresses this gap by introducing a parallel English-Akan maternal health dataset to support machine translation and text-to-speech systems. The dataset consists of 3487 unique English maternal health phrases (question-and-answer) generated from verified digital health sources and medically reviewed for Ghanaian contextual relevance. These phrases were translated into Akan by four linguistic experts, producing 12,000 Akan transcriptions and 12,000 corresponding Akan audio recordings, totalling 32.313 h. The dataset covers prenatal and postnatal maternal health themes that are aligned with the WHO-recommended domains. The audio recordings were captured in soundproof vocal booths and processed into .wav format. This dataset provides a valuable resource for advancing Akan health communication technologies, maternal health chatbots, and low-resource language AI research.