Related Experiment Video
Updated: Aug 1, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
643
Deep learning for Arabic healthcare: MedicalBot
Mohammed Abdelhay1, Ammar Mohammed1, Hesham A Hefny1
1Department of Computer Science, Faculty of Graduate Studies for Statistical Research, Cairo University, Giza, Egypt.
Summary
This study introduces MAQA, the largest Arabic healthcare question-answering dataset, to improve medical bots. Transformer models demonstrated superior performance, enhancing AI-driven healthcare accessibility in Arabic.
Area of Science:
- Natural Language Processing
- Artificial Intelligence in Healthcare
- Medical Informatics
Background:
- Remote and automated healthcare consultations have surged post-COVID-19.
- Medical bots offer 24/7 access, reduced wait times, and cost savings.
- Effective medical bots require high-quality, domain-specific training data.
Purpose of the Study:
- To address the lack of a comprehensive Arabic medical question-answering dataset.
- To introduce the MAQA dataset, the largest of its kind for Arabic healthcare.
- To benchmark deep learning models for Arabic medical chatbots.
Main Methods:
- Creation of the MAQA dataset: over 430,000 Arabic healthcare Q&A pairs across 20 specializations.
- Implementation and comparison of three deep learning models: LSTM, Bi-LSTM, and Transformers.
- Evaluation of model performance using cosine similarity and BLEU score.
Main Results:
- The Transformer model significantly outperformed LSTM and Bi-LSTM.
- The Transformer model achieved an average cosine similarity of 80.81%.
- The Transformer model obtained a BLeU score of 58%.
Conclusions:
- The MAQA dataset is a valuable resource for developing Arabic medical bots.
- Transformer models show strong potential for advancing Arabic natural language understanding in healthcare.
- This work facilitates improved AI-driven healthcare solutions for Arabic-speaking populations.

