Related Experiment Video
Updated: Sep 15, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
691
Menstrual Health Education Using a Specialized Large Language Model in India: Development and Evaluation Study of
Prottay Kumar Adhikary1, Isha Motiyani1, Gayatri Oke1
1Department of Electrical Engineering, Indian Institute of Technology Delhi, Room: 3B-7 (Block III 3rd Floor), Hauz Khas, New Delhi, 110016, India, 91 26591076 ext 011.
Journal of Medical Internet Research
|July 16, 2025
Summary
MenstLLaMA, a specialized AI model, significantly improves menstrual health education in India by providing accurate, culturally sensitive information. It outperforms general large language models (LLMs) in user satisfaction and expert evaluations.
Area of Science:
- Artificial Intelligence in Healthcare
- Digital Health Education
- Natural Language Processing
Background:
- Menstrual health education (MHE) in low- and middle-income countries, including India, faces challenges like poverty, stigma, and inequality.
- General-purpose large language models (LLMs) often lack accuracy and cultural sensitivity for MHE.
- A specialized LLM, MenstLLaMA, was developed to address these limitations for the Indian context.
Purpose of the Study:
- To develop and evaluate MenstLLaMA, a specialized LLM for accurate and culturally sensitive MHE.
- To compare MenstLLaMA's effectiveness against existing general-purpose LLMs.
Main Methods:
- Curated MENST dataset (23,820 Q&A pairs) with demographic and cultural metadata.
- Fine-tuned Meta-LLaMA-3-8B-Instruct using parameter-efficient techniques.
- Multi-layered evaluation: NLP metrics (BLEU, BERTScore), expert review, medical practitioner interaction (ISHA chatbot), and user study (N=200).
Main Results:
- MenstLLaMA achieved top scores in BLEU (0.059) and BERTScore (0.911), surpassing GPT-4o and Claude-3.
- Clinical experts preferred MenstLLaMA's culturally sensitive responses.
- User satisfaction ratings were high: understandability (4.7/5), relevance (4.3/5), correctness (4.1/5), and context sensitivity (3.9/5).
Conclusions:
- MenstLLaMA demonstrates superior accuracy, empathy, and user satisfaction for MHE compared to general LLMs.
- It offers a scalable solution to bridge MHE gaps in diverse populations.
- Future work includes expanding demographic representation and integrating multimodal interactions.

