Related Experiment Video
Updated: Jan 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Application of Large Language Models in Data Analysis and Medical Education for Assisted Reproductive Technology:
Noriyuki Okuyama1, Mika Ishii1, Yuriko Fukuoka1
1Kyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.
JMIR Formative Research
|October 1, 2025
Summary
Advanced AI models show promise in analyzing reproductive medicine data and answering infertility questions, significantly faster than human experts. However, current models struggle with accurate medical image diagnosis.
Area of Science:
- Artificial Intelligence in Medicine
- Reproductive Medicine Data Analysis
- Medical Knowledge Assessment
Background:
- Large language models (LLMs) demonstrate high performance in medical exams, but their application in specific domains like reproductive medicine and practical data analysis remains underexplored.
- Comparative analyses of different LLMs for specialized medical tasks are lacking.
- This study addresses the gap in understanding LLM capabilities for reproductive medicine data analysis and knowledge assessment.
Purpose of the Study:
- To evaluate the data analysis and reproductive medicine knowledge capabilities of advanced AI models.
- To compare the performance of GPT-4, GPT-4o, Claude 3.5 Sonnet, and Gemini Pro 1.5 in analyzing template-based reproductive data and answering expert-level questions.
- To assess the potential of AI as a tool for data-intensive tasks and decision support in reproductive medicine.
Main Methods:
- Four AI models were tested on their ability to perform pregnancy rate analysis and graph rendering using blastocyst grading data.
- Models were assessed on their knowledge of infertility treatment via 10 expert-developed examination questions.
- Performance was evaluated based on output accuracy, processing time, and diagnostic capabilities, with all procedures repeated 10 times per model.
Main Results:
- GPT-4o excelled in data analysis, achieving Grade A in 90% of trials with an average processing time of 26.8 seconds, significantly faster than embryologists.
- GPT-4o, Claude 3.5 Sonnet, and Gemini Pro 1.5 achieved perfect scores on multiple-choice knowledge questions, while GPT-4 had a 60% success rate.
- No AI model could reliably diagnose chromosomal abnormalities from karyotype images, with Claude and Gemini achieving the highest accuracy at 70%.
Conclusions:
- AI models can rapidly process reproductive medicine data, offering potential as educational tools or decision support systems.
- Despite strong performance in data analysis and knowledge recall, current AI models lack the capability for accurate medical image interpretation and diagnosis.
- Further development is needed to enhance AI's diagnostic accuracy in medical imaging for clinical application.

