Related Experiment Video
Updated: Jun 15, 2025

A Novel Method for Involving Women of Color at High Risk for Preterm Birth in Research Priority Setting
Published on: January 12, 2018
ChatGPT-3.5 Versus Google Bard: Which Large Language Model Responds Best to Commonly Asked Pregnancy Questions?
Keren Khromchenko1, Sameeha Shaikh2, Meghana Singh2
1Obstetrics and Gynecology, Hackensack Meridian Jersey Shore University Medical Center, Neptune, USA.
Abstract:
Large language models (LLM) have been widely used to provide information in many fields, including obstetrics and gynecology. Which model performs best in providing answers to commonly asked pregnancy questions is unknown. A qualitative analysis of Chat Generative Pre-Training Transformer Version 3.5 (ChatGPT-3.5) (OpenAI, Inc., San Francisco, California, United States) and Bard, recently renamed Google Gemini (Google LLC, Mountain View, California, United States), was performed in August of 2023. Each LLM was queried on 12 commonly asked pregnancy questions and asked for their references. Review and grading of the responses and references for both LLMs were performed by the co-authors individually and then as a group to formulate a consensus. Query responses were graded as "acceptable" or "not acceptable" based on correctness and completeness in comparison to American College of Obstetricians and Gynecologists (ACOG) publications, PubMed-indexed evidence, and clinical experience. References were classified as "verified," "broken," "irrelevant," "non-existent," and "no references." Grades of "acceptable" were given to 58% of ChatGPT-3.5 responses (seven out of 12) and 83% of Bard responses (10 out of 12). In regard to references, ChatGPT-3.5 had reference issues in 100% of its references, and Bard had discrepancies in 8% of its references (one out of 12). When comparing ChatGPT-3.5 responses between May 2023 and August 2023, a change in "acceptable" responses was noted: 50% versus 58%, respectively. Bard answered more questions correctly than ChatGPT-3.5 when queried on a small sample of commonly asked pregnancy questions. ChatGPT-3.5 performed poorly in terms of reference verification. The overall performance of ChatGPT-3.5 remained stable over time, with approximately one-half of responses being "acceptable" in both May and August of 2023. Both LLMs need further evaluation and vetting before being accepted as accurate and reliable sources of information for pregnant women.
More Related Videos
06:39Using a Murine Model of Psychosocial Stress in Pregnancy as a Translationally Relevant Paradigm for Psychiatric Disorders in Mothers and Infants
Published on: June 13, 2021
05:33Author Spotlight: Alleviating Nausea and Vomiting in Pregnancy with Safe and Effective Auricular Acupuncture
Published on: August 4, 2023
Related Concept Videos
Teratogenicity
Gonadal and Placental Hormones
In males, testosterone is the primary gonadal androgen. It plays a central role in the maturation of male reproductive organs — the penis and testes. Additionally, testosterone is instrumental in the development of secondary sexual characteristics — a deep voice as well as facial and pubic hair...
Diabetes Mellitus: Type 2 and Gestational
Fetal Circulation
Two umbilical arteries transport blood from the fetus to the placenta. At the placenta, the blood absorbs oxygen and nutrients while simultaneously eliminating waste products. This oxygen-enriched and nutrient-rich blood then returns to the fetus through one...
Fertilization
Hormones Regulating Blood Glucose
In addition to accelerating glucose uptake and utilization, insulin has...