大型语言模型在孕产妇健康方面的输出质量评估
Henrique A Lima1, Pedro H F S Trocoli-Couto1, Zorays Moazzam2
1Federal University of Minas Gerais Faculty of Medicine, Belo Horizonte, Brazil.
Scientific reports
|July 2, 2025
概括
大型语言模型 (LLM) 显示出在低收入国家提高孕产妇健康素养的潜力. GPT-4和GPT-3.5在提供清晰,高质量的健康信息方面表现出卓越的表现,尽管内容的完整性和可读性需要进一步开发.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 全球健康 全球健康
背景情况:
- 在低收入和中等收入国家 (LMICs) 优化医疗保健准入对于提高健康素养至关重要.
- 大型语言模型 (LLM) 为提供准确,可靠和文化相关的健康信息提供了一个潜在的途径.
- 孕产妇健康仍然是一个关键领域,需要加强信息可访问性.
研究的目的:
- 评估各种大型语言模型 (LLM) 产生的孕产妇健康信息的质量.
- 在技术和非技术上下文中评估LLM成果的准确性,清晰性,可读性和文化相关性.
- 为了比较不同LLM的性能,包括GPT-4,GPT-3.5,一个定制的GPT-3.5,以及Meditron-70b.
主要方法:
- 采用了混合方法,横截面调查方法.
- 对选定的LLM (GPT-4,GPT-3.5,定制GPT-3.5,Meditron-70b) 提出了孕产妇健康问题.
- 来自巴西,美国和巴基斯坦的专家评估了本国语言的LLM答案,评估了信息质量,清晰度,可读性和充分性.
主要成果:
- 与定制的GPT-3.5和Meditron-70b相比,GPT-4和GPT-3.5在整体质量,清晰度和内容方面获得了持续更高的分数.
- 响应质量因语言而异,不完整的内容是最常见的限制.
- 可读性分析表明需要更高的教育水平来理解,并且在回复中观察到性别偏见.
结论:
- GPT-4和GPT-3.5显示出在全球范围内提高获得高质量的孕产妇健康信息的巨大潜力.
- 翻译工具和上下文化方面的改进是必要的,以优化LLM的性能,特别是对于非英语内容.
- 解决内容完整性,可读性和性别偏见等局限性对于在医疗保健中负责任地部署人工智能至关重要.
相关概念视频
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Language Development
460
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
460
Nursing Evaluation
3.6K
The evaluation stage signals the end of the nursing process. The nurse gathers evaluative data to assess whether or not the patient has attained the expected results. Whereas the nurse collects data in the nursing assessment to identify the patient's health concerns, the evaluation stage data determines if the indicated health issues are resolved. Evaluative data collection includes two sections: the data acquired to evaluate patient outcomes and the time criteria for data collection.
3.6K
Nursing Assessment
8.2K
The two sources for collecting information are primary and secondary. After gathering information, interpretation and validation help to complete the data. The purpose of assessment is to establish data with the initial information, to interpret data about the patient's perceived needs and health problems, and to respond to these problems identified.
The nurse collects all aspects of the patient's health in the initial assessment, establishing priorities for ongoing focused assessments...
The nurse collects all aspects of the patient's health in the initial assessment, establishing priorities for ongoing focused assessments...
8.2K
Data Validation
5.4K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.4K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K


