众包与增强的数据质量保证:一个有效的方法,以减轻资源稀缺的挑战在培训医疗保健大型语言模型的训练
Prosanta Barai1, Gondy Leroy1, Prakash Bisht1
1The University of Arizona, Tucson 85721, U.S.A.
概括
本研究介绍了一个众包框架与质量控制,以改善大语言模型 (LLM) 的医疗保健数据. 实时质量检查显著提高了数据质量,有助于更好地预测自闭症症状.
科学领域:
- 人工智能的人工智能
- 医疗保健信息学 医疗保健信息学
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 在医疗保健方面具有前景,但需要高质量的标记数据.
- 数据获取是具有挑战性和昂贵的,特别是在资源较低的医疗保健环境中.
- 现有的方法在专门的AI应用中与数据质量作斗争.
研究的目的:
- 开发和评估一个众包框架,对医疗保健数据进行综合质量控制.
- 评估增强数据质量的LLM绩效对自闭症症状预测的影响.
- 解决资源有限的医疗保健领域的数据稀缺和质量问题.
主要方法:
- 实施众包框架,对数据收集前,实时和后进行质量控制.
- 使用定量指标对数据质量改进的评估.
- 一个针对医疗保健的LLM (Bio-BERT) 在众包数据上的微调,用于自闭症症状预测.
- 与基线模型对比LLM绩效的比较.
主要成果:
- 实时质量控制显示数据质量比质量前控制措施有19%的改善.
- 微调生物BERT与众包数据导致回忆力增加,但与基线相比精度降低.
- 众包方法显示了在数据有限的场景中提高LLM性能的潜力.
结论:
- 通过强有力的质量控制增强的众包是获得高质量的医疗保健数据的可行策略.
- 优化数据采集可以提高LLM在医疗保健中的有效性,特别是对于诸如症状预测等任务.
- 调查结果为开发更有效,更节省资源的医疗保健人工智能解决方案提供了见解.
相关概念视频
Improving Translational Accuracy
2.6K
2.6K
Data Validation
5.0K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.0K
Health Literacy
4.0K
Health literacy is an individual's or a community's capacity to comprehend, receive, read, and use relevant healthcare information and services. The World Health Organization (WHO, 2018) defines health literacy as the cognitive and social skills that determine the ability of individuals to gain access to, understand, and use information in ways that promote and maintain good health. As a result, the WHO helps individuals manage long-term health concerns, participate in preventative...
4.0K


