评估文本数据质量对特征表示和机器学习模型的影响:使用大型语言模型进行定量研究
Tabinda Sarwar1, Antonio José Jimeno Yepes1, Lawrence Cavedon1
1Royal Melbourne Institute of Technology University, Melbourne, Australia.
Journal of medical Internet research
|December 30, 2025
概括
机器学习模型在不到10%的错误率的健康数据集上表现良好,但在错误率较高的情况下,性能显著下降. 评估和纠正数据质量对于医疗保健中可靠的机器学习至关重要.
科学领域:
- 医疗信息学 医疗信息学
- 机器学习 机器学习
- 自然语言处理自然语言处理.
背景情况:
- 现实世界的数据收集通常会损害数据质量,影响机器学习 (ML) 模型的性能.
- 医疗保健中的文本数据,例如进度说明,由于错误的ML预测可能会导致危及生命的后果,因此需要高度准确.
- 评估文本数据质量的影响对于可靠的医疗保健ML应用程序至关重要.
研究的目的:
- 量化文本数据集质量,评估错误对ML模型的影响.
- 为了确定特征表示和ML模型对数据错误的容忍度.
- 评估投资资源改善医疗保健数据质量的合理性.
主要方法:
- 开发了用于文本数据集质量评估的代币级错误率指标.
- 利用Mixtral大型语言模型 (LLM) 来量化和纠正数据集中的错误.
- 分析了MIMIC-III (高质量) 和澳大利亚老年护理院 (AACHs;低质量) 数据集,将错误引入MIMIC-III并纠正AACHs.
- 评估的特征表示和ML模型使用接收器操作曲线 (AUC) 下的面积.
主要成果:
- 在63%的进度记录中,Mixtral检测到错误,其中17%是单一代币错误分类.
- 特征表示性能容忍的错误率低于10%,但在这个值以上显著下降.
- AACH数据集的错误率为8%,没有显示出重大性能下降.
- 术语频率反转的文档频率 (TF-IDF) 优于嵌入功能;ML模型的有效性因任务而异.
结论:
- ML模型对数据质量很敏感,在错误率为10%或更高的情况下,性能会显著降低.
- 在医疗保健中实施ML之前,数据集质量评估至关重要.
- 对于具有高错误率的数据集,纠正措施至关重要,以确保ML模型的可靠性和有效性.
相关概念视频
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Survival Tree
369
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
369
Language and Cognition
681
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
681
Stereotype Content Model
15.3K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.3K
Regression Analysis
7.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.8K


