大規模言語モデルを用いた定量的研究:テキストデータの品質が特徴量表現と機械学習モデルに与える影響の評価
Tabinda Sarwar1, Antonio José Jimeno Yepes1, Lawrence Cavedon1
1Royal Melbourne Institute of Technology University, Melbourne, Australia.
Journal of medical Internet research
|December 30, 2025
まとめ
機械学習モデルは、エラー率が10%未満のヘルスケアデータセットで良好に機能するが、エラー率が高くなるとパフォーマンスが大幅に低下する。信頼性の高いヘルスケアにおける機械学習のためには、データ品質の評価と修正が不可欠である。
科学分野:
- ヘルスインフォマティクス; 機械学習; 自然言語処理
背景:
- 実世界のデータ収集では、データ品質が損なわれることが多く、機械学習(ML)モデルのパフォーマンスに影響を与える。進捗記録などのヘルスケアにおけるテキストデータは、不正確なML予測が生命を脅かす可能性があるため、高い精度が必要とされる。テキストデータの品質の影響を評価することは、信頼性の高いヘルスケアMLアプリケーションに不可欠である。
研究 の 目的:
- テキストデータセットの品質を定量化し、エラーがMLモデルに与える影響を評価する。特徴量表現とMLモデルのエラーに対する耐性を決定する。ヘルスケアMLのためのデータ品質改善へのリソース投資の正当性を評価する。
主な方法:
- テキストデータセットの品質評価のためのトークンレベルのエラー率メトリックを開発した。Mixtral大規模言語モデル(LLM)を使用して、データセットのエラーを定量化および修正した。MIMIC-III(高品質)およびオーストラリア高齢者介護施設(AACHs、低品質)データセットを分析し、MIMIC-IIIにエラーを導入し、AACHsを修正した。受信者操作特性曲線下面積(AUC)を使用して、特徴量表現とMLモデルを評価した。
主要な成果:
- Mixtralは進捗記録の63%でエラーを検出し、17%で単一トークンの誤分類が見られた。特徴量表現のパフォーマンスは10%未満のエラー率では許容範囲であったが、この閾値を超えると著しく低下した。エラー率8%のAACHデータセットでは、パフォーマンスの大きな低下は見られなかった。TF-IDFは埋め込み特徴量よりも優れていた。MLモデルの有効性はタスクによって異なった。
結論:
- MLモデルはデータ品質に敏感であり、エラー率が10%以上になるとパフォーマンスが著しく低下する。ヘルスケアでのML実装前にデータセット品質評価が不可欠である。MLモデルの信頼性と有効性を確保するため、エラー率の高いデータセットには修正措置が不可欠である。
関連する概念動画
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Survival Tree
369
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
369
Language and Cognition
681
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
681
Stereotype Content Model
15.3K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.3K
Regression Analysis
7.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.8K


