大規模言語モデルを用いた合成データとオントロジーによる臨床情報抽出の促進
Yan Hu1, Huan He2, Qingyu Chen2
1Mcwilliam School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, Texas, USA.
AMIA ... Annual Symposium proceedings. AMIA Symposium
|February 23, 2026
まとめ
大規模言語モデルは、固有表現抽出を改善するために合成臨床データを生成できる。自己検証と意味マッピングはデータユーティリティを強化し、人間対合成データの1:1の比率がパフォーマンスを最適化する。
科学分野:
- 自然言語処理
- ヘルスケアにおける人工知能
背景:
- 電子カルテには、膨大な非構造化臨床テキストが含まれています。
- 情報抽出システムの開発は重要ですが、アノテーション付きデータが不足しているため、制限されています。
研究 の 目的:
- 固有表現抽出のための合成臨床データを生成するための大規模言語モデルの探求。
- モデルのパフォーマンスと一般化可能性に対する合成データの影響の評価。
主な方法:
- SNOMED-CTの意味マッピングによる自己検証合成データ生成を用いた新しいフレームワーク。
- データ作成のためのGPT-4o-miniとファインチューニングのためのLLaMA-3-8Bの活用。
- 合成データ品質を向上させるための反復検証と異常検出。
主要な成果:
- 自己検証と意味マッピングは、合成データの有用性を大幅に向上させました。
- 人間アノテーションデータと合成データの比率が1:1の場合に最適なパフォーマンス向上が得られました。
- 4つの多様な臨床データセット全体で、モデルの一般化可能性の向上が観察されました。
結論:
- 合成データ生成は、臨床NLPアノテーションの課題に対するスケーラブルなソリューションです。
- モデルパフォーマンスの向上には、人間データと合成データのバランスが鍵となります。
- 提案されたフレームワークは、臨床情報抽出機能を前進させます。
さらに関連する動画
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.7K
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
16.6K
関連する概念動画
Language Development
990
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
990
Language and Cognition
865
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
865
Natural and Artificial Concepts
601
In psychology, concepts can be divided into two categories: natural and artificial. Natural concepts are formed through direct or indirect experiences. For example, consider the concept of snow. If you live in a place with regular snowfall, such as Essex Junction, Vermont, you know snow through direct experiences. You’ve seen it fall, touched it, shoveled it, and played in it. You recognize its texture, appearance, and even its smell. In contrast, if you live on an island like Saint...
601
Synthetic Biology
5.7K
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
5.7K
Improving Translational Accuracy
3.7K
3.7K
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
