全面召回:通过人工智能驱动的元数据标准化来提高数据的公平性
Sowmya S Sundaram1, Rafael S Gonçalves1, Mark A Musen1
1Stanford Center for Biomedical Informatics Research, Stanford University, Stanford, California, USA.
GigaScience
|March 3, 2026
概括
我们使用生成预训练变压器4 (GPT-4) 和扩展数据注释和检索中心 (CEDAR) 模板开发了一种方法来标准化科学元数据. 这种方法显著改善了数据集检索性能和回忆,加速了科学发现.
科学领域:
- 数据科学数据科学数据科学
- 生物信息学是一种生物信息学.
- 科学数据管理科学数据管理
背景情况:
- 科学元数据经常表现为不完整,不一致和格式错误,阻碍了有效的数据发现和重用.
- 标准化元数据对于科学研究中可靠的数据检索和可重复性至关重要.
研究的目的:
- 通过使用生成预训练变压器4 (GPT-4) 和扩展数据注释和检索中心 (CEDAR) 模板,展示一种用于自动元数据标准化和合规性的新方法.
- 通过提高元数据质量来提高科学数据集的检索性能.
主要方法:
- 将GPT-4与结构化的CEDAR元数据模板结合起来,以指导标准化过程.
- 将该方法应用于国家生物技术信息中心 (NCBI) 的生物样本和基因表达总 (GEO) 存储库.
- 将GPT-4+CEDAR的性能与基线原始元数据和GPT-4与数据字典指导 (GPT-4+DD) 的性能进行了比较.
主要成果:
- 采用GPT-4+CEDAR方法显著改善了数据集检索回忆,从17.65% (基线) 提高到62.87%.
- 与GPT-4+DD和基线元数据相比,GPT-4+CEDAR表现优越.
- 与LLaMA-3和MedLLaMA2.2等其他大型语言模型相比,GPT-4显示出一致的性能优势.
结论:
- 将像GPT-4这样的高级语言模型与象征性元数据结构 (CEDAR模板) 结合起来,为元数据标准化提供了一个变革性的解决方案.
- 这种方法导致更有效和可靠的数据检索,加速科学发现和数据驱动的研究.
相关概念视频
Halo Effect
565
The halo effect is a cognitive bias in which an individual's overall impression influences judgments about their specific traits. This psychological phenomenon leads people to associate positive characteristics with those they perceive as generally good and negative characteristics with those they view as bad. This effect is particularly influential in social perception, professional evaluations, and decision-making processes.The Psychological Basis of the Halo EffectThe halo effect is rooted...
565
Weighted Mean
7.1K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
7.1K
One-Way ANOVA: Equal Sample Sizes
4.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.3K
Distribution Reliability and Automation
547
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
547
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
