利用序列对序列模型进行荷兰病理学报告的语义注释
M Siepel1,2, G T N Burger3,4, Q J M Voorham5
1Amsterdam UMC, University of Amsterdam, Department of Medical Microbiology and Infection Prevention, Amsterdam Public Health Research Institute, Digital Health & Methodology, Amsterdam, the Netherlands.
Journal of pathology informatics
|January 15, 2026
概括
使用T5变压器模型对荷兰病理学报告的自动注释显示出希望,在较短的文本上表现优于标准模型,但在复杂的报告中面临挑战. 为了更广泛的应用,需要进一步开发.
科学领域:
- 医疗信息学医学信息学
- 计算病理学计算病理学
- 自然语言处理 (NLP) 是一种自然语言处理.
背景情况:
- 病理学报告的注释对于患者护理和研究至关重要,但是手动的,耗时的,容易出错的.
- 帕尔加基金会使用结论文本中的手动注释对荷兰病理学数据进行索引,将其映射到帕尔加词典中.
- 自动化这个注释过程可以提高效率和准确性.
研究的目的:
- 研究基于Text-To-Text转移变压器 (T5) 的模型用于荷兰病理学报告的自动注释.
- 将标准的多语言T5模型 (mT5) 与在Palga数据上预先训练的自定义T5模型 (PaTh5.NL) 进行比较.
- 评估受约束解码 (CD) 与默认解码 (DD) 对注释性能的影响.
主要方法:
- 开发了一个定制的T5模型 (PaTh5.NL),在Palga数据上进行预训练.
- 使用默认解码 (DD) 和受约束解码 (CD) 微调了mt5和PaTh5.NL模型.
- 使用双语评估基础研究 (BLEU) 评分和基于病例的患者检索评估来评估绩效.
主要成果:
- 微调的PaTh5.NL模型在较短的组织学和细胞学报告中显著超过mT5.
- 所有模型都显示在较长或更复杂的病理报告中性能下降.
- 基于案例的评估表明,较高的BLEU分数并不总是与mT5.5相比,通过PaTh5.NL模型转化为更好的患者检索.
结论:
- 精心调整的T5模型可以增强荷兰病理学报告的注释,特别是对于特定的报告类型.
- 复杂的结论文本仍然存在挑战,特别是在组织学和尸检报告中.
- 未来的工作应该专注于更大的数据集和后处理算法,以改善注释通用化.
关键词:
自动回归实体的检索深度学习是一种深度学习.病理学 病理学 病理学语义注释 语义注释 语义注释T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T5 T1 T1 T1 T1 T1 T2 T1 T2 T1 T1 T1 T2 T1 T1 T2 T1 T1 T2 T1变压器变压器变压器相关概念视频
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Sanger Sequencing
773.2K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
773.2K
Next-generation Sequencing
97.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.7K
Maxam-Gilbert Sequencing
12.6K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
12.6K
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K


