Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Cross-reactivity00:42

Cross-reactivity

30.9K
Overview
30.9K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Revised Adaptive Immune Receptor Data in the Immune Epitope Database.

bioRxiv : the preprint server for biology·2026
Same author

Cancer epitope prediction tools and analysis pipelines in CEDAR.

Nucleic acids research·2026
Same author

VO: The Vaccine Ontology.

Scientific data·2026
Same author

The Cell Ontology in the age of single-cell omics.

Scientific data·2026
Same author

NetMHCIIphosPan: A Machine Learning Tool for Predicting HLA Class II Antigen Presentation of Phosphorylated Peptides.

Journal of proteome research·2026
Same author

Empiric azithromycin alters the upper respiratory microbiome and resistome without anti-inflammatory benefit in COVID-19.

Nature microbiology·2026

相关实验视频

Updated: May 21, 2025

A High Throughput MHC II Binding Assay for Quantitative Analysis of Peptide Epitopes
07:59

A High Throughput MHC II Binding Assay for Quantitative Analysis of Peptide Epitopes

Published on: March 25, 2014

14.9K

标准化自由文本数据,以免疫皮数据库中的两个字段为例.

Sebastian Duesing1, Jason Bennett2, James A Overton3

  • 1Center for Vaccine Innovation, La Jolla Institute for Immunology, La Jolla, CA, 92037, USA. sduesing@lji.org.

Journal of biomedical semantics
|March 23, 2025
PubMed
概括

本研究介绍了一种工具来规范非结构化生物医学文本数据,提高其用于自动化分析和数据查询的可用性. 规范化过程显著减少了数据的差异,提高了可搜索性,并使其与正式的本体学相集成.

关键词:
数据规范化的数据规范化.数据标准化数据标准化自由文本数据的数据.免疫表位基因数据库中的免疫表位基因.存在论 (Ontology) 是一种存在论.非结构化数据是非结构化数据.

更多相关视频

Synthetic Antigen Controls for Immunohistochemistry
09:30

Synthetic Antigen Controls for Immunohistochemistry

Published on: August 23, 2021

2.4K
Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing
08:51

Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing

Published on: March 15, 2019

12.3K

相关实验视频

Last Updated: May 21, 2025

A High Throughput MHC II Binding Assay for Quantitative Analysis of Peptide Epitopes
07:59

A High Throughput MHC II Binding Assay for Quantitative Analysis of Peptide Epitopes

Published on: March 25, 2014

14.9K
Synthetic Antigen Controls for Immunohistochemistry
09:30

Synthetic Antigen Controls for Immunohistochemistry

Published on: August 23, 2021

2.4K
Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing
08:51

Identification of Mouse and Human Antibody Repertoires by Next-Generation Sequencing

Published on: March 15, 2019

12.3K

科学领域:

  • 生物医学信息学 生物医学信息学
  • 数据科学数据科学数据科学
  • 自然语言处理自然语言处理.

背景情况:

  • 由于提取挑战,非结构化的生物医学自由文本未得到充分利用.
  • 数据规范化是使用结构化词汇和本体学的关键.
  • 本研究的重点是从免疫表皮层数据库 (IEDB) 中对"年龄"和"数据位置"字段进行规范化.

研究的目的:

  • 提出一个可适应的工具,用于规范化自由文本生物医学数据.
  • 评估该工具在IEDB内的特定领域的应用.
  • 展示规范化如何提高数据的可查性和可用性.

主要方法:

  • 开发了一个三步规范化过程:字符,单词和短语的规范化.
  • 创建了使用开发的工具应用的可通用规则.
  • 在IEDB中将该工具应用于4095个不同的"年龄"值和251,810个"数据位置"值.

主要成果:

  • 规范化工具在两个数据集的所有阶段都实现了高输出有效性.
  • 字符规范化产生了>99.97%的有效性;单词规范化>98.06%;短语规范化>83.81%.
  • 具体来说",年龄"数据在句子规范化后达到83.81%的有效性,而"数据位置"数据达到97.95%.

结论:

  • 成功开发了一种可通用的方法来规范自由文本数据库字段.
  • 规则创建的一次性努力可以应用于持续的数据策划.
  • 标准化显著减少了数据变异,改善了搜索功能和本体学链接.