有时果确实落在树上很远的地方:PubMed中关于自动索引精度错误的案例研究
1paije.wilson@wisc.edu, Health Sciences Librarian, University of Wisconsin-Madison School of Medicine and Public Health, Madison, WI.
Journal of the Medical Library Association : JMLA
|November 14, 2025
概括
自动索引错误发生在术语具有多个含义时. 这项研究发现,7.9%的用"Malus" (果属) 索引的记录是不正确的,通常是由于比喻语言.
科学领域:
- 图书统计学 图书统计学
- 信息科学 信息科学 信息科学
- 医疗信息学 医疗信息学
背景情况:
- 自动索引系统对于组织生物医学文献在PubMed.com等数据库中至关重要.
- 索引的精确性可以确保准确地检索相关的科学信息.
- 医学科目标题词汇库 (MeSH) 是一个用于索引的受控词汇库.
研究的目的:
- 在MEDLINE记录的自动索引中识别和量化精度错误.
- 分析与MeSH术语*Malus*相关的索引错误的流行率和类型.
- 在处理具有非字面意义的术语时,评估自动索引的准确性.
主要方法:
- 选择了一组由1705个自动索引为*Malus*的MEDLINE记录组成的子集.
- 记录经过了标题,摘要和全文的手动选,以进行正确的索引.
- 索引错误根据上下文和类型进行分类.
主要成果:
- 7.9% (1705个中135个) 的记录被错误地用*Malus*进行索引.
- 主要的错误来源是"果"在比喻,比喻和成语中的使用 (59.2%).
- 其他错误包括"果"在名称/术语 (37%) 和缩写.
结论:
- 自动索引系统在遇到具有比喻或替代意义的单词时可能会产生错误.
- 图书馆员和信息专业人员应该意识到这些潜在的索引不准确性.
- 可能需要策略来减轻这些错误对文献搜索的影响.
相关概念视频
Systematic Error: Methodological and Sampling Errors
8.6K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
8.6K
Uncertainty in Measurement: Accuracy and Precision
99.5K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
99.5K
Accuracy and Precision
13.9K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
13.9K
Random and Systematic Errors
14.3K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
14.3K
Bias in Epidemiological Studies
1.2K
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
1.2K
Accuracy and Errors in Hypothesis Testing
555
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
555


