使用物品响应理论作为一种方法来归纳分类缺失值.
Adrienne Kline1,2,3, Yuan Luo4,5
1Department of Surgery, Northwestern University, Chicago, USA. adrienne.kline@northwestern.edu.
Scientific reports
|November 5, 2025
概括
对分类归算的项目响应理论 (IRT) 有效地解决了缺失的数据,优于kNN和MICE等方法. 这种方法为提高机器学习和统计分析中的数据质量提供了一个强大的替代方案.
科学领域:
- 数据科学数据科学数据科学
- 统计 统计 统计 统计
- 机器学习 机器学习
背景情况:
- 缺少数据是数据集中常见的挑战,限制了统计推断和模型开发.
- 现有的归算技术对下游分析产生不同的影响,包括临床得分计算和模型测试.
研究的目的:
- 评估基于物品响应理论 (IRT) 的方法来进行分类数据归算.
- 为了将IRT方法与已建立的机器学习归算技术进行比较,例如k-最近邻居 (kNN),多重归算链式方程 (MICE) 和DataWig.
主要方法:
- IRT归算方法应用于有序,名义和二进制类别的三个数据集.
- 数据集被操纵以改变缺失的数据比例和系统化.
- 绩效通过价值再现的准确性和计算后的预测性绩效来评估.
主要成果:
- 在分类归算中,IRT方法表现出强的表现.
- 它在各种条件下超过了当前的几种多重归算方法.
- 该方法在复制缺失值和提高预测准确性方面表现特别有希望.
结论:
- 对分类归算的项目响应理论为现有方法提供了理论上合理和有效的替代方案.
- 它对填补缺失值的概率方法为数据质量和后续分析提供了优势.
- 对于研究人员来说,IRT方法是处理缺少分类数据的可行选择.
相关概念视频
How Data are Classified: Categorical Data
42.6K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
42.6K
Response Surface Methodology
595
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
595
Censoring Survival Data
516
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
516
Nominal Level of Measurement
36.9K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
36.9K
Systematic Error: Methodological and Sampling Errors
8.6K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
8.6K
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K


