基于机器学习的差异估计,采用两阶段采样,使用卫生和教育部门数据
Sanaa Al-Marzouki1, Ibrahim A Nafisah2, Mhassen E E Dalam3
1Statistics Department, Faculty of Science, King Abdul Aziz University, Jeddah, Kingdom of Saudi Arabia.
Scientific reports
|February 7, 2026
概括
本研究引入了用于双相采样的新差异估计器,使用最小的辅助数据提高了估计效率. 与现有估计器相比,该新方法显示出优越的分析和经验性能.
科学领域:
- 统计 统计 统计 统计
- 调查方法 调查方法
背景情况:
- 准确的差异估计对于可靠的统计推理至关重要.
- 当辅助信息有限时,传统方法可能缺乏效率.
- 两阶段采样提供了一个框架来纳入辅助数据.
研究的目的:
- 提出一种新的差异估计器,用于两相采样.
- 为了提高估计效率,使用一个辅助变量和一个二进制属性.
- 证明拟议估计者的分析和经验优势.
主要方法:
- 开发了一种包含辅助信息的新型差异估计器.
- 导出理论属性,包括偏差和平均平方误差 (MSE).
- 使用健康和教育数据集进行模拟研究.
- 训练并评估机器学习分类器 (回归树,随机森林,支持向量的回归).
主要成果:
- 拟议的估计器证明了分析优越性,具有被证明的偏差和MSE公式.
- 与经典和竞争性估计器相比,模拟结果显示MSE值始终较低.
- 机器学习模型显示出良好的预测能力,但拟议的估计器提供了更好的解释性.
- 三个参数的韦布尔分布被认为是最适合分析的.
结论:
- 新型差异估计器在两相采样中提高了估计精度.
- 最少的辅助信息,结构化采样和混合建模改善了差异估计.
- 拟议的估计器为应用研究提供了理论上健全和经验验证的方法.
相关概念视频
Variance
12.4K
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.
The standard deviation measures the spread in the same units as the data....
The standard deviation measures the spread in the same units as the data....
12.4K
Three-Phase Short Circuit—Unloaded Synchronous Machine
718
Conducting a three-phase short circuit test on an unloaded synchronous machine helps understand its impact on the system. The AC fault current's oscillogram, with the DC offset removed, reveals that the waveform amplitude decreases from an initially high value to a steady-state level for one phase of the machine.
This behavior occurs due to the magnetic flux produced by the short-circuit armature currents. Initially, these currents follow high-reluctance paths but eventually shift to...
This behavior occurs due to the magnetic flux produced by the short-circuit armature currents. Initially, these currents follow high-reluctance paths but eventually shift to...
718
Friedman Two-way Analysis of Variance by Ranks
511
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
511
Machines
581
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. One example of a machine is the cutting plier, which is used to cut wires by applying forces to its handles. When equal and opposite forces are exerted on the handles of the cutting plier, they cause the cutting edges to come together and apply equal and opposite reaction forces on the wire, which are greater than the applied forces.
A free-body diagram of the...
A free-body diagram of the...
581
Data Reporting and Recording
5.5K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.5K
What are Estimates?
8.8K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
8.8K


