在噪音,失衡,特征减少和可解释性下对经典和量子机器学习进行了强有力的评估
Savita Kumari Sheoran1, Vikesh Yadav2, Rakesh Kumar Sheoran3
1Dept. of Computer Science & Engineering, Indira Gandhi University Meerpur, Rewari, Haryana, India.
Scientific reports
|December 5, 2025
概括
量子机器学习 (QML) 分类器显示出对复杂数据挑战的前景. 量子支持向量机器证明了对杂,不平衡的数据集的弹性,为强大的AI系统提供了洞察力.
科学领域:
- 计算机科学 计算机科学
- 量子计算是一种量子计算.
- 人工智能的人工智能
背景情况:
- 现实世界的数据往往会带来诸如噪音,不平衡和高维度等挑战,影响传统机器学习 (ML) 模型的性能.
- 量子机器学习 (QML) 通过使用量子计算为复杂的分类任务提供了潜在的优势.
- 在现实的数据条件下评估QML模型对于实际部署至关重要.
研究的目的:
- 对传统的监督ML分类器和QML分类器进行广泛的实验比较.
- 评估不同数据集的模型性能,模拟现实世界的复杂性,如噪音,类失衡和高维度.
- 使用可解释AI (XAI) 工具调查ML模型的可解释性.
主要方法:
- 对比了五个监督的ML分类器 (决策树,K-NN,随机森林,线性回归,SVM) 与三个QML分类器 (量子SVM,量子K-NN,变量量子分类器).
- 使用了五个不同的数据集 (Iris,葡萄酒质量,乳腺癌,UCI HAR,Pima糖尿病).
- 引入了类不平衡 (SMOTE,ADASYN),特征噪声 (高斯式) 和维度减小 (ANOVA);为了可解释性,使用了SHAP和LIME.
主要成果:
- 后勤回归在各种复杂度上表现出一致的性能.
- 量子支持向量机器表现出了显著的弹性,以特征噪音和类不平衡.
- 可解释的人工智能工具为模型决策过程提供了洞察力.
结论:
- 量子ML模型,特别是量子SVM,显示出处理复杂,现实世界的数据挑战的潜力.
- 该研究强调了当前的QML能力和局限性,为可概括和可解释的ML系统的开发提供了信息.
- 这些发现对于在复杂的实际环境中部署强大的AI至关重要.
相关概念视频
Randomized Experiments
8.8K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
8.8K
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Propagation of Uncertainty from Systematic Error
1.2K
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
1.2K
Propagation of Uncertainty from Random Error
1.6K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.6K
Random and Systematic Errors
14.3K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
14.3K
Variability: Analysis
426
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
426
