NISQE:基于自然统计数据的非侵入性语音质量评估器,平均减去对比度,谱图的规范系数
Shakeel Zafar1, Imran Fareed Nizami2, Mobeen Ur Rehman3
1Department of Computer Engineering, University of Engineering and Technology, Taxila 47050, Pakistan.
Sensors (Basel, Switzerland)
|July 8, 2023
概括
本研究引入了一种使用自然光谱图统计性质的新型非侵入性语音质量评估方法. 该技术通过分析自然语音结构的偏差,有效地预测语音质量,优于现有的方法.
科学领域:
- 信号处理 信号处理
- 声学 声学 在声学上
- 机器学习 机器学习
背景情况:
- 在会议和VoIP等现代应用中,语音通信至关重要.
- 由于噪音和技术限制,语音信号的退化需要持续的质量评估.
- 非侵入性语音质量评估 (NI-SQA) 是一个挑战,因为缺乏原始的参考信号.
研究的目的:
- 提出一种新的非侵入性语音质量评估 (NI-SQA) 方法.
- 利用语音信号的自然结构进行更准确的质量预测.
- 根据最先进的技术来评估拟议的方法.
主要方法:
- 开发了一种基于自然光谱图统计 (NSS) 属性的NI-SQA方法.
- 从语音信号谱图中提取NSS特征,以近似自然语音结构.
- 通过测量原始和扭曲信号之间的NSS属性的偏差来量化语音质量.
主要成果:
- 在VCTK-Corpus上实现了高性能,Spearman排序相关常数 (SRC) 为0.902,Pearson相关常数 (PCC) 为0.960,根平均平方误差 (RMSE) 为0.206.
- 在NOIZEUS-960数据库中表现出优异的结果,SRC为0.958,PCC为0.960,RMSE为0.114.
- 提出的方法显著优于现有的最先进的NI-SQA技术.
结论:
- 拟议的NI-SQA方法有效地通过利用语音信号的自然结构来预测语音质量.
- 在没有原始信号的情况下,NSS属性提供了强大的功能来评估语音质量.
- 这种方法为在现实场景中可靠的语音质量评估提供了有希望的解决方案.
更多相关视频
相关概念视频
Root Mean Square
3.3K
If in an experiment, data values have a probability of being both positive and negative, neither the arithmetic mean, the geometric mean, nor the harmonic mean can be used to calculate the central tendency of the data set. In particular, if the positive and negative values are equally likely, the arithmetic mean is close to zero.
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
3.3K
Mean Absolute Deviation
2.7K
The mean absolute deviation is also a measure of the variability of data in a sample. It is the absolute value of the average difference between the data values and the mean.
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
2.7K
Pulse amplitude and quality
1.9K
Pulse amplitude is a crucial indicator of cardiac health because it provides valuable insights into the strength of left ventricular contractions and the overall uniformity of blood circulation within the vasculature. The strength of the pulse is directly related to the force with which the heart contracts and the volume of blood being pumped.
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
1.9K
Trimmed Mean
2.9K
While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
2.9K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Sound Intensity Level
4.2K
Humans perceive sound by hearing. The human ear helps sound waves reach the brain, which then interprets the waves and creates the perception of hearing. The loudness of the environment in which a person is located determines whether they can distinguish between different sound sources.
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
4.2K


