一种深度多视图不平衡学习方法,用于识别来自社交媒体的信息性COVID-19推文
Kok Kiang Long1, Stephen Wai Hang Kwok2, Jayne Kotz3
1School of Information Technology, Murdoch University, Perth, Australia.
Computers in biology and medicine
|August 2, 2023
概括
这项研究介绍了D-SVM-2K,这是一种新的深度学习方法,用于识别有信息的COVID-19推文. 它有效地处理不平衡的数据,并使用多个数据视图来提高机器学习性能.
科学领域:
- 计算机科学 计算机科学
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 像Twitter这样的社交媒体是快速传播COVID-19信息的主要来源.
- 机器学习应用程序需要从庞大的,不平衡的数据集中过信息性推特.
- 现有的方法往往忽视了阶级不平衡和单视图学习的挑战.
研究的目的:
- 提出一种新的深度失衡多视图学习方法 (D-SVM-2K) 来识别有信息的COVID-19推文.
- 解决现有解决方案在处理不平衡的培训数据和单视图学习方面的局限性.
主要方法:
- 开发了D-SVM-2K,这是一个基于SVM-2K多视图学习框架的深度学习模型.
- 整合了来自不同特征提取技术的多个视图.
- 利用一个堆叠的深层结构与多个SVM-2K分类器和k-最近邻近算法来管理类不平衡.
- 在多个视角实现全球和本地深度合体学习.
主要成果:
- 在现实世界标记的推特数据集上的经验实验验证了D-SVM-2K的有效性.
- 拟议的方法成功地解决了COVID-19推文识别中的多视图类失衡问题.
- 与处理不平衡数据的现有方法相比,表现出优异的性能.
结论:
- D-SVM-2K提供了一个强大的解决方案,用于从社交媒体数据中识别有信息的COVID-19推文.
- 这种方法有效地减轻了阶级不平衡带来的挑战,并利用了多视角学习.
- 这项工作有助于提高下游机器学习应用程序使用社交媒体数据的准确性和效率.
相关概念视频
Classification of Illness
7.6K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.6K
Improving Translational Accuracy
11.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.6K


