在Android应用程序的大数据集上对机器学习模型的性能评估 审查 审查
Ali Adil Qureshi1, Maqsood Ahmad2, Saleem Ullah1
1Department of Computer Science, Khwaja Fareed University of Engineering and Information Technology, Rahim Yar Khan, 64200 Pakistan.
概括
使用机器学习分析移动应用程序评论有助于改进应用程序. 将统计特征与TF-IDF结合起来,并使用支向量机器模型,在情绪分析中实现了最高的准确性.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 移动应用程序越来越受欢迎,产生大量的用户评论.
- 分析这些审查对于应用程序的改进和开发至关重要,但由于它们的数量和复杂性,它们会带来挑战.
- 分类审查有助于用户选择合适的应用程序.
研究的目的:
- 提出一个框架来分析八个不同类别的移动应用程序评论的情绪分析.
- 评估机器学习模型与TF-IDF相结合的有效性,用于应用程序审查分析中的特征提取.
主要方法:
- 通过使用正规表达式和美丽的,删除了251,661条用户评论的数据集.
- 用于特征提取的使用术语频率-反向文档频率 (TF-IDF).
- 利用各种机器学习模型,包括支持矢量机器,并通过预处理和统计功能评估性能.
主要成果:
- 将统计特征与TF-IDF相结合,显著提高了模型性能.
- 支持矢量机模型在情绪分析中表现出最高的准确性.
- 这项研究为进一步研究提供了大量,多样化和平衡的数据集.
结论:
- 拟议的框架有效地分析了移动应用程序评论中的情绪.
- 结合TF-IDF和统计特征,以及SVM,为情绪分析提供了一个强大的方法.
- 这些发现可以指导研究人员选择应用程序审查分析的最佳模型,并有助于开发更好的移动应用程序.
相关概念视频
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K


