在孟加拉语句子中导航时代:用于增强预测的堆叠组合模型
Umme Ayman1, Md Nahid Hasan2, Ms Nusrat Khan1
1Department of CSE, Daffodil International University Daffodil Smart City (DSC), Birulia, Savar, Dhaka, Bangladesh.
PloS one
|August 21, 2025
概括
这项研究为准确的孟加拉语时代分类引入了一个新的堆叠组合模型, 达到85%的准确性. 这使得孟加拉自然语言处理在机器翻译和语法校正等应用中取得了进步.
科学领域:
- 计算语言学
- 自然语言处理
- 机器学习
背景情况:
- 在孟加拉语自然语言处理 (NLP) 中,
- 精确的时态识别对于各种NLP应用至关重要,包括机器翻译,情感分析和语法校正.
- 现有的方法缺乏全面的方法来处理孟加拉语动词形态的复杂性.
研究的目的:
- 在孟加拉语句子中开发一个强大的堆叠组合模型.
- 解决孟加拉语中精确且大规模的时代分类差距.
- 为孟加拉语NLP任务建立一个新的绩效基准.
主要方法:
- 构建一个新的孟加拉语语库, BengaliTenseCorpus, 包含13,500个手动标记的句子, 跨越现在,过去和未来的时代.
- 实现一个堆叠组合框架,集成五个基本模型:随机森林,支持矢量机,XGBoost,LSTM和GRU.
- 严格的预处理技术应用于各种文本源,以确保数据质量和语言完整性.
主要成果:
- 拟议的堆叠组合模型在测试数据上实现了85%的分类准确性.
- 整体模型表现出优异的性能与单个基底模型相比.
- 这项工作代表了第一个将机器学习和深度学习结合在一起的大型系统,
结论:
- 开发的堆叠组合模型为孟加拉语时代分类提供了非常准确的解决方案.
- 这项研究为推进孟加拉语NLP应用提供了坚实的基础,
- 能够有效地捕捉孟加拉语动词形态的细微差别, 为未来的改进铺平道路.
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
End Point Prediction: Gran Plot
575
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
575
Per-Unit Sequence Models
115
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
115
Hindsight Biases
3.9K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.9K
Contingency Table
2.6K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.6K

