単語埋め込み子統計推定値
Neil Dey1, Matthew Singer1, Jonathan P Williams2
1Department of Statistics, North Carolina State University.
まとめ
この研究は、点相互情報量(PMI)を通じたWord2Vecの解釈を提供する単語埋め込みの統計的フレームワークを導入しています。新しい欠損値推定値は、Word2Vecと同等の性能を持つ統計的に健全な代替手段を提供します。
科学分野:
- 自然言語処理
- 統計理論
- 機械学習
背景:
- 単語埋め込みはNLPにおいて重要ですが、理論的な理解が不足しています。
- 現在の評価は、厳密な特性ではなく、経験的なパフォーマンスに依存しています。
- 形式的な推論と不確実性の定量化には、理論的な基盤が必要です。
研究 の 目的:
- 単語埋め込みの統計的理論的視点を提供すること。
- 古典的な方法、例えばWord2Vecを形式的な統計モデル内で解釈すること。
- 既存の単語埋め込み技術に代わる、統計的に扱いやすい新しいものを開発すること。
主な方法:
- テキストデータのためのコピュラベースの統計モデルを提案しました。
- Word2Vecを理論的な点相互情報量(PMI)の推定値として解釈しました。
- 以前の研究に基づいて、欠損値ベースの推定値を開発しました。
主要な成果:
- Word2Vecと理論的PMIの推定との関連を実証しました。
- 提案された欠損値推定値は、Word2Vecと同等の推定誤差を示します。
- 新しい推定値は、切り捨てベースの方法よりも優れた性能を発揮します。
- IMDb感情分析タスクでWord2Vecと同等の性能を達成しました。
結論:
- コピュラベースのモデルは、単語埋め込みの理論的な基盤を提供します。
- 欠損値推定値は、統計的に解釈可能で効果的な代替手段を提供します。
- この研究は、単語埋め込みにおける経験的な成功と理論的な理解との間のギャップを埋めます。
関連する概念動画
Estimating Population Standard Deviation
3.3K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.3K
Estimating Population Mean with Unknown Standard Deviation
8.7K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
8.7K
What are Estimates?
8.0K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
8.0K
Estimating Population Mean with Known Standard Deviation
9.6K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
9.6K
Statistical Significance
21.0K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
21.0K
Empirical Method to Interpret Standard Deviation
9.3K
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...
9.3K

