在类别事件数据中的变化点的概率比率与数字法医学的应用
Rachel Longjohn1, Padhraic Smyth2
1Department of Statistics, University of California, Irvine, Irvine, California, USA.
Journal of forensic sciences
|April 1, 2024
概括
本研究介绍了一个未知的变化点概率比率模型用于数字取证. 它准确地处理不确定的事件数据,优于具有固定变化点的模型.
科学领域:
- 数字法医学数字法医学
- 统计建模 统计建模
- 贝叶斯的推理是贝叶斯的推理.
背景情况:
- 数字取证通常分析时间标记的用户生成的事件数据.
- 区分来自单个所有者的数据与多个用户的数据 (例如,被黑客攻击的帐户) 是至关重要的.
- 现有的方法需要一个精确的,已知的变化点来计算概率比率.
研究的目的:
- 为数字法医学开发一种新的概率比率模型.
- 为了应对设备/帐户所有权变更 (变更点) 的确切时间不确定性的挑战.
- 为了利用贝叶斯的技术,一个未知的变化点概率比率模型.
主要方法:
- 开发了一个贝叶斯概率比率模型,适应未知的变化点.
- 获得了一个闭式,计算上简单的概率比表达式.
- 使用模拟变化点与现实世界数据集进行评估.
主要成果:
- 未知变化点模型表现出与已知变化点模型具有完美指定的变化点的性能相比.
- 拟议的模型显著优于已知的变化点模型,该模型具有错误的变化点.
- 结果强调了结合变化点不确定性的优势.
结论:
- 开发的未知变化点概率比率模型对数字取证有效.
- 这种贝叶斯式方法提供了一个强大的解决方案,当数据来源的确切时间是不确定的.
- 该模型在涉及用户生成数据的法医调查中提供了更高的准确性和可靠性.
相关概念视频
Hazard Rate
104
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
104
Censoring Survival Data
88
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
88
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Hazard Ratio
118
The hazard ratio (HR) is a widely used measure in clinical trials to compare the risk of events, such as death or disease recurrence, between two groups over time. It reflects the ratio of hazard rates—the instantaneous risk of the event occurring—between a treatment group and a control group. This measure provides valuable insights into the relative effectiveness of a treatment by assessing how the risk of an event differs between the two groups.
For example, in a clinical trial...
For example, in a clinical trial...
118
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K


