Related Experiment Video
Updated: Jun 29, 2025

12:39
A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
11.4K
Likelihood ratios for changepoints in categorical event data with applications in digital forensics
Rachel Longjohn1, Padhraic Smyth2
1Department of Statistics, University of California, Irvine, Irvine, California, USA.
Journal of Forensic Sciences
|April 1, 2024
Summary
This study introduces an unknown changepoint likelihood ratio model for digital forensics. It accurately handles uncertain event data, outperforming models with fixed changepoints.
Area of Science:
- Digital Forensics
- Statistical Modeling
- Bayesian Inference
Background:
- Digital forensics often analyzes time-stamped user-generated event data.
- Distinguishing data from a single owner versus multiple users (e.g., hacked accounts) is crucial.
- Existing methods require a precise, known changepoint for likelihood ratio calculations.
Purpose of the Study:
- To develop a novel likelihood ratio model for digital forensics.
- To address the challenge of uncertainty in the exact time of device/account ownership change (changepoint).
- To utilize Bayesian techniques for an unknown changepoint likelihood ratio model.
Main Methods:
- Developed a Bayesian likelihood ratio model accommodating an unknown changepoint.
- Derived a closed-form, computationally straightforward expression for the likelihood ratio.
- Employed simulated changepoints with real-world datasets for evaluation.
Main Results:
- The unknown changepoint model demonstrated comparable performance to a known changepoint model with a perfectly specified changepoint.
- The proposed model significantly outperformed a known changepoint model with a misspecified changepoint.
- The results highlight the advantage of incorporating changepoint uncertainty.
Conclusions:
- The developed unknown changepoint likelihood ratio model is effective for digital forensics.
- This Bayesian approach offers a robust solution when the exact time of data origin is uncertain.
- The model provides improved accuracy and reliability in forensic investigations involving user-generated data.
Related Concept Videos
Hazard Rate
104
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
104
Censoring Survival Data
88
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
88
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Hazard Ratio
118
The hazard ratio (HR) is a widely used measure in clinical trials to compare the risk of events, such as death or disease recurrence, between two groups over time. It reflects the ratio of hazard rates—the instantaneous risk of the event occurring—between a treatment group and a control group. This measure provides valuable insights into the relative effectiveness of a treatment by assessing how the risk of an event differs between the two groups.
For example, in a clinical trial...
For example, in a clinical trial...
118
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K

