Related Experiment Video
Updated: Jul 22, 2025

12:18
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.6K
Predictors of Sexual Harassment Using Classification and Regression Tree Analyses and Hurdle Models: A Direct
Jan-Louw Kotzé1, Patricia A Frazier1, Kayla A Huber1
1Department of Psychology, University of Minnesota Twin Cities.
Journal of Sex Research
|July 24, 2023
Summary
Sexual harassment is prevalent in US higher education. This study confirmed previous findings on risk factors, showing that younger age, alcohol use, and prior adversity predict peer harassment, while LGBQ+ identity predicts faculty/staff harassment.
Area of Science:
- Sociology
- Higher Education Studies
- Public Health
Background:
- Sexual harassment is a significant issue affecting numerous students in U.S. higher education institutions.
- Previous research identified key risk factors using statistical models.
Purpose of the Study:
- To validate and replicate previously identified risk factors for sexual harassment among college students.
- To assess the robustness of findings from prior hurdle model and classification and regression tree (CART) analyses.
Main Methods:
- Secondary data analysis of 9,552 students from two- and four-year colleges.
- Replication of hurdle models and CART analyses to assess statistical significance and effect size consistency.
- Evaluation of replicability criteria for model coefficients and variable predictions.
Main Results:
- The original findings on sexual harassment risk factors were robustly replicated.
- 91% of effects in hurdle models and 88% of variables in CART analyses met replication criteria.
- Key predictors included younger age, frequent alcohol consumption, four-year college attendance, prior victimization for peer harassment, and LGBQ+ identity for faculty/staff harassment.
Conclusions:
- The identified risk factors for sexual harassment are reliable across different student samples.
- Findings support the development of targeted prevention and intervention strategies in higher education.
- Further research is necessary to elucidate the mechanisms behind demographic and contextual associations with harassment risk.
More Related Videos
Related Concept Videos
Survival Tree
112
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
112
Hazard Rate
139
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
139
Stereotype Threat and Self-fulfilling Prophecies
37.7K
When we hold a stereotype about a person, we have expectations that he or she will fulfill that stereotype. A self-fulfilling prophecy is an expectation held by a person that alters his or her behavior in a way that tends to make it true. When we hold stereotypes about a person, we tend to treat the person according to our expectations. This treatment can influence the person to act according to our stereotypic expectations, thus confirming our stereotypic beliefs. Research by Rosenthal and...
37.7K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Analysis
5.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.8K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K

