自动驾驶汽车撞车严重性的分类:解决不平衡数据集和小样本大小的问题
Pei-Fen Kuo1, Wei-Ting Hsu1, Dominique Lord2
1Department of Geomatics, National Cheng Kung University, Taiwan.
Accident; analysis and prevention
|June 20, 2024
概括
这项研究解决了自动驾驶汽车 (AV) 碰撞分析中的数据局限性,发现车辆类型和学校附近的位置等因素会影响AV碰撞的严重程度. 数据平衡和特征选择增强了预测模型,使AV开发更安全.
科学领域:
- 运输安全运输安全
- 数据科学数据科学数据科学
- 人工智能的人工智能
背景情况:
- 关于环境和道路特征对自动驾驶汽车 (AV) 事故严重性的影响的研究有限.
- 现有研究往往忽略了关键数据限制,如小样本大小,不平衡的数据集和高维度.
研究的目的:
- 为了分析自动驾驶汽车 (AV) 2019-2021年的撞车数据,结合环境和道路特征.
- 为了解决数据的局限性,包括小样本大小,不平衡的数据集和高维特征在AV撞击严重性分析.
- 为了比较特征选择与数据平衡作为初始预处理步骤的有效性.
主要方法:
- 使用加利福尼亚州机动车辆部 (CA DMV) 的AV撞车数据集 (266份报告).
- 整合了来自OpenStreetMap (OSM) 和Data San Francisco的兴趣点 (POI) 和道路特征的外部数据.
- 应用随机过量采样示例 (ROSE) 和合成少数人过量采样技术 (SMOTE) 用于数据平衡.
- 员工相互信息,随机森林和XGBoost用于特征选择和维度减少.
主要成果:
- 自动驾驶汽车 (AV) 碰撞的严重程度与汽车制造商,损坏程度,碰撞类型,移动,参与方,速度限制和接近POI (交通,娱乐,公共空间,学校,医疗设施) 相关.
- 数据重新抽样和特征选择方法都改善了模型性能.
- 采用SMOTE和数据平衡作为最初步骤的模型表现出卓越的性能.
结论:
- 数据平衡和特征选择技术显著提高了AV碰撞严重程度模型的预测性能.
- 接近特定的感兴趣点 (POI) 成为影响AV事故严重性的新因素.
- 这些发现为开发更强大的自动驾驶汽车安全系统提供了基础,通过解决数据挑战和识别关键风险因素来开发更强大的自动驾驶汽车安全系统.
相关概念视频
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K
One-Way ANOVA: Unequal Sample Sizes
5.7K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.7K
One-Way ANOVA
7.9K
One-way ANOVA analyzes more than three samples categorized by one factor. For example, it can compare the average mileage of sports bikes. Here, the data is categorized by one factor - the company. However, one-way ANOVA cannot be used to simultaneously compare the sample mean of three or more samples categorized by two factors. An example of two factors would be sports bikes from different companies driven in different terrains, such as a desert or snowy landscape. Here, two-way ANOVA is used...
7.9K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Aggregates Classification
317
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
317
Classification of Systems-I
179
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
179


