改善されたバランスされたランダムフォレスト (iBRF):不均衡のデータセットにおけるクラッシュの重度の分類のための柔軟なハイブリッド再サンプリングバックの枠組み
Seyed Iman Mohammadpour1, Javadreza Vahedi1
1Department of Civil Engineering, Faculty of Engineering, University of Bojnord, Bojnord, Iran.
Traffic injury prevention
|February 18, 2026
まとめ
クラッシュデータにおけるクラス不均衡は課題です. 改善されたバランスされたランダムフォレスト (iBRF) モデルは,再サンプリングとアンサンブル方法を組み合わせて,重度の衝突結果の分類を改善することで,これを効果的に対処します.
科学分野:
- 機械学習 (Machine Learning) とは,機械学習 (Machine Learning) について学ぶことです.
- データサイエンス データサイエンス
- 交通安全について,交通安全について.
背景:
- 階級の不均衡は,重度の衝突の結果など,希少だが重要な出来事を正確に分類する上で大きな課題となっている.
- 伝統的な再サンプリング方法は,オーバーフィッティングと情報損失につながり,モデルの信頼性を損なう可能性があります.
- 重度の衝突の結果の正確な分類は,効果的な安全対策に不可欠です.
研究 の 目的:
- 重度のクラッシュ結果分類におけるクラス不均衡を克服するために設計された新しいアンサンブルフレームワークである改善されたバランスのとれたランダムフォレスト (iBRF) を導入する.
- 様々な再サンプリング技術とアンサンブル学習を組み合わせて,予測の精度を向上させる堅牢なモデルを開発する.
主な方法:
- iBRFのフレームワークは,SMOTE,NCR,RUSを用いたマイノリティクラス保存とマジョリティクラスアンダーサンプリングを統合しています.
- 意思決定ツリーは,各イテレーションでバランスのとれたデータで訓練され,最終的な予測はG-meanに基づいた加重ソフト投票で生成されます.
- 性能は,G-meanとMCCを使用して,iBRFを他の機械学習モデルと比較して,ホールドアウトデータで評価されました.
主要な成果:
- iBRFは,さまざまな再サンプリング技術 (SMOTE,RUS,NCR,CTGANs) と最先端のアンサンブル方法 (XGBoost,RF,LightGBMなど) に比べて優れたパフォーマンスを示しました. ) を実施する.
- iBRFアルゴリズムは,単純なRFよりもG平均を9.30%,SMOTE-RFよりも4.50%大幅に改善しました.
- このモデルは,バランスのとれたメトリックに基づいて,不均衡なデータセットに対して,より高い分類精度を達成しました.
結論:
- iBRFアルゴリズムは,特に重度のクラッシュ結果を特定するために,不均衡なデータセットの分類パフォーマンスを効果的に向上させます.
- このフレームワークは,マイノリティクラス (重度のクラッシュ) に関連する重要なリスク要因を特定するのに役立ちます.
- iBRFは,交通安全分析と介入戦略を改善するための有望なアプローチを提供します.
関連する概念動画
Survival Tree
440
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
440
Multiple Regression
4.1K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.1K
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
Randomized Experiments
9.1K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
9.1K
Classification of Systems-I
618
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
618
Classification of Systems-II
523
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
523

