一个集群意识的合成重新采样和机器学习框架,用于多类空气质量指数分类.
Gokulan Ravindiran1, K Karthick2, Sujatha Sivarethinamohan3
1Department of Civil Engineering, Dayananda Sagar College of Engineering, Bengaluru, 560111, Karnataka, India.
Environmental monitoring and assessment
|March 2, 2026
概括
机器学习模型有效地使用各种数据对空气质量进行分类. 一种新的集群意识合成过量采样方法显著提高了不平衡空气质量指数 (AQI) 数据的性能.
科学领域:
- 环境科学 环境科学
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 空气质量指数 (AQI) 对于评估污染和健康风险至关重要.
- 数据驱动的AQI分类面临着由于数据集不平衡,极端污染水平代表性不足的挑战.
- 现有的方法难以准确地分类稀疏,严重的空气污染类别.
研究的目的:
- 开发强大的机器学习模型,用于多类AQI分类.
- 通过使用先进的重新采样技术,解决AQI数据集中的严重类失衡问题.
- 评估集合学习模型在提高AQI分类准确性的有效性.
主要方法:
- 使用了包括主要空气污染物 (PM2.5,PM10,等) 在内的综合数据集. 和气象变量.
- 应用数据预处理:缺失值赋值,分布规范化和风向循环编码.
- 实施了一个集群意识的合成过量抽样 (CASO) 框架,集成随机过量抽样,ENN,KMeans-SMOTE和类均等.
主要成果:
- 集成的梯度增强模型 (LightGBM,XGBoost) 显示出卓越的性能.
- 在应用CASO框架后,获得了高测试平衡精度 (≥0.96).
- 拟议的重新抽样技术显著提高了不平衡的AQI数据的分类准确性.
结论:
- 将集群意识的合成重新采样与集体学习相结合,在严重数据不平衡的情况下大幅改善了AQI分类.
- 开发的框架为准确的空气质量评估提供了可靠和可解释的方法.
- 这项研究为处理机器学习应用中不平衡的环境数据集提供了强大的方法.
相关概念视频
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
Cluster Sampling Method
15.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.3K
Classification of Systems-I
644
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
644
Sampling Plans
1.1K
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
1.1K
Classification of Systems-II
540
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
540
Classification of Signals
1.5K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.5K

