阿尔卡:用于机器学习分类建模,风险评估和填补稀缺的环境毒性数据数据缺口的维度缩小框架
Arkaprava Banerjee1, Kunal Roy1
1Drug Theoretics and Cheminformatics Laboratory, Department of Pharmaceutical Technology, Jadavpur University, Kolkata 700 032, India. kunal.roy@jadavpuruniversity.in.
Environmental science. Processes & impacts
|May 14, 2024
概括
量化结构-活动关系 (QSAR) 模型在小数据集上扎. 在K组分析 (ARKA) 描述符中的新型算术余量改善了对小数据集的生态毒理学分类建模和预测,优于传统的QSAR方法.
科学领域:
- 计算毒理学和化学信息学.
- 环境风险评估和化学安全.
背景情况:
- 对环境化学品的实验性毒性数据很少,需要*in silico*方法.
- 量化结构-活动关系 (QSAR) 模型对于小数据集来说是常见的,但由于高的描述器-复合比率,因此存在可靠性问题.
- 在QSAR中减少模型自由度,限制小数据集的预测准确度.
研究的目的:
- 在K组分析 (ARKA) 描述符中引入新的算术余数,以提高QSAR模型的可靠性.
- 以监督的方式减少建模描述符的数量,防止化学信息丢失.
- 改善生态毒理学分类建模和预测小数据集的准确性.
主要方法:
- 通过将描述符分成K类 (K=2) 的新型ARKA描述符的计算,基于平均规范值.
- 应用ARKA框架对五个环境相关的终点进行分类建模,并分级响应.
- 开发一个基于Java的专家系统,从QSAR描述器计算ARKA描述器.
主要成果:
- 与传统的QSAR描述器相比,ARKA描述器显著提高了分类模型的预测质量.
- 来自ARKA描述符的模型在多个分级数据验证指标中表现出卓越的性能.
- 此外,ARKA描述器也提高了阅读交叉预测的准确性,超过了QSAR描述器.
结论:
- 在生态毒理学中,ARKA描述器提供了一种有希望的方法来克服QSAR建模的局限性,使用小型数据集.
- 该ARKA框架有效地减少了预测错误,并提高了环境化学毒性评估的模型可靠性.
- 开发的专家系统为用户方便ARKA描述符的实际应用.
相关概念视频
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
44
Flood risk assessment involves careful planning and analysis to ensure the safety of communities near water retention structures. Capacity contours are a vital tool in this process, as they illustrate the potential spread of water at specific levels in a given area. In the context of building a bund across a small valley, these contours play a critical role in evaluating the safety of nearby residential areas.In this example, the bund is intended to store stormwater in the valley. The engineers...
44
Statistical Methods for Analyzing Epidemiological Data
363
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
363
Statistical Software for Data Analysis and Clinical Trials
539
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
539
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
Aggregates Classification
317
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
317
Survival Tree
80
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
80


