DrivR-Base:用于变量效应预测模型构建的特征提取工具包
Amy Francis1, Colin Campbell2, Tom R Gaunt1
1MRC Integrative Epidemiology Unit, Bristol Medical School (PHS), University of Bristol, Bristol BS8 2BN, United Kingdom.
Bioinformatics (Oxford, England)
|April 11, 2024
概括
DrivR-Base简化了对人类基因组变异的分子特征的提取,有助于疾病预测. 这个资源简化了机器学习模型的输入,加速了对遗传变异病原性的研究.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 测序技术的进步已经确定了许多人类基因组变异.
- 了解这些变体在疾病发病过程中的功能作用是具有挑战性的.
- 目前用于预测变异病原性的方法需要从各种数据源中提取复杂的特征.
研究的目的:
- 引入DrivR-Base,这是一种用于高效提取和整合单核酸变体分子特征的新资源.
- 为机器学习应用程序提供一种用户友好和可重复的方法来获得变体相关的特性.
- 为了促进病原性人类基因组变异的预测,并支持其他基因组分析.
主要方法:
- DrivR-Base集成了来自多个数据库和工具的功能,包括AlphaFold,ENCODE和变量效应预测器.
- 特性包括基因组和蛋白质位置,结构性质,监管信息和预测的变异后果.
- 该资源可以通过Docker容器部署,以提高可访问性和可重复性.
主要成果:
- DrivR-Base有效地提取和整合单核酸变体的综合分子特征.
- 生成的特征集适合输入到机器学习模型中,用于病原性预测.
- 该资源支持各种应用,包括哈普洛缺陷预测和药物重定位.
结论:
- DrivR-Base为基因组变异分析的特征提取提供了显著的提高效率和可访问性.
- 该资源使研究人员能够构建更强大的遗传疾病预测模型.
- DrivR-Base有可能在未来扩展,包括额外的数据类型和分析功能.
相关概念视频
Extraction: Advanced Methods
446
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
446
Extraction: Partition and Distribution Coefficients
2.4K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.4K
Variability: Analysis
141
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
141
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
487
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
487
Sensitivity, Specificity, and Predicted Value
305
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
305
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


