一种基于NLP的技术,可以从药物SMILES中提取有意义的特征
Rahul Sharma1, Ehsan Saghapour1, Jake Y Chen1
1Informatics Institute, School of Medicine, The University of Alabama at Birmingham, Birmingham, AL, USA.
iScience
|March 8, 2024
概括
这项研究引入了一种新的自然语言处理 (NLP) 方法,用于从药物分子结构中提取特征,其表示为SMILES符号. 这种方法增强了机器学习模型,用于个性化药物查,并加速了药物发现.
科学领域:
- 计算化学是一种计算化学.
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 自然语言处理 (NLP) 模型的序列数据,类似于药物分子的SMILES符号.
- 在SMILES中的特殊字符具有特定的含义,对于药物结构分析至关重要.
- 现有的方法可能无法充分利用SMILES的顺序性质来提取特征.
研究的目的:
- 开发一种基于NLP的新方法,从药物SMILES标记中提取可解释的序列和特征.
- 将NLP衍生特征与传统的摩根指纹位向量进行比较.
- 在个性化药物查 (PSD) 案例研究中验证该方法的有效性.
主要方法:
- 利用NLP中的N-gram来提取药物SMILES中的序列特征.
- 采用基于UMAP的嵌入来比较NLP特征与摩根指纹.
- 集成基于NLP的功能,用于机器学习模型的基因表达和疾病表型数据.
主要成果:
- 基于NLP的功能对于PSD来说是稀疏而有效的.
- 结合的功能导致了改进的机器学习模型,用于个性化药物查.
- 通过两个成功的个性化药物查案例研究证明了它的实用性.
结论:
- 这种基于NLP的新方法提供了一种新的方法来分析SMILES标记中的药物分子结构.
- 这种方法可以加速药物发现和开发工作.
- 开发的方法可以作为一个可访问的Python库.
相关概念视频
Drug Discovery: Overview
7.9K
Drug discovery is a multifaceted process involving extensive screening, testing, and optimization of lead compounds to identify potential new drugs for therapeutic use. It combines several approaches, including screening large numbers of natural products, chemical modification of known active molecules, identification of new drug targets, and rational design based on biological mechanisms and drug-receptor structure. These approaches are carried out in both academic research laboratories and...
7.9K
Drug-Receptor Bonds
2.8K
Drug-receptor bonds are formed through various chemical forces when drugs interact with target cells. Covalent bonds, strong and irreversible, are exemplified by DNA-alkylating anticancer agents that inhibit cell division. However, such irreversible drug binding lacks selectivity and can modify the DNA of the surrounding healthy cells. Covalent binding often contributes to tissue toxicity, as seen with chloroform and paracetamol metabolites binding to the liver, causing hepatotoxicity.
In...
In...
2.8K


