一个基于denoising自编码器和深度生存回归的临床试验终止预测模型
Huamei Qi1, Wenhui Yang1, Wenqin Zou2
1School of Electronic Information Central South University Changsha Hunan China.
Quantitative biology (Beijing, China)
|February 12, 2026
概括
预测临床试验的完成是至关重要的. 一种新的自编码器和DeepSurv (DAE-DSR) 模型改进了生存预测的准确性,特别是在涉及孕妇的试验中稀少的数据.
科学领域:
- 生物统计学 生物统计学
- 临床试验方法论 临床试验方法论
- 医疗保健中的机器学习
背景情况:
- 有效的临床试验对于医学进步至关重要,但早期终止会导致资源浪费.
- 生存模型预测试验结果,但稀少的数据挑战了DeepSurv等现有模型,限制了准确性和概括性.
- 临床试验往往不包括孕妇,因此需要为这一群体提供专门的预测模型.
研究的目的:
- 使用稀疏的数据开发一个改进的生存预测模型来完成临床试验.
- 提高临床试验中生存分析的特征表示能力.
- 具体解决涉及孕妇的研究中试验完成的预测问题.
主要方法:
- 提出了一种混合模型,将无声自动编码器 (DAE) 与DeepSurv模型 (DAE-DSR) 结合起来.
- 利用DAE从原始临床试验数据中提取可靠的特征表示.
- 在clinicaltrials.gov的数据集上训练了DAE-DSR模型,专注于孕妇的试验.
主要成果:
- DAE-DSR模型有效地提取了生存分析的有意义和强大的特征.
- 在训练数据集上达到0.74的C指数,在测试数据集上达到0.75的C指数.
- 与传统的Cox比例危险和独立的DeepSurv模型相比,表现出优越的性能和稳定性.
结论:
- 拟议的DAE-DSR模型显著提高了对稀疏数据的临床试验的生存预测准确度.
- 模型捕捉强大的特征的能力提高了概括性和预测能力.
- 这种方法为预测临床试验完成提供了更可靠的工具,特别是在像孕妇这样的代表性不足的群体中.
相关概念视频
Clinical Trials
10.9K
Clinical trials are prospective experimental studies conducted on humans to determine the safety and efficacy of treatments, drugs, diet methods, and medical devices. Using statistics in clinical trials enables researchers to derive reasonable and accurate conclusions from the collected data, allowing them to make wise decisions in uncertain situations. In medical research, statistical methods are crucial for preventing errors and bias.
There are four phases in a clinical trial. A phase one...
There are four phases in a clinical trial. A phase one...
10.9K
Clinical Trials: Overview
5.0K
Clinical development focuses on how the drug will interact with the human body and encompasses four key phases of clinical trials, each serving a specific purpose in assessing the safety and effectiveness of new drugs. These phases overlap and build upon one another. Phase I involves a small group of healthy volunteers (typically 20-80 individuals) or, in cases where significant toxicity is expected, patients with the targeted disease, such as cancer or AIDS. The volunteers are tested for...
5.0K
Regression Toward the Mean
7.2K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.2K
Statistical Software for Data Analysis and Clinical Trials
1.6K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.6K
Termination of Translation
27.9K
The large ribosomal subunit has several important structures essential to translation. These include the peptidyl transferase center (PTC) - which is the site where the peptide bond is formed - and a large, internal, water-filled tube through which the nascent polypeptide moves. This latter structure is called the Peptide Exit Tunnel, and it begins at the PTC and spans the body of the large ribosomal subunit. During translation, as the nascent polypeptide chain is synthesized, it passes through...
27.9K
Multiple Regression
4.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.0K


