自动编码器的调整用于从临床笔记中减少稀疏性 代表性学习学习
Thanh-Dung Le1,2, Rita Noumeir1, Jerome Rambaud2
1Biomedical Information Processing Laboratory, Ecole de Technologie SuperieureUniversity of Quebec Montreal QC H3C 1K3 Canada.
本研究引入了一种自编码算法,以减少临床文本数据的稀疏性,提高小型数据集的分类准确性. 该方法通过压缩特征空间来提高神经网络性能,达到高达92%的准确性.
科学领域:
- 医疗保健中的机器学习
- 临床文本的自然语言处理.
- 数据挖掘和表示学习学习学习
背景情况:
- 在小型数据集上的临床文本分类通常与特征空间稀疏性作斗争.
- 传统的特征选择方法可能无法充分解决临床数据中的非线性依赖性和稀疏性.
- 多层感知子看起来很有希望,但可以通过有效的特征表示进一步优化.
研究的目的:
- 开发一种替代方法来解决临床代表性特征空间的稀疏性.
- 为了有效地压缩高维度,稀疏的临床数据,使有限的数据集的分析,包括法国临床笔记.
- 通过在压缩的特征空间中对神经网络分类器进行评估来提高其性能.
主要方法:
- 提出了一个自编码器学习算法来减少维度,并利用临床笔记表示中的稀疏性.
- 该研究的重点是压缩临床笔记的特征空间,以增强学习表征.
- 用在压缩特征空间上训练并测试的分类器来评估分类性能.
主要成果:
- 自动编码器方法在测试组评估中产生了高达3%的整体性能增长.
- 分类器实现了高性能指标:92%的准确性,91%的回忆,91%的精度和91%的F1分数.
- 理论信息瓶框架被用来证明自动编码器的压缩和预测机制.
结论:
- 自动编码器学习有效地解决了小型临床叙事数据集的稀疏性,改善了下游分类.
- 算法的无损压缩容量允许学习最佳数据表示,在这种情况下超过深度学习模型.
- 这种方法显著提高了稀疏,高维度的临床数据的分类能力.
更多相关视频
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
相关概念视频
Improving Translational Accuracy
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
Regression Toward the Mean
Local Anesthetics: Clinical Application as Spinal Anesthesia
Associative Learning
Classical conditioning, also known...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
