使用机器学习和字符串距离相似性标准化临床实验室分类测试结果的新方法
Syed Ahmmed1, M Rubaiyat Hossain Mondal1, Md Raihan Mia2,3
1Institute of Information and Communication Technology, Bangladesh University of Engineering And Technology, Dhaka, Bangladesh.
Heliyon
|November 30, 2023
概括
标准化临床实验室测试结果对于数据科学至关重要. 这项研究引入了一种具有Jaro-Winkler相似性的机器学习方法,实现了临床数据标准化的高准确性.
科学领域:
- 医疗信息学 医疗信息学
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 标准化临床实验室测试结果对于可靠的临床数据科学研究至关重要.
- 目前的数据处理工具和标准化指南不足.
研究的目的:
- 提出一种新的,可扩展的方法来标准化分类临床实验室测试结果.
- 提高临床大数据的可用性,用于研究,特别是在资源有限的环境中.
主要方法:
- 监督机器学习模型 (支持矢量分类) 用于对测试结果进行分类.
- 用Jaro-Winkler相似算法将文本测试结果映射到类别内的标准化临床术语中.
- 该方法在来自孟加拉国的75,062个测试结果和MIMIC-III数据集上得到了验证.
主要成果:
- 支持矢量分类在分类测试结果中实现了98%的准确性,超过了随机森林.
- 贾罗-温克勒相似性在大多数组的标准化测试结果中显示出99.93%的成功率.
- 与以前基于规则和距离相似性的方法相比,拟议的方法显示出更高的性能.
结论:
- 这种新的方法有效地使用机器学习和Jaro-Winkler相似性标准化了分类临床实验室测试结果.
- 这种技术显著有利于临床大数据研究和国家临床研究数据中心的发展.
- 该方法在地方和公共数据集上表现出色,突出了其广泛的适用性.
关键词:
数据质量数据质量数据质量数据科学是数据科学.电子健康记录是电子健康记录.在LOINC中,您可以使用LOINC.机器学习 机器学习这就是SNOMED CT.标准化 标准化 标准化字符串距离的相似性更多相关视频
06:22Standardization of Transfer across Labs between Flow Cytometers for Detection of Lymphocytes in Japanese Encephalitis Vaccinated Children
Published on: February 10, 2023
1.0K
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
8.7K
相关概念视频
Testing a Claim about Standard Deviation
2.5K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.5K
Spearman's Rank Correlation Test
810
Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates...
Spearman's test calculates...
810
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
How Data are Classified: Categorical Data
33.1K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
33.1K
