一个轻量级的基于混合的短文集群,用于对比学习
Qiang Xu1, HaiBo Zan1, ShengWei Ji1
1School of Artificial Intelligence and Big Data, Hefei University, Hefei, Anhui, China.
Frontiers in computational neuroscience
|February 13, 2024
概括
这项研究引入了一种新的对比集群方法,使用混合用于医学文本. 它通过优化特征空间和减少计算负载来提高聚类精度,特别是在罕见疾病中.
科学领域:
- 计算语言学计算语言学
- 医疗信息学医学信息学
- 机器学习 机器学习
背景情况:
- 传统的文本集群方法面临的挑战是医学数据中的重叠表示.
- 聚类对于从大型未标记的文本数据集中对疾病进行分类至关重要.
- 聚类稀疏的数据,如罕见疾病的数据,存在独特的困难.
研究的目的:
- 提出一种新的对比集群方法,并增强了用于医学文本分析的混合.
- 解决传统集群在处理重叠数据和稀疏的罕见疾病信息方面的局限性.
- 优化功能空间,减少无监督文本集群中的计算负担.
主要方法:
- 一个对比式学习模块通过处理正对和负样本来优化特征空间.
- 混合数据增强产生了具有成本效益的虚拟功能,隐式减少了计算负载.
- 该方法涉及选择小数据批量来模拟罕见疾病实验条件,以进行有效的集群.
主要成果:
- 拟议的方法实现了卓越的实验分数,特别是在小批量数据,超过现有技术.
- 它有效地减轻了数据重叠问题,并提高了稀疏医疗文本的集群准确性.
- 观察到资源使用和时间开支的显著减少.
结论:
- 与混合的对比分类为无监督医疗文本分类提供了一个有利和有效的策略.
- 这种方法表明了前沿的结果,特别有利于分析罕见疾病数据.
- 该方法优化了特征表示,并有效地提高了聚类性能.
更多相关视频
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
16.9K
08:56Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
Published on: January 13, 2023
2.2K
相关概念视频
Self-Discrepancy Theory
18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.
18.3K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Mismatch Repair
4.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
4.8K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Law of Independent Assortment
55.7K
While Mendel’s Law of Segregation states that the two alleles for one gene are separated into different gametes, a different question of how different genes are inherited remains. For example, is the gene for tall plants inherited with the gene for green peas? Mendel asked this question by experimenting with a dihybrid cross; a cross in which both parents are homozygous for two distinct traits resulting in an F1 generation that are heterozygous for both traits.
55.7K
Improving Translational Accuracy
2.6K
2.6K
