Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Self-Discrepancy Theory02:45

Self-Discrepancy Theory

18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.  
18.3K
Cluster Sampling Method01:20

Cluster Sampling Method

11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Mismatch Repair01:20

Mismatch Repair

4.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
4.8K
Multiple Comparison Tests01:13

Multiple Comparison Tests

3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Law of Independent Assortment02:03

Law of Independent Assortment

55.7K
While Mendel’s Law of Segregation states that the two alleles for one gene are separated into different gametes, a different question of how different genes are inherited remains. For example, is the gene for tall plants inherited with the gene for green peas? Mendel asked this question by experimenting with a dihybrid cross; a cross in which both parents are homozygous for two distinct traits resulting in an F1 generation that are heterozygous for both traits.
55.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Pentamethoxyflavanone regulates macrophage polarization and ameliorates sepsis in mice.

Biochemical pharmacology·2014
Same author

Generalized tonic-clonic seizures: aberrant interhemispheric functional and anatomical connectivity.

Radiology·2014
Same author

An unbalanced PD-L1/CD86 ratio in CD14(++)CD16(+) monocytes is correlated with HCV viremia during chronic HCV infection.

Cellular & molecular immunology·2014
Same author

Mussel-inspired polydopamine biopolymer decorated with magnetic nanoparticles for multiple pollutants removal.

Journal of hazardous materials·2014
Same author

Structural variation in Zn(II) coordination polymers built with a semi-rigid tetracarboxylate and different pyridine linkers: synthesis and selective CO2 adsorption studies.

Dalton transactions (Cambridge, England : 2003)·2014
Same author

High expression of sarcoplasmic/endoplasmic reticulum Ca(2+)-ATPase 2b blocks cell differentiation in human liposarcoma cells.

Life sciences·2014

Related Experiment Video

Updated: Jul 3, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K

A lightweight mixup-based short texts clustering for contrastive learning.

Qiang Xu1, HaiBo Zan1, ShengWei Ji1

  • 1School of Artificial Intelligence and Big Data, Hefei University, Hefei, Anhui, China.

Frontiers in Computational Neuroscience
|February 13, 2024
PubMed
Summary

This study introduces a novel contrastive clustering method using mixup for medical text. It improves clustering accuracy, especially for rare diseases, by optimizing feature spaces and reducing computational load.

Keywords:
contrastive learningdata augmentationmixupoverlappingtext clustering

More Related Videos

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
12:49

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition

Published on: July 13, 2019

16.9K
Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
08:56

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates

Published on: January 13, 2023

2.2K

Related Experiment Videos

Last Updated: Jul 3, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K
Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
12:49

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition

Published on: July 13, 2019

16.9K
Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
08:56

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates

Published on: January 13, 2023

2.2K

Area of Science:

  • Computational linguistics
  • Medical informatics
  • Machine learning

Background:

  • Traditional text clustering methods face challenges with overlapping representations in medical data.
  • Clustering is crucial for categorizing medical conditions from large unlabeled text datasets.
  • Clustering sparse data, like that for rare diseases, presents unique difficulties.

Purpose of the Study:

  • To propose a novel contrastive clustering method enhanced with mixup for medical text analysis.
  • To address the limitations of traditional clustering in handling overlapping data and sparse rare disease information.
  • To optimize feature spaces and reduce computational burden in unsupervised text clustering.

Main Methods:

  • A contrastive learning module optimizes feature spaces by treating positive pairs and negative samples.
  • Mixup data augmentation generates cost-effective virtual features, implicitly reducing computational load.
  • The method involves selecting small data batches to simulate rare disease experimental conditions for effective clustering.

Main Results:

  • The proposed method achieves superior experiment scores, particularly with small batch data, outperforming existing techniques.
  • It effectively mitigates data overlap issues and enhances clustering accuracy for sparse medical text.
  • Significant reductions in resource usage and time overhead were observed.

Conclusions:

  • The contrastive clustering with mixup offers a favorable and effective strategy for unsupervised medical text clustering.
  • This approach demonstrates cutting-edge outcomes, particularly beneficial for analyzing rare disease data.
  • The method optimizes feature representation and enhances clustering performance efficiently.