Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Predicting Products: Substitution vs. Elimination02:52

Predicting Products: Substitution vs. Elimination

11.6K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
11.6K
Prediction Intervals01:03

Prediction Intervals

2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.3K
Cluster Sampling Method01:20

Cluster Sampling Method

11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Predicting Products: SN1 vs. SN202:27

Predicting Products: SN1 vs. SN2

13.3K
Nucleophilic substitution reactions of alkyl halides can proceed via an SN1 or an SN2 mechanism. While in SN2 reactions, the nucleophile attacks the substrate simultaneously as the leaving group departs, in SN1 reactions, the substrate first dissociates to give the carbocation intermediate. Various factors such as the structure of the substrate, the strength of the nucleophile, and the nature of the solvent promote one mechanism over the other.
With increased substitution on the alkyl halide,...
13.3K
Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K
Upsampling01:22

Upsampling

229
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
229

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Attribution-based interpretable classification neural network with global and local perspectives.

Scientific reports·2025
Same author

Non-redundant implicational base of formal context with constraints using SAT.

PeerJ. Computer science·2024
Same author

FFP: joint Fast Fourier transform and fractal dimension in amino acid property-aware phylogenetic analysis.

BMC bioinformatics·2022
Same author

Multichannel Two-Dimensional Convolutional Neural Network Based on Interactive Features and Group Strategy for Chinese Sentiment Analysis.

Sensors (Basel, Switzerland)·2022

Related Experiment Video

Updated: Jun 25, 2025

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

389

Clustering swap prediction for image-text pre-training.

Sun Fayou1,2,3, Hea Choon Ngo4, Yong Wee Sek4

  • 1Guangxi University, Nanning, 530004, Guangxi, China. 314565679@qq.com.

Scientific Reports
|May 24, 2024
PubMed
Summary

This study introduces Clus, a novel multimodal pre-training approach using clustering swap prediction for improved image-text understanding. Clus achieves state-of-the-art results on various downstream tasks, enhancing representation learning.

Keywords:
Cluster numberClustering learningModel pre-trainingSwap prediction

More Related Videos

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.8K
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

2.4K

Related Experiment Videos

Last Updated: Jun 25, 2025

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

389
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.8K
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

2.4K

Area of Science:

  • Computer Science
  • Artificial Intelligence
  • Machine Learning

Background:

  • Multimodal model pre-training significantly impacts downstream task performance.
  • Clustering learning offers benefits but faces challenges with open image-text pairs.
  • Existing methods struggle with the scale and openness of web-scale alt-text data.

Purpose of the Study:

  • To propose a novel approach for learning image-text clustering embedding space.
  • To address the challenges of multimodal clustering with open image-text pairs.
  • To develop a method that allows for an open number of clusters for web-scale data.

Main Methods:

  • Introduced a clustering swap prediction strategy for interaction prediction between image and text features.
  • Employed a distillation learning approach for efficient training of image and text encoders.
  • Pre-trained the model end-to-end using large-scale image-text pairs with text and image as ground truth for swap prediction.

Main Results:

  • Achieved state-of-the-art performance on multiple downstream fine-tuning and zero-shot tasks.
  • Demonstrated effective representation learning through swap prediction.
  • Evaluated the image-encoder's performance on downstream visual tasks.

Conclusions:

  • The proposed Clus method effectively learns image-text clustering embeddings.
  • Clus offers a scalable solution for multimodal clustering with open-world data.
  • The approach shows strong generalization capabilities across diverse vision-language tasks.