Related Experiment Video
Updated: Oct 17, 2025

07:42
A Data-Driven Approach to Quantifying Immune States in Sepsis
Published on: February 7, 2025
343
Efficient Self-Supervised Metric Information Retrieval: A Bibliography Based Method Applied to COVID Literature
Gianluca Moro1, Lorenzo Valgimigli1
1Department of Computer Science and Engineering (DISI), University of Bologna, Via dell'Università 50, I-47521 Cesena, Italy.
Sensors (Basel, Switzerland)
|October 13, 2021
Summary
SUBLIMER is a novel self-supervised information retrieval (IR) engine that efficiently searches scientific papers without needing labeled data. It outperforms existing methods on the COVID-19 Open Research Dataset (CORD19).
Area of Science:
- Information Retrieval
- Computational Linguistics
- Bibliometrics
Background:
- The vast and growing body of scientific literature, particularly on coronaviruses (over 300,000 publications), necessitates efficient methods for knowledge discovery.
- Current information retrieval (IR) systems often rely on deep learning with supervised training, requiring extensive labeled datasets and expert input, which are resource-intensive and slow to create, especially during urgent research periods like a pandemic.
- The COVID-19 Open Research Dataset (CORD19) presents a significant corpus for evaluating IR systems.
Purpose of the Study:
- To develop a novel, self-supervised information retrieval (IR) engine capable of searching scientific paper corpora for relevant documents against arbitrary queries without requiring pre-labeled datasets.
- To evaluate the performance of this new system, named SUBLIMER, against state-of-the-art competitors on the CORD19 dataset.
Main Methods:
- Developed SUBLIMER, a self-supervised IR engine utilizing deep metric learning.
- Trained SUBLIMER on the unsupervised CORD19 dataset.
- Exploited the citation network of papers to create a latent space where proximity indicates semantic similarity, eliminating the need for labeled query-document pairs.
Main Results:
- SUBLIMER demonstrates superior performance compared to state-of-the-art competitors on the CORD19 dataset, specifically in Precision@5 (P@5) and Bpref metrics.
- The self-supervised approach of SUBLIMER significantly outperforms supervised methods that require labeled data.
- SUBLIMER achieves this high performance with an order of magnitude fewer trainable parameters than competing approaches.
Conclusions:
- SUBLIMER offers an efficient and effective self-supervised solution for information retrieval in large scientific literature corpora.
- The methodology, leveraging citation networks for semantic similarity, is adaptable to other scientific domains beyond CORD19.
- This approach addresses the limitations of traditional supervised IR methods, particularly in resource-constrained or time-sensitive research scenarios.
Related Concept Videos
MALDI-TOF Mass Spectrometry
5.8K
Mass spectrometry is a powerful characterization technique that can identify and separate a wide variety of compounds ranging from chemical to biological entities, based on their mass-to-charge ratio (m/z). The instruments that allow this detection, known as mass spectrometers, have three components: an ion source, a mass analyzer, and a detector. These spectrometers differ based on the nature of their ion source and analyzers.
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
5.8K
Single Nucleotide Polymorphisms-SNPs
16.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
16.8K
Statistical Methods for Analyzing Epidemiological Data
612
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
612
Improving Translational Accuracy
3.0K
3.0K

