Related Experiment Video
Updated: May 9, 2025

10:41
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
Published on: May 9, 2017
9.2K
Contrastive learning and mixture of experts enables precise vector embeddings in biological databases.
Logan Hallee1, Rohan Kapur2, Arjun Patel3
1Center for Bioinformatics and Computational Biology, University of Delaware, Newark, USA.
Scientific Reports
|April 29, 2025
Summary
This study enhances scientific text vector embeddings using a novel Mixture of Experts (MoE) approach on BERT models. The method improves representation learning for biomedical documents, boosting search and compilation efficiency.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Machine Learning
Background:
- Transformer neural networks excel at sentence similarity but struggle with complex scientific documents.
- Effective vector representations are crucial for retrieval augmentation and searching diverse research literature.
- Existing models generate sub-optimal embeddings for structurally and thematically varied scientific texts.
Purpose of the Study:
- To improve vector embeddings for scientific text, particularly in biomedical domains.
- To develop a novel Mixture of Experts (MoE) extension pipeline for pretrained BERT models.
- To enhance the representation learning of heterogeneous biomedical inputs for efficient search and compilation.
Main Methods:
- Assembled domain-specific datasets using co-citations as a similarity metric in biomedical domains.
- Introduced a novel Mixture of Experts (MoE) extension pipeline applied to pretrained BERT models.
- Trained MoE variants to classify co-cited publications based on scientific abstracts, utilizing a unique routing scheme.
Main Results:
- The Mixture of Experts (MoE) variants demonstrated improved vector embeddings for scientific text.
- The unique routing scheme ensured the MoE system maintained the same throughput as regular transformers.
- The methodology shows promise for encoding heterogeneous biomedical inputs efficiently.
Conclusions:
- The proposed MoE extension pipeline offers a versatile and efficient solution for encoding diverse biomedical documents.
- Advancements in representation learning can significantly enhance vector database search and compilation for scientific literature.
- This approach paves the way for One-Size-Fits-All transformer networks in scientific information retrieval.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
5.6K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.6K
Convergent Evolution
27.1K
Evolution shapes the features of organisms over time, ensuring that they are suited for the environments in which they live. Sometimes, selection pressure leads to the rise of similar but unrelated adaptations in organisms with no recent common ancestors, a process known as convergent evolution.
27.1K

