Related Experiment Video
Updated: Jun 14, 2026

DeepOmicsAE: Representing Signaling Modules in Alzheimer's Disease with Deep Learning Analysis of Proteomics, Metabolomics, and Clinical Data
Published on: December 15, 2023
PEPE: scalable extraction of multi-modal protein language model representations
Jahn Zhong1, Niccolò Cardente1, Geir Kjetil Sandve2,3
1Department of Immunology, University of Oslo, Oslo, Norway.
Protein language models (PLMs) generate valuable embeddings, but current methods are inefficient. PEPE (Parallel Extraction for Protein Embeddings) offers a faster, memory-efficient solution for generating protein embeddings.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning in Biology
Background:
- Protein language models (PLMs) capture complex amino-acid relationships, yielding embeddings rich in structural, functional, and evolutionary data.
- Existing workflows for extracting these embeddings often involve arbitrary choices, leading to suboptimal representations and hindering downstream analyses.
- Large-scale embedding generation faces computational and memory bottlenecks due to inefficient accumulation and redundant computations.
Purpose of the Study:
- To introduce PEPE (Parallel Extraction for Protein Embeddings), a tool designed for efficient, high-throughput, and multimodal extraction of protein embeddings.
- To overcome the limitations of current methods in terms of speed, memory consumption, and scalability for generating protein representations.
- To enable researchers to efficiently generate large, informative embedding datasets for discovering optimal representations for various biological tasks.
Main Methods:
- PEPE utilizes a parallelized and streaming-based architecture to significantly accelerate embedding extraction.
- The tool maintains stable, low memory consumption, overcoming the linear scaling issues of conventional methods.
- PEPE offers a flexible interface supporting a wide range of state-of-the-art and custom PLMs.
Main Results:
- PEPE achieves runtimes several orders of magnitude faster than sequential embedding extraction approaches.
- The tool enables multimodal embedding extraction with minimal memory footprint, even exceeding available RAM.
- PEPE streamlines the generation of diverse embedding configurations, facilitating the identification of high-performing latent states.
Conclusions:
- PEPE provides a scalable, robust, and user-friendly solution for efficient protein embedding generation.
- The tool empowers researchers to create massive, information-rich embedding datasets for downstream structural, functional, and evolutionary analyses.
- By optimizing embedding generation, PEPE facilitates the discovery of context-specific latent states without additional computational cost.
Related Concept Videos
Multi-pass Transmembrane Proteins and β-barrels
α-Helix containing multi-pass transmembrane proteins
Multi-pass transmembrane proteins such as G-protein-linked receptors (GPCRs) and...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Complex Assembly
Many viruses self-assemble into a fully functional unit using the infected host cell to...