Related Experiment Video
Updated: Jan 28, 2026

Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Probabilistic variable-length segmentation of protein sequences for discriminative motif discovery (DiMotif) and
Ehsaneddin Asgari1,2, Alice C McHardy2, Mohammad R K Mofrad3,4
1Molecular Cell Biomechanics Laboratory, Departments of Bioengineering and Mechanical Engineering, University of California, Berkeley, CA, 94720, USA.
We introduce peptide-pair encoding (PPE), a novel method for segmenting protein sequences into variable-length sub-sequences. This approach enhances protein motif discovery and sequence embedding for machine learning applications in bioinformatics.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- Protein sequence analysis is crucial for understanding biological function.
- Existing methods for sequence segmentation and feature extraction have limitations.
- Subsequence-based approaches, inspired by natural language processing, offer new possibilities.
Purpose of the Study:
- To introduce peptide-pair encoding (PPE) as a general-purpose probabilistic segmentation method for protein sequences.
- To demonstrate the utility of PPE in protein motif discovery (DiMotif) and protein sequence embedding (ProtVecX).
- To provide a versatile representation for downstream machine learning tasks in protein bioinformatics.
Main Methods:
- Developed a modified byte-pair encoding (BPE) algorithm for probabilistic segmentation of protein sequences into variable-length sub-sequences (PPE).
- Implemented DiMotif, an alignment-free method for discriminative motif discovery using PPE.
- Extended k-mer based protein vector (ProtVec) embedding to variable-length sub-sequences using PPE, creating ProtVecX.
Main Results:
- DiMotif achieved high recall scores in discovering experimentally verified protein motifs and showed strong performance in classification tasks for integrins and biofilm formation.
- ProtVecX demonstrated marginal improvements over ProtVec in enzyme and toxin prediction tasks.
- Combined PPE-derived embeddings with raw amino acid k-mer features improved protein classification performance.
Conclusions:
- Peptide-pair encoding (PPE) provides a flexible and effective representation for protein sequences.
- PPE-based methods like DiMotif and ProtVecX offer valuable tools for motif discovery, sequence embedding, and various machine learning applications in protein bioinformatics.
- The PPE representation can enhance the performance of downstream machine learning tasks by capturing variable-length sub-sequence information.
Related Concept Videos
Cis-regulatory Sequences
Cis-regulatory Sequences
Sanger Sequencing
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...

