Related Experiment Video
Updated: Sep 11, 2025

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
BPA: a BERT-based priority annotation strategy for assessing the rationality of aquatic algal protein sequences
Rui-Hua Huang1, Jun-Ze Liang1, Zheng-Hua Sun1
1MOE Key Laboratory of Tumor Molecular Biology and State Key Laboratory of Bioactive Molecules and Druggability Assessment, Guangdong Basic Research Center of Excellence for Natural Bioactive Molecules and Discovery of Innovative Drugs, College of Life Science and Technology, Jinan University, No. 601, West Huangpu Avenue, Guangzhou 510632, China.
Abstract:
Database searching remains the main approach for mass spectrometry-based proteomics, where protein identification fundamentally requires prior inclusion in the reference database. For aquatic algal species lacking annotated genomes, six-frame translation of species-specific transcriptomes has emerged as a prevalent method. However, this approach results in databases that encompass all potential translation products, substantially increasing the database size and search space. Here, we introduce BERT-based Protein Annotation (BPA), a deep learning strategy that combines a pretrained BERT model for contextual patterns, Pseudo Amino Acid Composition for physicochemical properties, and InterProScan for functional domain prediction, to optimize reference proteome construction. These features are integrated by using a Random Forest classifier to generate dynamic Sequence Reliability Scores, enabling adaptive filtering thresholds tailored to diverse experimental designs. Based on the validation across three distinct test species, this study demonstrates a robust performance of BPA with sustained high classification accuracy (AUC > 0.95). In the application to Karenia mikimotoi, BPA achieved 90% proteome compression while maintaining 40% identification coverage, effectively resolving the peptide ambiguity from redundant translations. This framework provides a scalable and efficient solution for constructing and optimizing reference libraries, facilitating proteomic research in aquatic algae and other genomically understudied species. Source code and executables are available at (https://github.com/huangruihua/BPA.git).

