Related Experiment Video
Updated: Oct 4, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Evaluation of protein sequence retrieval via protein language model embeddings and FAISS
Can Wu1, Haoran Zhang1, Liqiong Chen1
1Shanghai Institute of Technology, Shanghai, China.
Abstract:
Protein sequence similarity search underpins function annotation, evolutionary analysis, and metagenomic mining, yet conventional alignment-based methods face two persistent difficulties. First, their computational cost grows linearly with database size, making them impractical for billion-scale repositories. Second, their sensitivity declines below 30% sequence identity (the "twilight zone"), limiting remote homolog detection. Protein language models (PLMs) produce evolutionarily informative embeddings through self-supervised learning; however, the systematic integration of PLM embeddings with vector search into online retrieval services, together with experimental characterisation across model variants and index types, has received limited attention. We present ProtFaiss, combining ESM2 embeddings with FAISS approximate nearest neighbour search. On SCOPe40, we evaluate classification accuracy, ranking quality, remote homology sensitivity, retrieval speed, and storage footprint. Ablations spanning index type, model capacity, database scale, pooling strategy, and layer selection characterise the design space. ProtFaiss vector search completes in ∼0.02ms per query - roughly 3×105-fold faster than a BLASTp database scan - and, including ESM2 encoding, the end-to-end latency is ∼30ms, approximately 200-fold faster than BLASTp. ProtFaiss achieves remote homology sensitivity of 7.7% (BLASTp: 6.7%), concentrated in the 15%-25% identity range. A notable finding is that intermediate Transformer layers substantially outperform the final layer: ESM2-650M layer 22 achieves Recall@10 of 0.780, compared with 0.431 for the final layer, suggesting later layers trade global semantic signal for local structural detail. A learned attention-pooling analysis and a preliminary Flat-index evaluation on multi-domain Swiss-Prot proteins further qualify this interpretation: learned pooling partially recovers final-layer performance, whereas intermediate layers remain advantageous in both evaluations. These results support the feasibility of PLM-based retrieval for large-scale protein search; however, generalisability is constrained by the single-domain benchmark.
More Related Videos
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence.
Protein and Protein Structures
Protein-protein Interfaces
Protein-Protein Interfaces
Protein Families
Protein Families

