Related Experiment Video
Updated: Sep 14, 2026

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
Evaluating transformer-based models for structural characterization of orphan proteins
Ercan Seçkin1,2, Dominique Colinet1, Etienne G J Danchin1
1Institut Sophia Agrobiotech, INRAE, Université Côte d'Azur, CNRS, Sophia-Antipolis 06903, France.
Motivation:
Transformer-based models (TBMs) are state-of-the-art deep learning architectures that predict protein structural features with high accuracy. Despite methodological differences, they all rely on large datasets structured in families of homologous sequences. However, 5%-30% of eukaryotic proteomes consist of orphan proteins, which are sequences without detectable similarity to known families. Although they may share structural traits with characterized proteins, their lack of homology makes them an ideal dataset for evaluating TBM generalization beyond familiar sequence space.
Results:
We compared predictions from several widely used TBM architectures on an expert-curated set of orphan proteins from the Meloidogyne genus, comprising some of the most destructive plant-parasitic nematodes. Multiple sequence alignment-based approaches such as AlphaFold2 performed poorly on orphan proteins, as did single-sequence or embedding-based language models ESMFold, OmegaFold, and ProtT5. This limited performance cannot be fully attributed to intrinsic disorder, as confirmed by independent non-TBM disorder predictors. While accurate tertiary structure prediction remains out of reach, secondary structure is more reliably captured: predictors share about 70% of secondary structure elements, regardless of global fold similarity, and these elements are consistently identified by dedicated secondary structure tools.
Availability:
All data and analysis scripts are available at https://doi.org/10.5281/zenodo.18788931.

