Related Experiment Video
Updated: Jun 11, 2025

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Dataset from a human-in-the-loop approach to identify functionally important protein residues from literature
Melanie Vollmar1, Santosh Tirunagari2, Deborah Harrus3
1Protein Data Bank in Europe, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge, CB10 1SD, UK. melaniev@ebi.ac.uk.
This study introduces a novel system combining human curators and machine learning to create a residue-level protein annotation dataset. The approach achieved high precision, recall, and F1-measure, enhancing protein research.
Area of Science:
- Biochemistry
- Bioinformatics
- Computational Biology
Background:
- Accurate residue-level functional annotations and protein structure features are crucial for understanding protein function.
- Extracting this information from scientific literature is challenging due to the complexity and volume of data.
Purpose of the Study:
- To develop a novel system for creating a dataset and model for residue-level protein structure feature and functional annotation detection.
- To integrate human expertise with machine learning for improved data curation.
Main Methods:
- A human-in-the-loop annotation system was employed, integrating data from PDBe, EuropePMC, PubMedCentral, and PubMed.
- Annotation guidelines from UniProt and tools like LitSuggest and HuggingFace models were utilized.
- Seven annotators manually curated ten articles, training a PubmedBert model.
Main Results:
- The iterative human-in-the-loop system achieved high performance metrics: 0.90 precision, 0.92 recall, and 0.91 F1-measure.
- Demonstrated a successful synergy between machine learning and human expertise.
- Successfully curated a dataset for residue-level functional annotations and protein structure features.
Conclusions:
- The proposed system effectively bridges the gap between advanced machine learning models and domain expert insights.
- This approach holds significant potential for broader applications in protein research.
- Highlights the value of integrating human curation in developing robust biological datasets and models.
Related Concept Videos
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Ligand Binding and Linkage

