Related Experiment Videos
The Proteins API: accessing key integrated protein and genome information
Andrew Nightingale1, Ricardo Antunes1, Emanuele Alpi1
1EMBL-EBI, Wellcome Genome Campus, Hinxton, Cambridgeshire CB10 1SD, UK.
Nucleic Acids Research
|April 7, 2017
Summary
The Proteins API offers researchers easy access to protein and genomics data, including annotations and variations. It provides a user-friendly interface and programmatic access for biological data exploration.
Area of Science:
- Bioinformatics
- Proteomics
- Genomics
Background:
- Accessing integrated protein and genomics data is crucial for biological research.
- Existing data sources are often fragmented, requiring complex queries.
- Programmatic access to curated annotations and large-scale data is needed.
Purpose of the Study:
- To introduce the Proteins API for searching and accessing protein and associated genomics data.
- To provide programmatic access to UniProtKB annotations and large-scale data (LSS).
- To offer a user-friendly interface for querying and retrieving biological data.
Main Methods:
- Development of a RESTful API for protein and genomics data.
- Implementation of a coordinates service for retrieving genomic sequence coordinates.
- Integration of data from UniProtKB and large-scale data sources (LSS).
- Provision of a Swagger UI for interactive querying and code generation.
Main Results:
- The Proteins API enables searching and programmatic access to curated protein sequence annotations, variations, and proteomics data.
- Researchers can retrieve genomic sequence coordinates for UniProtKB proteins.
- Dynamically generated source code in multiple programming languages is available.
- Search results are returned in standard formats like JSON, XML, and GFF.
Conclusions:
- The Proteins API serves as a scalable, reliable, and fast resource for protein information.
- It facilitates an integrated overview of protein annotations to aid knowledge gain in biological processes.
- The API empowers researchers with diverse expertise to explore protein data effectively.