Related Experiment Video
Updated: May 12, 2026

10:41
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved (Non-model) Organisms
Published on: May 9, 2017
Alembic: a framework for converting disparate biological data into structured resources
I V Bezdvornykh1, K I Yuditskiy1, N A Cherkasov1
1Institute for Translational Biomedicine, Saint Petersburg State University, St. Petersburg, Russia.
Vavilovskii Zhurnal Genetiki I Selektsii
|May 11, 2026
Summary
Re-analyzing public sequencing data is crucial but hindered by inconsistent metadata. Alembic, a new software package, uses advanced AI and Natural Language Processing to structure this data, enabling efficient searching and analysis of biological datasets.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Public sequencing data repositories contain vast biological information.
- Data heterogeneity and non-standardized metadata impede effective data retrieval and analysis.
- Existing tools lack integrated workflows for accessing and processing diverse sequencing datasets.
Purpose of the Study:
- To develop a comprehensive solution for overcoming data heterogeneity in public sequencing repositories.
- To enable systematic search, integration, and comparative analysis of sequencing data.
- To streamline the workflow from data discovery to local repository construction.
Main Methods:
- Utilized the Entrez database system and its API for programmatic access to sequencing data and metadata.
- Applied Natural Language Processing (NLP) techniques, specifically transformer-based AI algorithms (PubMedBERT via AIONER platform), to analyze biomedical text.
- Developed the Alembic software package, a client-server system for automated data structuring and analysis.
Main Results:
- Transformed unstructured textual metadata into structured, computable information.
- Enabled efficient keyword-based searching and identification of relevant datasets, including gene-specific searches.
- Provided a curated list of datasets with structured metadata and direct links for downloading primary sequencing files.
Conclusions:
- Alembic offers a universal solution to the fragmented approach of existing tools for public sequencing data.
- The software facilitates efficient identification of high-value targets and construction of tailored local repositories.
- Alembic enhances the accessibility and utility of accumulated public sequencing data for biological research.
Related Concept Videos
Applications of Molecular Taxonomy
Molecular taxonomy has revolutionized the understanding and classification of bacteria, providing precise insights into their diversity, evolutionary relationships, and ecological roles. By utilizing molecular techniques such as DNA sequencing and fingerprinting, researchers have made significant strides in various fields related to bacterial studies.Resolving Taxonomic AmbiguitiesMolecular taxonomy has been instrumental in distinguishing closely related bacterial species initially thought to...
Synthetic Biology
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.

