Related Experiment Videos
Omics Data Discovery Agents: Agent-Supported Retrieval, Reanalysis, and Synthesis of Published Omics Data
Alexandre Hutton1, Jesse G Meyer1
1Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, USA.
Abstract:
The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for reproducing reported results is dispersed across main text, supplementary files, and code repositories, and in the rarer cases where intermediate data (e.g. protein abundance files) are shared, their location is irregular. Here we present an agentic framework for the agent-supported retrieval, reanalysis, and synthesis of published omics data. The system employs large language model (LLM) agents with access to tools for fetching omics studies, extracting article metadata, identifying and downloading published data, executing containerized quantification pipelines, and synthesizing results across studies. Applied at corpus scale, the pipeline catalogued dataset references across thousands of PubMed Central articles; we report these as descriptive system outputs rather than as a validated measure of extraction accuracy. Using model context protocol (MCP) servers to expose containerized analysis tools, the agents retrieved and re-quantified data in five end-to-end reanalyses spanning data-dependent and data-independent proteomics and bulk RNA-seq. All five reanalyses completed, each with documented human guidance and workflow accommodations, and reproduced the authors' deposited abundances with high per-sample correlation (0.85-0.997) and strongly concordant differentially expressed features (fold-change Spearman 0.88-0.91), with no direction reversals among features called differentially expressed in both analyses; residual differences in significant-feature lists were attributable to threshold placement, tool-version, and preprocessing differences rather than to the underlying quantities. We further demonstrate that agents can identify semantically similar studies, judge data compatibility, and synthesize findings across studies, including a random-effects meta-analysis that recovered consistent protein regulation in liver fibrosis. Rather than a validated benchmark of literature-wide performance, this work is a feasibility demonstration together with an auditable, reusable toolset, establishing a foundation for prospective evaluation of automated omics-data reuse.
Related Concept Videos
Genomics
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Evolutionary Relationships through Genome Comparisons
DNA Microarrays