Related Experiment Video
Updated: May 13, 2026

13:01
Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment
Published on: June 3, 2022
scHilda: Hierarchical Integration of LLM with KG database for single cell type annotation
Yilang Li1, Yidi Sun2, Aoyun Geng2
1School of Cyberspace Security (School of Cryptology), Hainan University, Haikou, China.
Plos Computational Biology
|May 11, 2026
Summary
scHilda enhances single-cell RNA sequencing annotation by integrating knowledge graphs with Large Language Models (LLMs). This framework reduces LLM hallucinations and improves cell type identification accuracy and interpretability.
Area of Science:
- Computational Biology
- Genomics
- Artificial Intelligence
Background:
- Single-cell RNA sequencing (scRNA-seq) cell type annotation is crucial but faces accuracy and generalization challenges.
- Existing automated methods struggle with novel cell types and LLM hallucinations.
- Integrating external knowledge bases with LLMs is difficult due to data imperfections.
Purpose of the Study:
- To develop a novel framework, scHilda, for robust and accurate cell type annotation in scRNA-seq data.
- To address LLM hallucinations and knowledge base deficiencies in biological data analysis.
- To improve the interpretability and generalization of automated cell annotation.
Main Methods:
- scHilda integrates a Knowledge Graph (KG) with Large Language Models (LLMs).
- A hierarchical arbitration strategy identifies major cell lineages using a global KG, then refines subtypes with focused subgraph information.
- Dynamic, knowledge-enhanced reasoning constrains LLM decision-making, mitigating hallucinations and knowledge base errors.
Main Results:
- scHilda achieves state-of-the-art performance on multiple benchmark datasets, outperforming existing methods.
- The framework demonstrates robustness in complex mixed samples and enables lightweight LLMs to approach top-tier performance.
- Statistical evaluations and interpretability case studies confirm scHilda's efficiency and transparent decision-making.
Conclusions:
- scHilda offers a new paradigm for trustworthy biological AI by integrating LLMs with structured biological knowledge.
- The framework significantly improves accuracy, interpretability, and robustness in scRNA-seq cell annotation.
- scHilda effectively constrains LLM reasoning, reducing hallucinations and enhancing biological insights.
