Related Experiment Video
Updated: Sep 3, 2026

Multiplex Immunohistochemical Analysis of the Spatial Immune Cell Landscape of the Tumor Microenvironment
Published on: August 18, 2023
LLMs build inconsistent knowledge graphs from one immune-checkpoint corpus: a five-axis benchmark and consensus
Qiang Xie1, Yang Fang2, Yanan Hu3
1Department of Surgical Oncology, The First Affiliated Hospital of Bengbu Medical University, Bengbu, Anhui, China.
Abstract:
Large language models (LLMs) are increasingly used to mine biomedical literature into knowledge graphs (KGs), yet whether independent models agree when given the same text is largely untested. We passed 9,978 immune-checkpoint PubMed abstracts through four LLMs (DeepSeek, GLM, Kimi, and Qwen) under an identical prompt and an eight-entity-type schema. This yielded 140,462 triples and a 95,981-edge union graph. We benchmarked the four graphs along five axes: extraction, correctness against a human expert, topology, embedding-based link prediction, and external validation. Extraction volume varied more than twofold (20,122-45,957 triples), and pairwise edge overlap was low (Jaccard, 0.06-0.11). Against a 151-triple dual-annotated gold set, expert precision was 0.914 (95% CI = 0.858-0.949) with overlapping per-model intervals, while an LLM judge systematically under-rated precision on the 51 triples it scored (raw agreement, 0.65). External support rose with cross-model consensus, but naive majority agreement did not beat the best single model (DeepSeek, 0.377); only unanimous four-model agreement exceeded it (0.423, + 12% relative) at ~ 4% edge coverage. The unanimity interval (95% CI = 0.362-0.486) overlaps that of the best single model (0.358-0.397), so unanimity is best read as a conservative reliability filter rather than a general fix for graph quality. Consensus is a tunable precision-coverage control, not a free-lunch gain; under the conditions tested here-four models, one domain, one prompt, and schema-the choice of extraction model mattered more than naive aggregation.

