Related Experiment Video
Updated: Feb 24, 2026

Author Spotlight: Integrating 2D-HPLC-MS and Molecular Networking in Natural Medicine Analysis
Published on: December 8, 2023
LDComKG: an LLM-powered dual-enhanced framework for community-aware knowledge graph completion in traditional Chinese
Xing Zeng1, Ziyan Wang1,2, Jingxian Chai1
1School of Artificial Intelligence and Information Technology, Nanjing University of Chinese Medicine, Nanjing, 210023 China.
This study introduces LDComKG, an LLM-powered framework for building knowledge graphs from ancient Chinese medical texts. It significantly improves accuracy in knowledge graph construction and completion, unlocking valuable historical medical information.
Area of Science:
- Bioinformatics and Health Informatics.
- Natural Language Processing (NLP) in Traditional Chinese Medicine (TCM).
- TCM knowledge graph completion and graph representation learning.
Background:
Ancient Chinese medical corpora contain intricate semantics and diverse terminology that resist standard digitization efforts due to their non-linear linguistic structures. Prior research has shown that rule-based methodologies exhibit limited efficacy when processing these non-standardized historical texts which often lack modern punctuation or consistent syntax. The absence of standardized linguistic structures in classical scripts complicates the extraction of meaningful entities and relationships required for robust digital modeling. Existing computational frameworks often struggle to maintain the semantic integrity of traditional medical wisdom during graph construction because they cannot interpret archaic context. Researchers have sought more robust ways to bridge the gap between unstructured ancient data and structured digital repositories to preserve cultural heritage. Traditional methods frequently fail to capture the nuanced relationships inherent in historical healthcare documents, leading to fragmented or inaccurate knowledge representations. This absence of evidence motivated the exploration of advanced neural architectures to unlock knowledge within complex historical medical documents through automated semantic processing.
Purpose Of The Study:
This research introduces LDComKG, a dual-enhanced framework designed to automate the construction and completion of knowledge graphs for Traditional Chinese Medicine (TCM) by leveraging neural language understanding. The system addresses the critical challenge of extracting structured data from unstructured raw texts found in ancient Chinese medical corpora through a specialized auto-generation module. By utilizing Large Language Models (LLMs), the investigators seek to improve the efficiency of graph data auto-generation while maintaining high levels of semantic accuracy. The framework aims to integrate structural analysis with semantic comprehension to refine link prediction tasks within the resulting knowledge network. The study evaluates how community-aware algorithms can enhance the recognition of complex relationships within medical networks to provide a more holistic view of TCM. Developers intended to create a scalable solution for digitizing and utilizing traditional medical knowledge in modern information systems to support clinical decision-making. This effort focuses on overcoming the obstacles presented by diverse terminology and lack of standardization in historical texts to ensure data reliability.
Main Methods:
The investigators developed a graph data auto-generation module that employs Large Language Models (LLMs) to process hierarchical text chunks for entity and relation extraction. This module extracts entities and relationships from classical Chinese architectures to form the initial graph structure without requiring extensive manual annotation. The team integrated the Leiden algorithm to identify and enhance community structure recognition within the network, allowing for better clustering of related medical concepts. Graph Sample and Aggregated (GraphSAGE) was utilized for effective node representation learning to facilitate downstream tasks like link prediction and graph completion. Researchers performed extensive experiments using annotated ancient medical text datasets to validate the LDComKG framework against established benchmarks in the field. Comparative evaluations were conducted against state-of-the-art fusion methods, including Knowledge Embedding and Pre-trained Language Representation (KEPLER) and Contextualized Language and Knowledge Embedding (CoLAKE). Systematic assessments of different backbones, such as Generative Pre-trained Transformer 4o (GPT-4o) and GPT-4o-mini, were executed across various Graph Retrieval-Augmented Generation (GraphRAG) stages to ensure robustness.
Main Results:
The LDComKG framework achieved an Accuracy of 95.94% and an F1-Score of 95.07% when utilizing the Generative Pre-trained Transformer 4o (GPT-4o) backbone during the final evaluation phase. Statistical analysis revealed an Area Under the Curve (AUC) of 97.72% for the GPT-4o model, indicating high precision in link prediction across the medical dataset. When using the Generative Pre-trained Transformer 4o (GPT-4o) mini variant, the system maintained strong performance with an Accuracy of 93.33% and an F1-Score of 93.69% despite the reduced parameter count. The AUC for the mini architecture reached 95.87%, illustrating the robustness of the Large Language Model (LLM) powered dual-enhanced framework across different computational scales and model sizes. Experimental data confirmed that the synergistic combination of Large Language Models (LLMs) and graph algorithms surpasses traditional baseline models in both entity recognition and relationship mapping. Results highlight that the auto-generation module significantly reduces manual effort while improving the quality of the initial graph by capturing complex semantic links. The framework exhibited superior performance in navigating intricate semantics compared to existing state-of-the-art fusion methods like Knowledge Embedding and Pre-trained Language Representation (KEPLER) and Contextualized Language and Knowledge Embedding (CoLAKE) in all tested scenarios.
Conclusions:
The integration of advanced Large Language Models (LLMs) with graph representation learning offers a transformative approach for processing semantically complex medical corpora from ancient eras. By harnessing the ability of these models to navigate intricate terminology, the LDComKG framework provides a scalable solution for Traditional Chinese Medicine (TCM) knowledge graph completion in modern research. The researchers conclude that this methodology addresses the critical challenge of digitizing and utilizing traditional medical knowledge effectively for future generations of practitioners. Findings suggest that community-aware frameworks can significantly improve the accuracy of intelligent information systems in healthcare by identifying hidden patterns in medical data. This work serves as a practical application of LLMs in healthcare knowledge discovery, specifically for ancient historical texts that were previously inaccessible to automated tools. Future developments may leverage these insights to build high-quality digital repositories for diverse traditional medical traditions across different cultures and languages. The study validates the potential of LLM-empowered solutions to preserve and utilize historical medical wisdom in modern digital contexts while maintaining scientific rigor.
Frequently Asked Questions
The framework utilizes a graph data auto-generation module powered by Large Language Models (LLMs) to process hierarchical text chunks. This approach allows the system to intelligently identify entities and relationships within unstructured classical Chinese corpora, significantly reducing manual effort while enhancing the quality of the initial graph.
According to the study's findings, the framework utilizing Generative Pre-trained Transformer 4o (GPT-4o) achieved an Accuracy of 95.94%, an F1-Score of 95.07%, and an Area Under the Curve (AUC) of 97.72%. These values represent a significant improvement over baseline models and the smaller GPT-4o-mini variant in link prediction tasks.
The researchers integrated the Leiden algorithm to enhance community structure recognition within the knowledge graph. By identifying these clusters, the framework improves the accuracy of node representation learning and link prediction, specifically addressing the complex semantic relationships found in traditional Chinese medical texts.
The framework is specifically designed for ancient Chinese medical corpora characterized by intricate semantics and diverse terminology. Its effectiveness is demonstrated on these challenging historical texts, and the authors focus on its application for digitizing and utilizing traditional medical knowledge within intelligent information systems.
The study's authors propose that integrating advanced Large Language Models (LLMs) with graph representation learning offers a scalable solution for building high-quality Traditional Chinese Medicine (TCM) knowledge graphs. They state that this approach has transformative potential for healthcare knowledge discovery.
More Related Videos
07:16Preparation of Gynura bicolor DC samples for High-Resolution Tandem Mass Spectrometry
Published on: February 2, 2024
13:18Network Pharmacology Prediction and Experimental Validation of Trichosanthes-Fritillaria thunbergii Action Mechanism Against Lung Adenocarcinoma
Published on: March 3, 2023