Related Experiment Video
Updated: Sep 9, 2026

In Vivo Modeling of the Morbid Human Genome using Danio rerio
Published on: August 24, 2013
[Exploring the feasibility and limitations of using large language model to interpret bibliometric findings: a case
1State Key Laboratory of Experimental Hematology, National Clinical Research Center for Blood Diseases, Haihe Laboratory of Cell Ecosystem, Institute of Hematology & Blood Diseases Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, Tianjin 300020, China Tianjin Institutes of Health Science, Tianjin 301600, China.
Abstract:
Objective: To evaluate the feasibility and limitations of using a large language model to assist in the interpretation of bibliometric findings. Methods: This study retrieved 1599 English-language articles on hemophilia gene therapy published between 1970 and 2024 from the Web of Science Core Collection, comprising 1019 articles from the United States, 174 from China, and 172 from the United Kingdom. VOSviewer was used for keyword co-occurrence and clustering analyses of the global and country-specific datasets, whereas CiteSpace was used for keyword co-occurrence, clustering, and burst-detection analyses of the global dataset. Structured data exported from the two tools were interpreted with DeepSeek-R1-0528 to summarize research hotspots and their temporal evolution. Two domain experts independently interpreted the global data and reviewed the country-specific interpretations generated using DeepSeek. Additional literature searches on CRISPR/Cas9, stem-cell-based approaches, and in utero gene therapy were conducted to cross-validate the national research characteristics inferred by the model. Results: DeepSeek summarized global hemophilia gene therapy research into four major themes: the design and in vivo expression regulation of gene therapy vectors; Adeno-associated virus vector delivery strategies and breakthroughs in immunogenicity; the clinical translation and real-world application of gene therapy; the mechanisms of immune tolerance induction and inhibitor formation. Based on seven major CiteSpace clusters and the temporal and burst information of the keywords, DeepSeek further outlined four stages of development: early exploration of vectors and animal models, optimization of delivery and therapeutic strategies, clinical validation, and increasing attention to clinical application and patient benefits. The core research themes and overall developmental trajectory for the global dataset identified using DeepSeek were broadly consistent with expert interpretations, and the model rapidly generated well-structured summaries. However, expert calibration remained necessary for professional terminology, stage delineation, and historical milestone recognition. In country-specific datasets, DeepSeek demonstrated risks of conceptual conflation, overgeneralization of emerging research directions, and inappropriate cross-country comparisons according to within-country keyword frequencies or cluster strengths. Additional searches revealed that the relative prominence or absence of a keyword within a national dataset could not be directly interpreted as international leadership or a lack of research activity. Conclusions: Large language models can assist in interpreting global bibliometric findings and improve the efficiency of identifying research hotspots and outlining developmental trajectories; however, they cannot replace domain experts. Interpretations of emerging research directions and cross-country differences require expert review and calibration using globally comparable data obtained under a unified search strategy.
Related Concept Videos
Pharmacogenomics: Identification of New Drug Targets
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...