Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Chi-square Analysis02:46

Chi-square Analysis

44.7K
The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
44.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Digital simulation of kidney collecting system and stone: geometry and texture mapping.

Translational andrology and urology·2026
Same author

Impact of Perioperative Opioid Prescriptions on Early and Late Complications After Rhinoplasty: A Propensity-Matched Study.

The Journal of craniofacial surgery·2026
Same author

Clinically Significant Prostate Cancer in Patients Undergoing Holmium Laser Enucleation of Prostate for Benign Hyperplasia: A Preoperative Nomogram and a Postoperative Surveillance Protocol.

Journal of endourology·2026
Same author

Perioperative Opioid Exposure and Postoperative Complications Following Facial Fracture Repair: A Propensity-Matched Analysis of 71,738 Patients.

The Journal of craniofacial surgery·2026
Same author

A validated custom pipeline for three-dimensional kidney stone renderings tocreate an open access repository.

Urolithiasis·2026
Same author

Effectiveness of pre-incision vs. post-incision PECS II block for breast reconstruction pain: A retrospective cohort study.

Journal of plastic, reconstructive & aesthetic surgery : JPRAS·2026

Related Experiment Video

Updated: Mar 31, 2026

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
05:02

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases

Published on: October 24, 2019

34.4K

Benchmarking Large Language Models Against Web of Science: A Comparative Bibliometric Analysis on Panniculectomy.

Rohan Mangal1, Mark Jessup1, Anshumi Desai2

  • 1University of Miami Miller School of Medicine.

Annals of Plastic Surgery
|March 30, 2026
PubMed
Summary

Large language models (LLMs) show significant limitations in generating accurate bibliometric analyses for scientific literature. ChatGPT-4o performed best among tested LLMs, but caution is advised when using AI-generated data for research.

Keywords:
LLMartificial intelligencebibliometricpanniculectomy

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.8K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

Related Experiment Videos

Last Updated: Mar 31, 2026

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases
05:02

Comparing Bibliometric Analysis Using PubMed, Scopus, and Web of Science Databases

Published on: October 24, 2019

34.4K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.8K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

Area of Science:

  • Bibliometrics
  • Artificial Intelligence in Scientific Research
  • Medical Informatics

Background:

  • Large language models (LLMs) are increasingly explored for scientific writing applications.
  • LLM performance, accuracy, and reliability vary significantly.
  • This study assesses LLM capabilities in bibliometric analysis.

Purpose of the Study:

  • To compare the bibliometric analysis capabilities of three leading LLMs (ChatGPT-4o, Deepseek, Claude 3.7).
  • To evaluate LLM-generated lists of highly cited panniculectomy articles against a manually curated reference dataset.
  • To identify limitations and potential biases in LLM outputs for bibliometric tasks.

Main Methods:

  • Manually extracted the 50 most cited panniculectomy articles from Web of Science (WoS) as a reference.
  • Prompted ChatGPT-4o, Deepseek, and Claude 3.7 to generate their own lists of the top 50 cited articles.
  • Compared LLM outputs against the manual dataset based on citation counts, publication years, journal distribution, author co-occurrence, and article authenticity.

Main Results:

  • LLM outputs showed limited overlap with the manual reference list (ChatGPT-4o: 14%, Claude 3.7: 4%, Deepseek: 0%).
  • Deepseek generated 100% hallucinated articles, while ChatGPT-4o had 40% hallucinations and Claude 3.7 had 70% hallucinations.
  • ChatGPT-4o produced the most accurate results among the LLMs, but still demonstrated significant inaccuracies and confabulations.

Conclusions:

  • Current LLMs exhibit substantial challenges in accurately replicating bibliometric data.
  • ChatGPT-4o demonstrated superior performance compared to Deepseek and Claude 3.7, yet its limitations necessitate caution.
  • Web of Science remains the gold standard for bibliometric analysis; LLM-generated data requires critical evaluation.