Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Classification of Systems-I01:26

Classification of Systems-I

533
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
533
Aggregates Classification01:29

Aggregates Classification

947
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
947

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Effect of video-based game therapy on activation of scapular muscles in children with thoracic hyperkyphosis.

PM & R : the journal of injury, function, and rehabilitation·2026
Same author

Sacral Dysmorphism as an Independent Risk Factor for Cortical Breach During Percutaneous S1 Iliosacral Screw Fixation.

Journal of clinical medicine·2026
Same author

Sagittal Alignment Reciprocal Changes After Thoracolumbar/Lumbar Anterior Vertebral Body Tethering.

Journal of clinical medicine·2026
Same author

Association of the Hemoglobin-Albumin-Lymphocyte-Platelet (HALP) Score with 3-Month Outcomes After Lumbar Medial Branch Radiofrequency Ablation: A Retrospective Cohort Study.

Diagnostics (Basel, Switzerland)·2025
Same author

Comparative Effectiveness of Ultrasound-Guided Corticosteroid Injection, Radiofrequency Ablation, and Their Combination for Recalcitrant Plantar Fasciitis: A Retrospective Cohort Study.

Journal of foot and ankle research·2025
Same author

Sagittal spinopelvic alignment in healthy Turkish adults: establishing normative radiographic reference values.

European spine journal : official publication of the European Spine Society, the European Spinal Deformity Society, and the European Section of the Cervical Spine Research Society·2025

Related Experiment Video

Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

983

Reliability of Large Language Model-Based Artificial Intelligence in AIS Assessment: Lenke Classification and

Cemil Aktan1, Akın Koşar1, Melih Ünal1

  • 1Department of Orthopedics and Traumatology, Antalya Training and Research Hospital, Antalya 07100, Turkey.

Diagnostics (Basel, Switzerland)
|December 30, 2025
PubMed
Summary

Current large language models (LLMs) are unreliable for adolescent idiopathic scoliosis (AIS) surgical planning. These AI tools show poor agreement with expert surgeons for deformity classification and fusion-level selection, making them unsuitable for clinical use without further validation.

Keywords:
Lenke classificationadolescent idiopathic scoliosisartificial intelligencedeep learningmultimodal large language models

More Related Videos

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

5.2K

Related Experiment Videos

Last Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

983
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

5.2K

Area of Science:

  • Spine Surgery
  • Artificial Intelligence in Medicine
  • Radiographic Assessment

Background:

  • Accurate adolescent idiopathic scoliosis (AIS) classification and fusion-level planning are critical for surgical success.
  • Traditional methods rely on Cobb angle and the Lenke system, requiring expert interpretation.
  • Multimodal large language models (LLMs) are emerging for image analysis but lack validation in orthopedic decision-making.

Purpose of the Study:

  • To evaluate the agreement and reproducibility of contemporary multimodal LLMs for AIS radiographic assessment compared to expert spine surgeons.
  • To determine the clinical reliability of LLMs in Lenke classification and fusion-level selection for AIS patients.

Main Methods:

  • Retrospective analysis of 125 AIS patients' spinal radiographs (AP, lateral, side-bending).
  • Independent classification and fusion-level selection by two expert spine surgeons, with consensus for reference standard.
  • Analysis of radiographs by four multimodal LLMs using zero-shot prompts; assessment of agreement (Cohen's κ) and test-retest reproducibility.

Main Results:

  • Expert surgeons demonstrated high agreement (κ=0.913 for Lenke, κ=0.879 for fusion).
  • All LLMs showed chance-level reproducibility and very low agreement with expert consensus (Lenke: κ=0.001-0.036; fusion: κ=0.003-0.053).
  • LLMs were significantly faster (~seconds vs. ~11-12 min) but lacked clinical reliability; one LLM produced missing outputs.

Conclusions:

  • General-purpose multimodal LLMs currently fail to provide reliable Lenke classification or fusion-level planning for AIS.
  • The poor agreement and internal inconsistency of LLM interpretations preclude their use in surgical decision-making without task-specific validation.
  • Further research is needed to develop and validate specialized AI tools for accurate AIS assessment.