Related Experiment Video
Updated: Aug 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Models for AI-Assisted Decision Support in Legal Capacity Assessment: A Comparative Study
Halit Canberk Aydogan1, Muhammet Sevindik2, Zeynep Unat Öztürk1
1Department of Forensic Medicine, Ordu University Training and Research Hospital, Ordu 52200, Türkiye.
Healthcare (Basel, Switzerland)
|August 13, 2026
Summary
Large language models (LLMs) show promise in assisting legal capacity assessments, aligning with expert recommendations in standardized cases. Further research is needed to ensure safe integration into medico-legal practice under expert supervision.
Area of Science:
- Medical Law
- Artificial Intelligence
- Clinical Decision Support
Background:
- Legal capacity assessment is complex, requiring multidisciplinary input.
- The utility of large language models (LLMs) in supporting legal capacity assessments is not well-defined.
- This study investigates LLMs as AI-assisted decision-support tools for legal capacity evaluations.
Purpose of the Study:
- To evaluate the performance of LLMs in assisting legal capacity assessments.
- To compare LLM outputs with expert interdisciplinary medical board (IMBD) recommendations.
- To assess the reliability and usability of LLMs in a medico-legal context.
Main Methods:
- Retrospective analysis of 234 court-referred adult cases (2018-2024).
- Standardized medico-legal case vignettes evaluated by ChatGPT-5.2, Gemini 3 Pro, and Claude 4.5 Sonnet.
- Comparison of LLM outputs with IMBD recommendations using accuracy, macro-F1, Cohen's κ, AUC, and other metrics.
Main Results:
- Gemini 3 Pro showed the highest accuracy (86.3%), ChatGPT-5.2 the highest macro-F1 (0.79) and AUC (0.91).
- LLM agreement with IMBD recommendations ranged from κ = 0.60 to 0.73, with high temporal stability (κ = 0.87-0.95).
- Performance varied by outcome, with limitations noted for Article 408 recommendations; minimal safety flags were identified.
Conclusions:
- LLMs demonstrated agreement with expert recommendations in standardized legal capacity assessments.
- LLMs show potential as AI-assisted decision-support tools in medico-legal settings, requiring expert supervision.
- Careful, supervised integration is crucial due to the high stakes involved in legal capacity decisions.