Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models in German Continuing Medical Education Assessments: Protocol for a Fully Crossed Experimental
Leyla Özmen1, Christian Burisch2,3, Daniel Gödde4
1Faculty of Health, Witten/Herdecke University, Alfred-Herrhausen-Straße 50, Witten, North Rhine-Westphalia, 58455, Germany, 49 23029260.
JMIR Research Protocols
|July 28, 2026
Summary
Large language models (LLMs) can cheat on continuing medical education (CME) tests. This study examines how PDF format affects LLM performance on CME assessments, exploring safeguards against AI-assisted cheating.
Area of Science:
- Medical Education
- Artificial Intelligence
- Document Engineering
Background:
- Continuing Medical Education (CME) is a mandatory requirement for physicians in Germany.
- The emergence of advanced Large Language Models (LLMs) poses a significant threat to the integrity of CME assessments.
- LLMs have demonstrated the capability to pass German CME tests, raising concerns about academic honesty.
Purpose of the Study:
- To investigate the impact of different PDF document formats on LLM performance in answering CME test questions.
- To determine if specific PDF formats can impede LLMs from achieving passing scores on CME assessments.
- To evaluate the influence of various LLMs on their ability to solve CME test questions.
Main Methods:
- A within-subjects repeated-measures design using 18 expired CME articles from 3 German publishers across 6 specialties.
- Conversion of articles into 4 PDF formats: searchable, protected, raster, and vector.
- Testing 4 current LLMs (GPT-5, Claude Sonnet 4, Grok-4, Gemini 3) on each article-format combination, with each model answering each article 3 times.
- Primary outcome: proportion of correctly answered questions; Secondary outcome: pass/fail rate.
Main Results:
- Data collection is planned for June 2026, with results expected later that year.
- Analyses will focus on quantifying performance differences across document formats.
- Findings will inform the feasibility of non-searchable document formats as a measure to mitigate LLM-enabled cheating.
Conclusions:
- Document format significantly influences LLM performance on CME assessments.
- Non-searchable PDF formats may serve as a temporary technical safeguard against AI-assisted cheating in CME.
- The study aims to guide regulators and CME providers in balancing assessment validity, accessibility, and responsible AI integration.