Related Experiment Video
Updated: Aug 6, 2026

On-Site Sampling and Extraction of Brain Tumors for Metabolomics and Lipidomics Analysis
Published on: May 31, 2020
Optimizing GPT-5 for Operation-Procedure-Code-Extraction from Operative Reports in Meningioma Surgery: Feasibility
Sebastian Lehmann1, Florian Wilhelmy1, Frederic V Schwaebe1
1University Hospital Leipzig, Neurosurgery, Sachsen, Germany, Leipzig.
Objective:
In our recent study we showed that GPT's ability to perform OPS-code extraction from operational reports was equivalent to coding neurosurgeons. In this study we aim to evaluate the effect of context enhancement on GPT-5's code extraction abilities.
Methods:
We provided OpenAIs GPT-5 with 100 operational reports (OR) of patients that underwent meningioma surgery. A chat prompt was generated instructing GPT to generate the correct OPS coding. Five groups were formed and provided with different context-enhancing modalities: 1. No additional context (GPT-5s), 2. the current OPS-catalogue (GPT-5o), 3. specified rules on Meningioma coding (GPT-5r), 4. code-triggering sample phrases based on 50 additional ORs (GPT-5e) and 5. all enhancements combined (GPT-5c). We analyzed code-extraction abilities, mistakes and hallucinations.
Results:
Non context enhanced GPT-5s showed the lowest rate of correct coding (44%) and highest number of hallucinations (105) and total mistakes (132). GPT-5o showed identical accuracy to GPT-5s (44%, p = 1.0), but fewer hallucinations (51, p = 0.002) and mistakes (78, p = 0.008). All other models were significantly superior to GPT-5s and GPT-5o in accuracy (GPT-5r 79%, GPT-5e 70%, GPT-5c 86%, p < 0.001), hallucinations (GPT-5r 8, GPT-5e 2, GPT-5c 6, p < 0.001) and mistakes (GPT-5r 21, GPT-5e 35, GPT-5c 14, p < 0.001). Highest coding accuracy was achieved by GPT-5c (GPT-5c vs GPT-5r, p = 0.143, GPT-5c vs GPT5e p = 0.008). GPT-5r and GPT-5e performed equally regarding accuracy, GPT-5r produced fewer mistakes (p = 0.034), while GPT-5e hallucinated less (p = 0.057).
Conclusion:
We show, that specific context enhancement significantly improves GPT-5s ability to perform OPS-code extraction from operational reports, while significantly lowering mistake- and hallucination rates.