Related Experiment Video
Updated: Jul 4, 2026

P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
Published on: September 8, 2023
Benchmarking Prompting Strategies for Open-Source Language Models in ICHD-3 Headache Classification
1Eberswalde University for Sustainable Development, Germany.
None:
We benchmarked five open-source LLMs (12-24B parameters) on ICHD-3 headache classification using 305 synthetic vignettes from the HeadAI dataset. Five standard prompting techniques achieved 56-61% top-1 accuracy with no significant pairwise differences. A three-tier oracle ablation that progressively provided the correct diagnosis reached 81-94%, revealing a 36-percentage-point gap attributable to retrieval quality. When the retriever placed the correct diagnosis at rank 1 (Hit@1 = 23%), accuracy reached 94.9%. The gap decomposed into retrieval quality (∼24 pp), noise (∼6 pp), and intra-category differentiation (∼6 pp). Rule application is not the bottleneck; knowledge delivery is.