Prospective multi-centre evaluation of guideline-based artificial intelligence to streamline multidisciplinary tumour
Oleksandr Ivashchuk1, Serhiy Hovornyan1
1Department of Oncology and Radiology, Bukovinian State Medical University, Chernivtsi, Ukraine.
Introduction:
Multidisciplinary tumour board (MTB) or tumour board (TB) is the "gold standard" for providing oncological patients' diagnosis and treatment. Often, MTBs are time-intensive, capacity-constrained, and absent or inconsistent in many routine hospitals. Despite that, in MTB, a few experienced specialists in the oncological field, sometimes the final decision is mistaken (4-47%), the necessity of treatment plan changing (12-64%), etc. Using artificial intelligence (AI) for cancer patients' treatment is a modern approach, but the main idea of our investigation should be to possibly replace MTB. On the other hand, the choice of correct treatment in oncological patients for some doctors, like surgeons, gynaecologists, and family medicine doctors, is a big trouble. They need to address the guidelines, which are so complicated quick changing. We have offered our AI algorithm, which connects doctors with patients' data, guidelines (only official sources), and then uses LLM for decision-making. Main research question - assess the role of our AI algorithm in diagnosis and treatment for cancer patients, compare with or together with MTB.
Methods:
Prospective, multicentre study at four specialised regional oncological hospitals. Trail period 2023-2025. Additionally, we included 37 doctors (DR) from general and rural hospitals (15 surgeons, 11 gynaecologists, 11 family medicine doctors). 728 cases in eight tumour groups were included in the study: breast 18.13% (132), colorectal 22.12% (161), lung 30.08% (219), prostate 8.79% (64), gynaecological 8.1% (59), gastric 6.59% (48), urothelial 3.99% (29), head & neck 2.2% (16). They were divided into 4 groups: 1) Diagnosis (D) and treatment plan (TP) with MTB only - 226 patients, 2) D and TP with AI only - 206 patients, 3) D and TP with MTB + AI - 147 patients, 4) D and TP with DR + AI - 149 patients. NCCN guidelines were used for all groups. Patients' data were ingested from the e-health system, de-identified, and normalized prior to inference. Evaluation criteria. Immediate - preparation time (PT), decision time (DT), time to recommendation (TTR), where TTR = PT + DT; Midterm - change diagnosis, change treatment plan (1 month, 6 months). We have collected MTB reports and for group N4 - questionnaires.
Results:
Our proposed AI algorithm works as an intermediary between the patient, the doctor, MTB, guidelines, and LLMs. It processes data at each stage for structuring, unification, conversion, machine-readable representation, creates appropriate Python classes, and ultimately a structured supervisory layer with a decision threshold 80%. For immediate criteria we have next results (min): group 1 - PT - 21.3 ± 3.6; DT - 15.6 ± 6.2; TTR - 36.9 ± 7.8; group 2 - PT - 19.7 ± 4.1; DT - 2.1 ± 0.02; TTR - 22.4 ± 3.3; group 3 - PT - 26.9 ± 4.2; DT - 4.9 ± 0.9; TTR - 29.2 ± 3.1; for midterm criteria (change D) results were: on 1 month - group 1 (6.79%), group 2 (5.36%), group 3 (2.78%); on 6 month - group 1 (18.22%), group 2 (19.8%), group 3 (12.4%); midterm criteria (change TP) results were: on 1 month - group 1 (11.31%), group 2 (8.78%), group 3 (4.17%); on 6 month - group 1 (25.12%), group 2 (22.84%), group 3 (15.5%); midterm criteria (change D + TP) results were: on 1 month - group 1 (13.12%), group 2 (8.78%), group 3 (4.86%); on 6 month - group 1 (28.08%), group 2 (24.3%), group 3 (14.73%). The combination of the MTB with AI gives the most effective result in reducing the change in D and TP from 28.08% to 14.73% within 6 months. For group 4, we have found next results. Before the implementation of AI, the time for TP decision in 78.4% of cases was from 30 to 60 minutes, and in 21.6% - more than 60 minutes. At 6 months, DR has time for TP decision, no more than 30 minutes in 67.7%, and among 30 and 60 minutes, 32.3%. Overall assessment using AI for TP in 3 months (good and very good) - 82.8%, in 6 months - 97.1%. Convenience AI in 3 months (good and very good) - 54.2%, in 6 months - 94.1%. Use AI after the end of the study in 3 months (good and very good) - 54.3%, in 6 months - 54.7%. Use in clinical practice the combination of MTB and AI, optimize TTR, decrease the percentage of change D and TP in 1,6 months, decrease the negative sides of single MTB and AI. We can use this combination for first-detected cases without comorbidities at a basic level. AI cannot replace MTB, but it empowers non-oncological clinicians and DR to generate evidence-based, auditable recommendations in settings with limited specialist capacity.
Discussion:
The decrease in the percentage of change D and TP after MTB using AI in the first detected cases highlights the importance of conducting further studies to optimize MTB performance. The positive results of using DR assistance in the form of AI indicate the need to create other practical tools, such as an app, a pocket version, etc. To reach expert-level autonomy, further work is required: a local guideline adaptation, multimodal data ingestion, and longitudinal validation across tumour types, tumour recurrence, multiple cancers, comorbidity, and assessment of the financial efficiency of implementing AI in the work of the MTB.


