Related Experiment Video
Updated: Jun 26, 2025

Comprehensive Endovascular and Open Surgical Management of Cerebral Arteriovenous Malformations
Published on: October 20, 2017
Reliability of study endpoint adjudication in a pragmatic trial on brain arteriovenous malformations
Tim E Darsaut1, Anass Benomar2, Elsa Magro3
1University of Alberta Hospital, Division of Neurosurgery, Department of Surgery, Mackenzie Health Sciences Centre, Edmonton, Alberta, Canada.
Background:
The results of a clinical trial are given in terms of primary and secondary outcomes that are obtained for each patient. Just as an instrument should provide the same result when the same object is measured repeatedly, the agreement of the adjudication of a clinical outcome between various raters is fundamental to interpret study results. The reliability of the adjudication of study endpoints determined by examination of the electronic case report forms of a pragmatic trial has not previously been tested.
Methods:
The electronic case report forms of 62/434 (14%) patients selected to be observed in a study on brain AVMs were independently examined twice (4 weeks apart) by 8 raters who judged whether each patient had reached the following study endpoints: (1) new intracranial hemorrhage related to AVM or to treatment; (2) new non-hemorrhagic neurological event; (3) increase in mRS ≥1; (4) serious adverse events (SAE). Inter and intra-rater reliability were assessed using Gwet's AC1 (κG) statistics, and correlations with mRS score using Cramer's V test.
Results:
There was almost perfect agreement for intracranial hemorrhage (92% agreement; κG = 0.84 (95%CI: 0.76-0.93), and substantial agreement for SAEs (88% agreement; κG = 0.77 (95%CI: 0.67-0.86) and new non-hemorrhagic neurological event (80% agreement; κG = 0.61 (95%CI: 0.50-0.72). Most endpoints correlated (V = 0.21-0.57) with an increase in mRS of ≥1, an endpoint which was itself moderately reliable (76% agreement; κG = 0.54 (95%CI: 0.43-0.64).
Conclusion:
Study endpoints of a pragmatic trial were shown to be reliable. More studies on the reliability of pragmatic trial endpoints are needed.

