Methods of Documentation V: CBE
Methods of Documentation VII: EMR
Data Collection I
Study Designs in Epidemiology
Observational Studies
Methods of Documentation IV: Focus Charting
You might also read
Articles linked to this work by shared authors, journal, and citation graph.
Updated: Feb 18, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Jessica Watson1, Brian D Nicholson2, Willie Hamilton3
1Centre for Academic Primary Care, Bristol Medical School, University of Bristol, Bristol, UK.
This study introduces a standardized method for creating clinical codelists from primary care electronic health records. The method includes three steps: defining the clinical feature, searching for relevant codes using statistical software, and reaching consensus among practitioners through a modified Delphi process. The approach was demonstrated by developing a codelist for shortness of breath in a lung cancer patient cohort. The method allows for sensitivity analysis and ensures reproducibility. The authors recommend that published codelists include quality ratings, definitions, and detailed methods to improve transparency and replication in EHR research.
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Area of Science:
Background:
Electronic health records are widely used in primary care research. However, defining clinical features from these records requires accurate codelist development. Prior research has shown that codelist creation often lacks transparency and reproducibility. This gap motivated the need for standardized methods to ensure scientific rigor. Existing approaches may not fully address the complexity of clinical coding. Researchers have explored various coding strategies, but consensus remains limited. No prior work had resolved how to systematically define and validate clinical codes. This paper introduces a reproducible framework to address these limitations.
Purpose Of The Study:
The study aimed to develop a standardized method for clinical codelist creation. It focused on improving transparency and replicability in EHR research. The authors sought to provide a three-stage process for defining clinical features. They wanted to ensure that codelists could be shared and replicated. The motivation came from the need for consistent and auditable coding practices. The study also aimed to demonstrate the method using a real-world example. Shortness of breath was selected as the clinical feature for illustration. The goal was to show how the method could be applied in primary care EHR studies.
Main Methods:
The first step involved defining the clinical feature using reliable resources. The second step used statistical software to search for all relevant codes. The third step employed a modified Delphi process with primary care practitioners. This process included generating an 'uncertainty' variable for sensitivity analysis. The authors demonstrated the method using shortness of breath as an example. They applied the method to a dataset from the Clinical Practice Research Datalink. The process included identifying candidate codes and reaching consensus on their relevance. The final step involved documenting the syntax for replication and sharing.
Main Results:
A codelist for shortness of breath was developed using the described method. Out of 78 candidate codes, 29 were excluded as inappropriate. Complete agreement was reached for 44 (90%) of the remaining codes. Partial disagreement occurred for 5 (10%) of the codes. The codelist was used to analyze a cohort of 28,216 patients with lung cancer. A total of 13,091 episodes of shortness of breath were identified. Sensitivity analysis showed that uncertain codes were rarely used in practice. The method enabled the creation of a reproducible and auditable codelist.
Conclusions:
The authors propose that standardized codelist development improves EHR research quality. Their method includes a three-stage process for defining and validating clinical codes. The use of a modified Delphi process with primary care practitioners is recommended. The inclusion of an 'uncertainty' variable allows for sensitivity analysis. Published codelists should be badged by quality and include detailed methods. The method supports transparency and replication in primary care EHR studies. The authors suggest that this approach future-proofs findings and enhances auditability. They emphasize the importance of documenting definitions, syntax, and code categorization.
The method produced a codelist for shortness of breath with 44 codes reaching full agreement among practitioners.
It allows primary care practitioners to reach consensus on code relevance and identify uncertain codes for sensitivity analysis.
It enables sensitivity analysis by flagging codes with partial disagreement among practitioners.
It is used to comprehensively search all available codes and generate modifiable syntax for replication.
A total of 13,091 episodes were identified in 28,216 patients with incident lung cancer diagnoses.
They propose badging codelists by quality and reporting definitions, syntax, candidate codes, and Delphi categorization.