Related Experiment Video
Updated: Jan 12, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Can Contemporary Large Language Models Provide the Domain Knowledge Needed for Causal Inference? Evaluating Automated
Maryam Aziz1, M Alan Brookhart1
1Department of Population Health Sciences, Duke University School of Medicine, Durham, NC, USA.
Purpose:
Directed acyclic graphs (DAGs) are critical in epidemiology and public health research for guiding study design and minimizing bias. Yet, developing DAGs for causal inference requires substantial domain knowledge. Given the vast amounts of training data for large language models (LLMs), this study assesses the effectiveness of prompt engineering for LLMs to generate DAGs that depict causal relationships in population health using OpenAI's GPT-4o and GPT-o1.
Methods:
We consider a hypothetical study on statins vs no treatment for prevention of cardiovascular disease in a general adult population. We assessed four types of prompt engineering strategies: zero-shot, one-shot, instruction based, and chain of thought (CoT) prompts. Generated DAGs were assessed based on consistency, acyclicity, accuracy of sources, completeness (based on ASCVD risk score criteria), and adherence to the prompt.
Results:
We found that all generated DAGs were acyclic, except for one run using the instruction-based prompt. Additionally, more than half of the DAGs included 6/7 of the ASCVD criteria, though race was absent from all. Overall, CoT resulted in the most complete DAGs and one-shot provided the most consistency across runs and adherence to the task in the prompt. The zero-shot prompt performed notably better on GPT-o1 compared to GPT-4o, consistently providing justifications and sources for variable inclusion.
Conclusion:
While the findings suggest that LLMs have a baseline capacity to generate DAGs that adhere to basic epidemiological conventions, we also found several limitations including lack of justification, systematic omission of race, and frequent source hallucination, highlighting the need for human oversight and expertise. We conclude that contemporary LLMs cannot replace a domain expert's judgment but may serve as a brainstorming or pre-analysis tool for DAG development when guided by well-engineered prompts.
Related Concept Videos
Vector Algebra: Graphical Method
We use the laws of geometry to construct resultant vectors, followed by trigonometry to find vector magnitudes and directions. For a geometric construction of the sum of two vectors in a plane, we follow the parallelogram rule. Suppose two vectors are at arbitrary positions. Translate either one of...
Causality in Epidemiology
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Correlation and Causation
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
Graphs of Equations in Two Variables

