Related Experiment Video
Updated: Jan 17, 2026

06:48
Emergency Undocking in Robotic Surgery: A Simulation Curriculum
Published on: May 20, 2018
10.1K
Hierarchical Reinforcement Learning for Fair Operating Room Scheduling in Plastic Surgery
Christoph Wallner1, Flemming Puscz1, Sonja Verena Schmidt1
1From the Department for Plastic and Hand Surgery, BG University Hospital Bergmannsheil Bochum, BG University Hospital Bergmannsheil, Ruhr University Bochum, Bochum, Germany.
Plastic and Reconstructive Surgery. Global Open
|September 22, 2025
Summary
A new hierarchical reinforcement learning (HRL) framework optimizes operating room (OR) scheduling, reducing bias and improving surgical training equity. This data-driven approach balances efficiency with fair resource allocation for surgeons.
Area of Science:
- Surgical Operations Research
- Machine Learning in Healthcare
- Health Equity
Background:
- Equitable operating room (OR) scheduling is vital for surgical training and resource optimization.
- Traditional scheduling methods often contain hidden biases, leading to unequal resource distribution.
- A hierarchical reinforcement learning (HRL) framework is proposed to address these biases.
Purpose of the Study:
- To develop and evaluate an HRL framework for optimizing OR assignment.
- To analyze and improve fairness in surgical scheduling.
- To maintain institutional constraints while enhancing efficiency and equity.
Main Methods:
- Retrospective analysis of 1.5 years of OR assignments for a plastic surgery service (24 surgeons).
- Utilized a random forest model with Shapley Additive Explanations to predict OR demand and quantify feature importance.
- Implemented an HRL architecture (Proximal Policy Optimization and Deep Q-Learning) to reallocate OR days, considering various constraints.
Main Results:
- Baseline scheduling exhibited significant inequality (Gini = 10.39).
- HRL reduced inequality (Gini = 0.22) and mean absolute deviation (33.6 to 25.9).
- OR exposure increased for fellows (+50%) and junior residents (-14%), preserving total OR capacity with 100% rule compliance.
Conclusions:
- The HRL framework provides data-driven, fair, and constraint-adherent OR schedules.
- This approach balances operational efficiency with educational equity in surgical training.
- Multicenter prospective validation is recommended for broader applicability.
More Related Videos
Related Concept Videos
Reinforcement Schedules
460
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
460
Operant Conditioning Intervention
461
Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...
In operant conditioning, behaviors that are...
461

