Related Experiment Video
Updated: Oct 27, 2025

Comprehensive Characterization of Extended Defects in Semiconductor Materials by a Scanning Electron Microscope
Published on: May 28, 2016
Learning with Delayed Rewards-A Case Study on Inverse Defect Design in 2D Materials
Suvo Banik1,2, Troy David Loeffler1,2, Rohit Batra1
1Center for Nanoscale Materials, Argonne National Laboratory, Lemont, Illinois 60439, United States.
We developed a reinforcement learning (RL) method using Monte Carlo Tree Search (MCTS) with delayed rewards to efficiently design material defects. This approach overcomes energy barriers, enabling targeted material functionality for technologies like catalysis and electronics.
Area of Science:
- Materials Science
- Computational Materials Science
- Artificial Intelligence in Materials
Background:
- Defect dynamics are crucial for material functionality in technologies like catalysis and microelectronics.
- Understanding and designing defect arrangements is challenging due to intermediate states and energy barriers.
- Current inverse design methods struggle with these complexities.
Purpose of the Study:
- To introduce a novel reinforcement learning (RL) strategy for efficient defect configurational space search.
- To overcome energy barriers in defect transformation for inverse defect design.
- To identify optimal defect arrangements for targeted material functionality.
Main Methods:
- Utilized Monte Carlo Tree Search (MCTS) with delayed rewards, a type of RL.
- Applied the method to a representative case of two-dimensional (2D) molybdenum disulfide (MoS2) with sulfur (S) vacancies.
- Compared the MCTS approach with genetic algorithms for global optimization.
Main Results:
- The RL strategy efficiently sampled the defect configurational space and overcame energy barriers for S vacancies in MoS2 (1.5-8% concentration).
- Observed evolution from random S vacancies to extended S line defects, consistent with experimental findings.
- MCTS with delayed rewards required fewer evaluations and yielded higher quality solutions than genetic algorithms.
Conclusions:
- The developed RL strategy accelerates the inverse design of defects in materials.
- Identified optimal pathways for defect transformation and arrangement in 2D MoS2.
- Discussed implications for 2H to 1T phase transitions in MoS2, impacting material properties.
Related Concept Videos
Plastic Deformations
Purposive Learning
Bending of Material: Problem Solving
Plastic Behavior
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Woodward–Hoffmann Selection Rules and Microscopic Reversibility

