Related Experiment Video
Updated: Jan 3, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Preventing undesirable behavior of intelligent machines
Philip S Thomas1, Bruno Castro da Silva2, Andrew G Barto3
1University of Massachusetts, Amherst, MA, USA. pthomas@cs.umass.edu.
We developed a new framework to design machine learning (ML) algorithms, simplifying the regulation of undesirable ML behaviors. This approach successfully prevented dangerous actions in experimental ML systems, promoting safer AI applications.
Area of Science:
- Artificial Intelligence
- Machine Learning Engineering
Background:
- Intelligent machines and machine learning (ML) algorithms are increasingly prevalent in diverse applications, from data analysis to complex task execution.
- Ensuring that ML systems do not exhibit harmful or undesirable behavior is a critical challenge in AI development.
Purpose of the Study:
- To propose a general and flexible framework for designing ML algorithms that simplifies the specification and regulation of undesirable behaviors.
- To demonstrate the practical viability of this framework in creating safer ML systems.
Main Methods:
- Development of a novel framework for ML algorithm design.
- Implementation of ML algorithms based on the proposed framework.
- Experimental evaluation of the designed ML algorithms against standard algorithms to assess the preclusion of dangerous behaviors.
Main Results:
- The proposed framework effectively simplifies the process of defining and controlling undesirable ML behaviors.
- ML algorithms designed using the framework successfully precluded dangerous behaviors observed in standard ML algorithms during experiments.
- The framework proved viable for creating safer ML applications.
Conclusions:
- The developed framework offers a streamlined approach to designing responsible and safe machine learning algorithms.
- This work contributes to the safe and responsible application of machine learning by providing a method to mitigate risks associated with undesirable behaviors.
Related Concept Videos
Stereotype Content Model
Machines: Problem Solving I
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
Machines: Problem Solving II
Behavior Modification
A real-world application of operant conditioning principles is applied...
Non-equilibrium in the Cell
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...

