Machines that halt resolve the undecidability of artificial intelligence alignment

Gabriel A Melo1, Marcos R O A Máximo2, Nei Y Soma2

  • 1Department of Computer Science, Instituto Tecnológico de Aeronáutica, São José dos Campos, SP, Brazil. gam@ita.br.

Scientific Reports
|May 4, 2025
PubMed
Summary

The inner alignment problem for artificial intelligence (AI) is undecidable. We propose AI alignment should be an intrinsic architectural property, incorporating a halting constraint for guaranteed safety.

Related Concept Videos

Introduction to Cognitive Psychology01:20

Introduction to Cognitive Psychology

Cognitive psychology is the field of psychology dedicated to examining how people think. It attempts to explain how and why we think the way we do by studying the interactions among human thinking, emotion, creativity, language, and problem-solving, as well as other cognitive processes. Cognitive psychology studies how information is processed and manipulated in remembering, thinking, and knowing.
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
229
Natural and Artificial Concepts01:24

Natural and Artificial Concepts

In psychology, concepts can be divided into two categories: natural and artificial. Natural concepts are formed through direct or indirect experiences. For example, consider the concept of snow. If you live in a place with regular snowfall, such as Essex Junction, Vermont, you know snow through direct experiences. You’ve seen it fall, touched it, shoveled it, and played in it. You recognize its texture, appearance, and even its smell. In contrast, if you live on an island like Saint...
77
Stereotype Content Model02:16

Stereotype Content Model

The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Reason and Intuition01:37

Reason and Intuition

The human brain processes information for decision-making using one of two routes: an intuitive system and a rational system (Epstein, 1994; popularized by Kahneman, 2011 as System 1 and System 2, respectively). The intuitive system is quick, impulsive, and operates with minimal effort, relying on emotions or habits to provide cues for what to do next, while the rational system is logical, analytical, deliberate, and methodical. Research in neuropsychology suggests that the...
6.3K
Hypothesis: Accept or Fail to Reject?01:17

Hypothesis: Accept or Fail to Reject?

The outcome of any hypothesis testing leads to rejecting or not rejecting the null hypothesis. This decision is taken based on the analysis of the data, an appropriate test statistic, an appropriate confidence level, the critical values, and P-values. However, when the evidence suggests that the null hypothesis cannot be rejected, is it right to say, 'Accept' the null hypothesis?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
27.3K
ortho–para-Directing Deactivators: Halogens01:24

orthopara-Directing Deactivators: Halogens

Halogens are ortho–para directors. They are more electronegative than carbon. Therefore, as ring substituents, they can withdraw electrons through the inductive effect and deactivate the aromatic ring towards electrophilic substitution. Halogens also have an electron-donating resonance effect on the ring, which influences the orientation of the incoming electrophile. If an electrophile attacks at the ortho or the para position, the halogen donates electrons and stabilizes the intermediate...
5.2K