Related Experiment Video
Updated: May 9, 2025

One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
Machines that halt resolve the undecidability of artificial intelligence alignment
Gabriel A Melo1, Marcos R O A Máximo2, Nei Y Soma2
1Department of Computer Science, Instituto Tecnológico de Aeronáutica, São José dos Campos, SP, Brazil. gam@ita.br.
The inner alignment problem for artificial intelligence (AI) is undecidable. We propose AI alignment should be an intrinsic architectural property, incorporating a halting constraint for guaranteed safety.
Area of Science:
- Artificial Intelligence
- Theoretical Computer Science
- AI Safety
Background:
- The inner alignment problem questions if AI models satisfy alignment functions.
- This problem is proven undecidable via Rice's theorem, linked to Turing's Halting Problem.
- Current alignment strategies often apply post-hoc adjustments to AI models.
Purpose of the Study:
- To demonstrate the undecidability of the inner alignment problem.
- To advocate for AI alignment as an inherent architectural feature.
- To introduce a halting constraint for AI judge functions to ensure finite execution.
Main Methods:
- Proof sketch reduction to Turing's Halting Problem.
- Analysis of enumerable sets of provably aligned AI operations.
- Conceptual modeling of AI architectures with intrinsic alignment and halting constraints.
Main Results:
- Rigorous proof of the inner alignment problem's undecidability.
- Identification of an enumerable set of provably aligned AI operations.
- Demonstration of AI architectures with built-in alignment and halting guarantees.
Conclusions:
- AI alignment should be a guaranteed property designed into AI architectures from inception.
- Outer alignment judge functions require a halting constraint for reliable AI behavior.
- An intrinsically hard-aligned AI approach ensures safety and predictable termination.
More Related Videos
Related Concept Videos
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
Natural and Artificial Concepts
Stereotype Content Model
Reason and Intuition
Hypothesis: Accept or Fail to Reject?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
ortho–para-Directing Deactivators: Halogens

