抽象化レベルを超えた機械と人間の視覚表現の並べ替え
Lukas Muttenthaler1,2,3,4, Klaus Greff5, Frieda Born6,7,8
1Google DeepMind, Berlin, Germany. lukas.muttenthaler@tu-berlin.de.
Nature
|November 12, 2025
まとめ
ディープニューラルネットワーク (DNN) は人間のように一般化できず,その表現には階層構造が欠けています. この研究はDNNを人間の知識で強化し 人間の認知能力と整合させ 機械学習の性能を向上させます
科学分野:
- 人工知能
- 認知科学
- コンピュータ・ビジョン
背景:
- ディープニューラルネットワーク (DNN) は,人間の行動とニューラル表現のモデルとしてますます使用されています.
- しかし,DNNのトレーニングと人間の学習には大きな違いがあり,モデルの汎用性には欠陥がある.
- 現在のビジョンモデルは,人間の概念的知識の階層的な組織を捉えることができません.
研究 の 目的:
- 人間の概念的知識とDNNの表現の不一致を特定し,対処する.
- より一般化され,より堅固な視覚モデルを開発する.
- 人工知能 (AI) システムに人間の知識を注入する方法を調査する.
主な方法:
- 概念の類似性について 人間の判断を模倣する モデル教師を訓練しました
- 教師モデルから 最先端の視覚基盤モデルに 微調整によって人間に合わせた表現構造を 移した.
- 複数の意味論抽象化レベルにおける人間の判断を用いた類似性タスクのモデルを評価した.
主要な成果:
- 人間の行動と不確実性を より正確に推定することが示されました
- 様々な機械学習のタスクのパフォーマンスが向上し,一般化と配送外での堅実性が向上しました.
- モデルは人間の認知に存在する階層的な意味論的抽象を成功裏に捉えました
結論:
- 人間の知識をDNNに統合することで ヒトの認知判断とよりよく一致する ハイブリッド表現が生まれます
- このアプローチにより より堅牢で 解釈しやすく 人間に合わせた AI システムが生まれます
- 人工知能に人間の知識を注入することで より有能で信頼性の高い 人工知能への道を切り開くことができます
関連する概念動画
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
High-Level and Low-Level Awareness
618
Controlled processes in human consciousness represent high-alert mental states where individuals deliberately focus their attention on achieving specific goals. Controlled processes can be seen in situations like mastering new technology, where a person might become so absorbed that they ignore surrounding distractions. Such processes involve selective attention, requiring one to concentrate on particular elements of experience while disregarding others. These are governed by executive...
618
Machines: Problem Solving II
632
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. Consider a lifting tong carrying a 100 kg load. It comprises movable sections DAF and CBG linked together with member AB.
632
Machines: Problem Solving I
670
A toggle clamp is a mechanical device commonly used for holding and clamping objects in various applications, such as woodworking, metalworking, and assembly operations. Consider a toggle clamp subjected to a force of 200 N at the handle. The vertical clamping force can be calculated, provided the dimensions of the toggle clamp are known.
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
670
Modeling and Similitude
598
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
598
Gestalt Principles of Perception
1.0K
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
1.0K


