通过来自歧视者的自我监督来实现GANs的空间方向性
概括
本研究介绍了在生成对抗网络 (GAN) 中进行空间控制的自我监督方法. 它通过使用热图来实现直观的图像编辑,改善对图像合成的控制,而无需额外的注释.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 生成型模型实现了光现实的图像合成.
- 通过潜伏空间操纵控制图像生成是一个活跃的研究领域.
- 现有的方法通常需要注释,并专注于全局属性.
研究的目的:
- 开发一种自我监督的方法,以提高GAN的空间方向性.
- 为了实现直观的图像编辑,而不需要额外的人类注释或搜索特定的潜在方向.
- 为了改善对图像生成的控制,以获得定制的输出.
主要方法:
- 引入随机采样的高斯热图作为中间GAN层中的空间感应偏差.
- 采用自主监督的学习策略,在培训期间将热图与GAN区分器的注意力保持一致.
- 集成了DragGAN框架,用于细粒度,粗到细的图像处理.
主要成果:
- 在各种图像领域展示了有效的空间编辑能力:人脸,动物脸,户外场景和复杂的室内场景.
- 通过热图实现了直观的用户交互,通过热图调整场景布局,对象放置和对象移除.
- 展示了整体图像合成质量的改进,以及增强的空间控制.
结论:
- 拟议的自我监督方法显著提高了GANs的空间方向性.
- 这种方法提供了一种无注释和直观的方式来控制图像生成,适用于各种场景类型.
- 与DragGAN的集成允许高效和详细的图像编辑,推进生成模型定制.
相关概念视频
Gyroscope
3.6K
A gyroscope is defined as a spinning disk in which the axis of rotation is free to assume any orientation. When spinning, the orientation of the spin axis is unaffected by the orientation of the body that encloses it. The body or vehicle enclosing the gyroscope can be moved from place to place, while the orientation of the spin axis remains the same. This makes gyroscopes very useful in navigation, especially where magnetic compasses cannot be used, such as in crewed and crewless spacecraft,...
3.6K
State Space Representation
785
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
785
Depth Perception and Spatial Vision
2.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.7K


