城市环境的语义细分:利用U-Net深度学习模型进行城市景观图像分析
T S Arulananth1, P G Kuppusamy2, Ramesh Kumar Ayyasamy3
1Department of Electronics and Communication Engineering, MLR Institute of Technology, Hyderabad, India.
PloS one
|April 5, 2024
概括
本研究使用U-Net深度学习模型探索城市景观的语义细分. 该U-Net架构实现了最先进的结果,证明了城市图像分析的高准确性.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 城市信息学 城市信息学
背景情况:
- 城市景观的语义细分对于自动驾驶和智能城市开发等应用至关重要.
- 深度学习模型提供了先进的能力来解释复杂的城市环境.
- 现有的方法在准确细分各种城市景观元素方面面临挑战.
研究的目的:
- 调查U-Net深度学习模型对城市景观语义细分的有效性.
- 探索U-Net架构对于详细的城市场景理解的适用性.
- 根据已建立的基准来评估模型的性能.
主要方法:
- 实现一个U-Net架构,采用编码器-解码器结构,并具有卷积层和上样层.
- 利用批量规范化和放弃用于模型稳定和规范化.
- 在综合Cityscapes数据集上进行的实验和评估.
主要成果:
- 拟议的U-Net模型在城市景观语义细分方面实现了最先进的性能.
- 证明了高精度,平均交叉点在欧盟 (mIOU) 和平均子系数.
- 在细分复杂的城市场景方面表现优于现有模型.
结论:
- 对于城市景观的语义细分,U-Net模型非常有效,提供了显著的改进.
- 建筑物捕捉层次特征的能力是其在城市图像分析中的成功的关键.
- 这项研究有助于推进自动驾驶,城市规划和智能城市技术.
相关概念视频
Visual System
580
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
580
Parallel Processing
150
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
150
Depth Perception and Spatial Vision
644
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
644


