绿色ViT:一个愿景变换器,用于城市绿色空间细分和可审计面积估计的单路径渐进式取样
Ziqiang Xu1, Young Choi2, Changyong Yi3
1Department of Robot and Smart System Engineering, Kyungpook National University, 80, Daehak-ro, Buk-gu, Daegu 41566, Republic of Korea.
Journal of imaging
|February 26, 2026
概括
绿色ViT使用视觉变压器 (ViT) 精确量化城市绿色空间. 该框架平衡了可靠的绿地指标的准确性和效率,以支持城市规划和环境管理.
科学领域:
- 遥感 遥感 遥感 遥感
- 计算机视觉 计算机视觉
- 城市规划 城市规划
背景情况:
- 城市绿地监测面临着准确性和效率的权衡.
- 现有的方法缺乏综合的可审计面积估计.
- 密集的城市景观对精确的绿色空间量化提出了挑战.
研究的目的:
- 介绍GreenViT,一个基于Vision Transformer (ViT) 的框架,用于精确的城市绿色空间细分和量化.
- 解决城市绿地准确性,效率和可审计面积估计的局限性.
- 开发一个可靠的工具,以决策为导向的长期监测和管理评估.
主要方法:
- 使用了一个ViT-L/14骨干,带有渐进式上采样解码器 (绿头).
- 在高分辨率卫星图像上采用了滑窗采样方案.
- 在手动注释的数据集上进行了实验,该数据集涵盖了五个土地覆盖类别.
主要成果:
- 实现了高性能指标:0.9200 mIoU,0.9580 子和0.9570 PA.
- 经过校准的估计器证明了1.10%的低相对面积误差.
- 展示了适用于薄或边界丰富的绿色空间类别的适用性.
结论:
- 在城市绿色空间监测中,GreenViT在准确性和效率之间提供了强大的平衡.
- 该框架为城市减热和污染控制提供可靠的绿地指标.
- 绿色ViT支持各种应用,包括规划评估,城市更新和生态验证.
相关概念视频
Upsampling
668
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
668
Downsampling
724
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
724
Topographic Surveying and Contours
1.1K
Topographic surveying is critical for documenting the Earth's surface, focusing on capturing elevations, slopes, and natural and man-made features. It is essential in construction planning, water resource management, and land-use analysis. The primary outcome of such surveys is a topographic map, which uses contour lines to visually represent the shape and slope of the terrain, providing valuable insights into the landscape's characteristics.Contour lines are fundamental to understanding the...
1.1K
Deconvolution
639
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
639
Vision
60.7K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
60.7K
Depth Perception and Spatial Vision
2.3K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.3K


