フォトリアリスティックな交通現場データセット:視覚ベースの分析のための現実的および合成的視点
Khulan Khalzaa1, Stephen Karungaru1, Kenji Terada1
1B1 Research Laboratory, Department of Computer Science, Tokushima University, Tokushima 770-8502, Japan.
Data in brief
|September 2, 2025
まとめ
この研究は,自動運転の知覚のための多様な交通状況データセットを導入します. マルチ高度と深さのビューを含む合成データと現実世界のデータは,コンピュータビジョンモデルのトレーニングを強化します.
科学分野:
- コンピュータ・ビジョン
- 自動運転システム
- 機械学習
背景:
- 認識は自動運転や 単眼カメラからの交通場面の解釈に不可欠です
- 現存するデータセットには 多様性や現実の対称性が欠けているかもしれません
研究 の 目的:
- 交通現場データセットの包括的なコレクションを提示する.
- 自動運転のためのコンピュータビジョンモデルの開発と訓練を支援する.
主な方法:
- トラフィックシーン,トップビュー,マルチハイトビュー,深さという4つのグループに分けられています.
- 新しいデータセットが導入されました:SupporterReal,SupporterVirtual,RealTop,SynthTop.
- RealSenseのステレオカメラを使って RGBと深さの画像を同期した.
主要な成果:
- 感知相似性分析 (LPIPSメトリック) は,合成画像が現実のシーンに非常に似ていることを確認しました.
- データセットはゼロショット単眼深度モデルサポートのために設計されています.
- データの収集と生成は2023年から2025年の間に行われた.
結論:
- 提示されたデータセットは,合成から実際の領域への移行学習を容易にする.
- 標準化されたフォーマットは ディープラーニングのフレームワークとの互換性を保証します
- このコレクションは,コンピュータビジョンのアプリケーションでより広範な再利用を促進します.
関連する概念動画
Vision
48.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
48.6K
Photoreceptors and Visual Pathways
8.5K
At the molecular level, visual signals trigger transformations in photopigment molecules, resulting in changes in the photoreceptor cell's membrane potential. The photon's energy level is denoted by its wavelength, with each specific wavelength of visible light associated with a distinct color. The spectral range of visible light, classified as electromagnetic radiation, spans from 380 to 720 nm. Electromagnetic radiation wavelengths exceeding 720 nm fall under the infrared category,...
8.5K
Depth Perception and Spatial Vision
2.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.7K
Imaging Studies III: Computed Tomography
893
DefinitionComputed Tomography (CT) of the genitourinary (GU) tract is a non-invasive imaging modality that utilizes X-rays and computer processing to generate detailed cross-sectional images of the urinary system, encompassing the kidneys, ureters, bladder, and adjacent structures such as the adrenal glands.PurposeCT scans of the GU tract serve several diagnostic and therapeutic purposes, including:Diagnosis of Urinary Tract Diseases: Detects kidney stones, tumors, cysts, and congenital...
893


