多任务视觉语言模型用于使用VehiclePaliGemma识别车辆车牌.
Nouar AlDahoul1, Myles Joshua Toledo Tan2, Raghava Reddy Tera3
1Computer Science, New York University Abu Dhabi, Abu Dhabi, UAE.
Scientific reports
|July 18, 2025
概括
本研究介绍了VehiclePaliGemma,这是一个微调的视觉语言模型 (VLM),可以显著提高对扭曲图像的车牌识别 (LPR) 准确性. 它的性能优于现有方法,在具有挑战性的条件下达到87.6%的准确性.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 车牌识别 (LPR) 传统上使用光学字符识别 (OCR),但与噪音,模糊和近距离字符等图像扭曲作斗争.
- 现有的LPR方法需要大幅改进,以获得准确的识别,特别是具有挑战性的图像质量.
研究的目的:
- 评估各种视觉语言模型 (VLM) 在克服LPR挑战方面的有效性.
- 引入和验证"VehiclePaliGemma",这是一个专门的VLM,用于强大的车牌识别.
主要方法:
- 评估了多个VLM (GPT-4o,Gemini 1.5,PaliGemma,Llama 3.2,Claude 3.5 Sonnet,LLaVA,VILA,moonream2) 用于识别车牌. 这是一个非常好的方法.
- 在复杂的条件下使用马来西亚车牌数据集开发和微调"VehiclePaliGemma".
- 车辆PaliGemma与最先进的方法和其他VLM进行了比较.
主要成果:
- 在一个具有挑战性的数据集上,VehiclePaliGemma实现了87.6%的卓越精度.
- 该模型在A100-80GB GPU上以每秒7的速度展示了高效的处理.
- 探索了VehiclePaliGemma的多任务能力,用于识别不同方向和条件的多辆车的车牌.
结论:
- 车辆PaliGemma显著提高了车牌识别的准确性,特别是在扭曲和复杂的图像.
- 比起传统的基于OCR的LPR系统,VLM提供了一个有前途的进步.
- 微调的VehiclePaliGemma模型显示了需要高精度LPR的现实应用的潜力.
更多相关视频
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.1K
08:13SwarmSight: Real-time Tracking of Insect Antenna Movements and Proboscis Extension Reflex Using a Common Preparation and Conventional Hardware
Published on: December 25, 2017
8.3K
相关概念视频
Parallel Processing
230
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
230
Vision
55.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.4K
Multi-input and Multi-variable systems
150
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
150
Force Classification
1.6K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.6K
