Related Experiment Video
Updated: Sep 14, 2025

Author Spotlight: Efficient Image Recognition Using Directional Gradient Histogram Technique and Support Vector Machines
Published on: January 5, 2024
Multitasking vision language models for vehicle plate recognition with VehiclePaliGemma
Nouar AlDahoul1, Myles Joshua Toledo Tan2, Raghava Reddy Tera3
1Computer Science, New York University Abu Dhabi, Abu Dhabi, UAE.
This study introduces VehiclePaliGemma, a fine-tuned visual language model (VLM) that significantly improves license plate recognition (LPR) accuracy for distorted images. It outperforms existing methods, achieving 87.6% accuracy in challenging conditions.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- License Plate Recognition (LPR) traditionally uses Optical Character Recognition (OCR) but struggles with image distortions like noise, blurring, and close characters.
- Existing LPR methods require substantial improvements for accurate recognition, particularly with challenging image quality.
Purpose of the Study:
- To evaluate the efficacy of various Visual Language Models (VLMs) in overcoming LPR challenges.
- To introduce and validate "VehiclePaliGemma", a specialized VLM for robust license plate recognition.
Main Methods:
- Evaluated multiple VLMs (GPT-4o, Gemini 1.5, PaliGemma, Llama 3.2, Claude 3.5 Sonnet, LLaVA, VILA, moondream2) for license plate recognition.
- Developed and fine-tuned "VehiclePaliGemma" using a dataset of Malaysian license plates under complex conditions.
- Compared VehiclePaliGemma against state-of-the-art methods and other VLMs.
Main Results:
- VehiclePaliGemma achieved a superior accuracy of 87.6% on a challenging dataset.
- The model demonstrated efficient processing at 7 frames per second on an A100-80GB GPU.
- Explored VehiclePaliGemma's multitasking ability in identifying plates from multiple vehicles with varied orientations and conditions.
Conclusions:
- VehiclePaliGemma significantly enhances license plate recognition accuracy, especially for distorted and complex images.
- VLMs offer a promising advancement over traditional OCR-based LPR systems.
- The fine-tuned VehiclePaliGemma model shows potential for real-world applications requiring high-accuracy LPR.
More Related Videos
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
08:13SwarmSight: Real-time Tracking of Insect Antenna Movements and Proboscis Extension Reflex Using a Common Preparation and Conventional Hardware
Published on: December 25, 2017
Related Concept Videos
Parallel Processing
Vision
Multi-input and Multi-variable systems
In the absence...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...