Related Experiment Video
Updated: Feb 25, 2026

Holistic Facial Composite Creation and Subsequent Video Line-up Eyewitness Identification Paradigm
Published on: December 24, 2015
FaceScanPaliGemma multi-agent vision language models for facial attribute recognition
Nouar AlDahoul1, Myles Joshua Toledo Tan2, Harishwar Reddy Kasireddy2
1Computer Science Department, New York University Abu Dhabi, Abu Dhabi, UAE.
Abstract:
Technologies for recognizing facial attributes such as race, gender, age, and emotion from images of human faces have several applications, including personalized advertising, sentiment analysis, and the study of demographic trends and social behaviors. Analyzing face images and facial expressions presents several challenges due to the complexity of human facial attributes and the diversity in representation. While numerous attempts have been made to improve facial attribute classification performance, there remains a strong demand for enhanced accuracy. In this paper, we propose "FaceScanPaliGemma," a multi-agent vision language model (VLM) system consisting of four fine-tuned Google PaliGemma models, each specialized for a specific facial attribute classification. To evaluate the proposed solution, we used the public "FairFace" and "AffectNet" datasets. The results show high accuracy, reaching up to 81.1%, 95.8%, 80.0%, and 59.4% for race, gender, age group, and emotion classification, respectively, outperforming other VLMs such as OpenAI GPT, Google Gemini, LLaVA, and Google PaliGemma under zero-shot evaluation.
Related Concept Videos
Facial Feedback Hypothesis
Association Areas of the Cortex
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Modeling and Similitude
Muscles for Facial Expressions
Prosopagnosia
