Related Experiment Video
Updated: Jun 2, 2026

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
Published on: January 23, 2017
Beyond Accuracy: On the Effects of Fine-tuning Towards Vision-Language Model's Prediction Rationality
Qitong Wang1, Tang Li1, Kien X Nguyen1
1DeepREAL Lab, Department of Computer & Information Sciences, University of Delaware.
Fine-tuning Vision-Language Models (VLMs) can improve accuracy but may rely on invalid evidence. New metrics reveal that while fine-tuned VLMs are more accurate with valid evidence, their trustworthiness requires careful evaluation.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Vision-Language Models (VLMs) like CLIP are widely used.
- Fine-tuning VLMs is common in safety-critical domains.
- Prediction rationality (correctness and valid evidence) is vital in these domains.
Purpose of the Study:
- Investigate the impact of fine-tuning on VLM prediction rationality.
- Introduce novel metrics: Prediction Trustworthiness and Inference Reliability.
Main Methods:
- Conducted extensive experiments across various settings.
- Evaluated fine-tuned VLMs using the proposed metrics.
- Assessed model performance under distributional shifts.
Main Results:
- Fine-tuning improved prediction accuracy but sometimes relied on invalid evidence.
- Fine-tuned VLMs showed higher accuracy when using valid evidence.
- Findings remained consistent across different settings and shifts.
Conclusions:
- Standard fine-tuning may decrease VLM trustworthiness by increasing reliance on invalid evidence.
- Valid evidence identification is key for reliable predictions from fine-tuned VLMs.
- Research offers new insights into VLM fine-tuning for critical applications.
Related Concept Videos
Accuracy, limits, and approximation
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Hindsight Biases
Improving Translational Accuracy
Improving Translational Accuracy
Language and Cognition
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...