Related Experiment Video
Updated: Jun 11, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Prompt-guided and multimodal landscape scenicness assessments with vision-language models.
Alex Levering1,2, Diego Marcos3, Nathan Jacobs4
1Laboratory of Geo-Information Science and Remote Sensing, Wageningen University, Wageningen, the Netherlands.
Vision-Language Models (VLMs) offer efficient landscape scenicness prediction using few-shot learning. These models outperform traditional methods, even with limited data, enabling flexible aesthetic assessments.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
- Environmental Science
Background:
- Deep learning and Vision-Language Models (VLMs) facilitate efficient transfer learning for downstream tasks with limited labeled data.
- VLMs enable direct comparison between textual descriptions and image content, opening new avenues for image analysis and annotation.
- Assessing landscape scenicness, or aesthetic quality, is crucial for environmental planning and has traditionally required extensive human annotation.
Purpose of the Study:
- To evaluate the potential of VLMs for landscape scenicness prediction using zero- and few-shot learning methodologies.
- To compare the performance of VLM-based approaches against traditional fully supervised methods in landscape aesthetic assessment.
- To introduce and validate a novel annotation method, Landscape Prompt Ensembling (LPE), for acquiring landscape scenicness ratings.
Main Methods:
- Few-shot learning was implemented by fine-tuning a single linear layer on pre-trained VLM representations.
- Zero-shot prediction was explored using contrastive prompting with positive and negative landscape aesthetic concepts.
- Landscape Prompt Ensembling (LPE) was developed as an annotation technique utilizing rated text descriptions without requiring an image dataset.
Main Results:
- A model fine-tuned on a few hundred samples demonstrated performance comparable to, or exceeding, a fully supervised model trained on hundreds of thousands of examples.
- Zero-shot contrastive prompting outperformed few-shot linear probing when optimizing prompt configurations with a small number of samples.
- LPE successfully generated landscape scenicness assessments that showed concordance with established image-based rating datasets.
Conclusions:
- VLMs, particularly through zero- and few-shot learning, offer a highly efficient and flexible approach to landscape scenicness assessment.
- The developed LPE method provides a scalable and effective alternative for landscape aesthetic data acquisition, reducing reliance on extensive image datasets.
- These findings highlight the transformative potential of VLMs in landscape analysis, enabling more accessible and adaptable scenicness evaluations.
More Related Videos
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
08:25Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019