Related Experiment Videos
Landmark Recognition Beyond Curated Benchmarks: Cross-Domain Evaluation of a Multi-Threshold Selective YOLO11
Ulugbek Hudayberdiev1, Abdimumin Alikulov1, Adkham Israilov2
1Faculty of Economics, Samarkand State University of Veterinary Medicine, Livestock and Biotechnologies, Samarkand 140103, Uzbekistan.
Abstract:
Landmark recognition for smart tourism is usually validated on curated benchmark images. In deployment, however, the classifier must handle user-generated photographs whose viewpoint, lighting, resolution, occlusion, and compression differ sharply from curated data. This paper evaluates a previously published multi-threshold enhancement and selective YOLO11n-cls ensemble under this shift, and provides a preliminary zero-shot comparison of three general-purpose multimodal large language models (MLLMs) on the same task. To measure the shift, we build Samarkand v2-SNS, a 300-image out-of-distribution test set of social-media photographs of 12 Samarkand landmarks, disjoint from the training and validation data. Under the shift, four supervised baselines fall by 12.73-22.08 percentage points to 73-80% accuracy, and their in-distribution ranking does not hold. The selective ensemble degrades least (99.24% to 93.00%, -6.24 points) and outperforms the strongest baseline by 13 points. A capacity-matched ablation shows that most of this robustness comes from enhancement diversity, not from generic ensembling. In a preliminary comparison, zero-shot MLLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5) reach only 24.81-54.26%, far below deployment needs. The results argue for reporting out-of-distribution accuracy alongside curated benchmarks, and for hybrid systems that pair compact specialised recognisers with MLLM-based interpretation.