Related Experiment Video
Updated: Jan 9, 2026

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
Write Sentence with Images: Revisit the Large Vision Model with Visual Sentence
Quan Liu1, Can Cui1, Ruining Deng1
1Department of Computer Science, Vanderbilt University, Nashville, TN.
Abstract:
This paper introduces a novel framework for generating high-quality images from "visual sentences" extracted from video sequences. By combining a lightweight autoregressive model with a Vector Quantized Generative Adversarial Network (VQGAN), our approach achieves a favorable trade-off between computational efficiency and image fidelity. Unlike conventional methods that require substantial resources, the proposed framework efficiently captures sequential patterns in partially annotated frames and synthesizes coherent, contextually accurate images. Empirical results demonstrate that our method not only attains state-of-the-art performance on various benchmarks but also reduces inference overhead, making it well-suited for real-time and resource-constrained environments. Furthermore, we explore its applicability to medical image analysis, showcasing robust denoising, brightness adjustment, and segmentation capabilities. Overall, our contributions highlight an effective balance between performance and efficiency, paving the way for scalable and adaptive image generation across diverse multimedia domains.
Related Concept Videos
Vision
Visual System
Once through the pupil, the light passes through the lens, a...
Depth Perception and Spatial Vision
Visual Agnosia

