Related Experiment Videos
Q-PETAL: Quality-aware Parameter-Efficient Text-to-Image Assessment Learning for Perceptual and Alignment
Abstract:
The rapid proliferation of Text-to-Image (T2I) models across domains such as advertising, education, art, and entertainment has created an urgent need for robust, scalable evaluation metrics. However, existing quality assessment methods rely on specialized models for each quality dimension, lacking robustness and scalability. In this study, we propose Q-PETAL, a parameter-efficient and holistic framework for T2I quality assessment. Q-PETAL includes: (1) quality-aware adaptation to attention projections via low-rank matrix decomposition; (2) patch-level embeddings to evaluate fine-grained perceptual features; (3) semantic-aware text-image alignment using a structured prompt ensemble and cross-modal cosine scoring. A learnable scalar balances the trade-off between text-image alignment and perceptual quality to produce an overall score. Despite having significantly fewer trainable parameters, Q-PETAL outperforms state-of-the-art methods, generalizes well across a wide range of quality variations, and achieves inference latency comparable to that of zero-shot CLIP. Hence, Q-PETAL is an effective and scalable solution for the evolving challenges in T2I generation.