Projection-Free CLIP-Scale EEG Latents via a U-Net-Style Autoencoder
Jeyoung Lee1,2, Jaekwan Ahn2, Jaeseung Sim2
1School of Computer Science and Engineering, Soongsil University, 369 Sangdo-ro, Dongjak-gu, Seoul 06978, Republic of Korea.
None:
Electroencephalography is emerging as a promising conditioning modality for generative visual models. However, existing representation learning approaches often rely on high-capacity masked autoencoders and complex projection networks. When constrained to compact embedding dimensions to match vision-language models, these heavy transformer-based bottlenecks frequently suffer from representation collapse and lose critical signal dynamics. To address this, we propose a lightweight and projection-free autoencoder that directly outputs compact, Contrastive Language-Image Pre-training (CLIP)-scale latent vectors trained toward the CLIP embedding space. Our model adopts a U-Net-style architecture combining one-dimensional convolutional residual blocks for temporal dynamics and inter-channel attention modules for spatial dependencies, alongside skip connections to ensure stable reconstruction. Extensive experiments on visual perception datasets demonstrate that our approach successfully tracks complex signal amplitudes without collapsing. Under strict dimensional constraints, the proposed model achieves superior signal reconstruction fidelity across time and frequency domains using significantly fewer parameters than traditional masked autoencoder baselines. Furthermore, latent space visualizations and zero-shot retrieval tasks reveal that while the baseline collapses toward unstructured, near-chance representations, our architecture preserves emerging, partial semantic organization and retrieves several times above chance. This indicates that the proposed design preserves signal structure while exhibiting preliminary, above-chance semantic alignment, enabling integration into brain-driven generative pipelines.

