Related Experiment Video
Updated: Sep 7, 2026

Deep-Learning Based Multi-Joint Synchronous Tracking for Objective Quantification of Hindlimb Locomotor Kinematics in Rats
Published on: April 3, 2026
SynNeura: Event-driven liquid-spiking dynamics for weakly supervised multimodal temporal alignment
1Information Engineering University, Zhengzhou, 450001, China; Henan Key Laboratory of Cyberspace Situation Awareness, Zhengzhou, 450001, China; Key Laboratory of Cyberspace Security, Ministry of Education, Zhengzhou, 450001, China.
Abstract:
Multimodal temporal alignment is a critical task for applications such as audiovisual understanding, lip reading, and instruction following. However, in weakly supervised settings, challenges like asynchronous sampling, irregular event triggers, and coarse labels hinder precise cross-modal alignment. Frame-based methods rely on fixed time grids, leading to redundant computation in sparse-event scenarios and reducing event-level precision. Differentiable time-warping methods typically require high-resolution inputs, resulting in high computational costs and sensitivity to numerical parameters. To address these challenges, we propose SynNeura, an event-driven continuous-time liquid-spiking neural framework for fine-grained alignment under weak supervision. SynNeura models alignment as a continuous-time latent-state process, analytically propagating states between events and updating only when spikes occur. This results in computational complexity that scales with the number of events rather than the sequence length. SynNeura introduces a piecewise-analytic update using matrix exponentials and trace variables, along with a hierarchical contrastive alignment objective at spike, trajectory, and state levels, enhancing robustness and consistency. Experiments with the AVE, LRS2, and YouCook2 datasets show SynNeura consistently outperforms frame-based and continuous-time baselines in alignment accuracy, temporal consistency, and efficiency. SynNeura achieves 0.291 MAE and 0.713 CAS on AVE, 0.648 TC on LRS2, and 0.829/0.794 EP/ER on YouCook2, with an overall score of 0.738. These results show SynNeura is a scalable, interpretable, and efficient solution for event-driven multimodal temporal alignment.