EchoVid:Stable DiffusionとCNN拡張トランスフォーマーによる動的コンテンツ作成のためのAI駆動型オーディオ・ツー・ビデオ生成
Deepak Dharrao1, Madhuri Dharrao2, Sneha Padgaonkar1
1Department of Computer Science and Engineering, Symbiosis Institute of Technology, Pune Campus, Symbiosis International (Deemed University), Pune, 412115, India.
Abstract:
Translating spoken words into emotionally and contextually aligned video content remains an open challenge in generative AI. Subtle vocal patterns-such as pauses and pitch modulations-often obscure emotional cues, resulting in visuals that feel emotionally disconnected or flat. While several models excel at text-to-image generation, they struggle with interpreting speech-based inputs, often misreading paralinguistic cues and contextual intent. To address these limitations, this research introduces EchoVid, an audio-to-video synthesis model designed to prioritize contextual fidelity and emotional alignment. A scalable web interface built with React.js and TypeScript connects to Node.js backend with the MongoDB Atlas for near-real-time generation (≈ 1.2× input-duration latency at 512² frames on RTX A4000) interaction. Using PyAudio input, EchoVid guides Hugging Face's Stable Diffusion v2.1 via emotion-aware prompts, with CNN-enhanced diffusion transformers supporting the video generation process. Preliminary results show that EchoVid can generate visuals that reflect both emotional tone (e.g., joyful imagery for upbeat speech) and context. The proposed EchoVid model is compared with MoCoGAN and Stable Video Diffusion variants based on metrics like, FVD, FID-VID and CLIPScore. Further this research introduces two novel evaluation metrics namely Temporal Semantic Stability (TSS) and Perceptual Flicker Index (PFI) that scores the semantic consistency and frame-to-frame change in the generated video. The results show that EchoVid outperforms the other models and can generate relatively better videos.
関連する概念動画
Non-equilibrium in the Cell
Transformers
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Source Transformation
It is essential to note that when...
Reconstruction of Signal using Interpolation
Upsampling

