Related Experiment Videos
MeshSeqGen: Zero-shot Mesh Animation based on a Large Video Generation Model
Abstract:
Motion generation for 3D meshes is a fundamental task in computer animation, yet traditional methods like keyframe animation and motion capture are often costly and resource-intensive. Motivated by the progress in large-scale 2D video generation models, we introduce MeshSeqGen, a zero-shot framework that simplifies this process by using the motions from generated video frames to guide mesh deformation. Our method significantly differs from prior video-to-4D techniques lacking mesh-specific structural consistency, and other meshbased methods that require extensive 4D training data. To bridge the gap between 2D videos and 3D meshes, we design a mesh deformation network that learns from a pre-trained large video generation model and a multi-view synthesis model as spatialtemporal consistency guidance, while our mesh-based Gaussian representation ensures the final animations are both plausible and structurally consistent. Extensive evaluations demonstrate the performance of our method over existing approaches by producing diverse motions and high-quality mesh animations from simple text prompts, and making the process user-friendly and highly efficient without the need for complex rigging or keyframing.