"Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue" by Chengqian Ma , Wei Tao , Haoyu Zhang , Yiwen Guo
TLDR:
Motion-Omni is a groundbreaking framework that seamlessly integrates spoken dialogue generation and full-body co-speech motion by leveraging shared hidden states. By implementing scalable pseudo-labeling and a unified evaluation method, Motion-Omni achieves real-time, aligned responses. Unlike existing methods where speech and motion are generated separately, Motion-Omni generates explicit facial expressions and full-body motion directly from the speech-producing hidden states. The framework involves joint training of the spoken dialogue model and motion generator to ensure alignment, with supervisory labels created using a scalable pipeline. Motion-Omni also introduces SwDA-500 and a new public evaluation protocol for stochastic open-ended full-body spoken dialogue, showcasing superior performance metrics in comparison to other omni-modal systems.
Read Paper / Blog

huggingface.co
Paper page - Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
Join the discussion on this paper page
September 7, 2026 36