English

Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

Computer Vision and Pattern Recognition 2025-07-02 v2 Artificial Intelligence

Abstract

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can both comprehend and generate dyadic behavioral dynamics. To this end, we introduce the Seamless Interaction Dataset, a large-scale collection of over 4,000 hours of face-to-face interaction footage from over 4,000 participants in diverse contexts. This dataset enables the development of AI technologies that understand dyadic embodied dynamics, unlocking breakthroughs in virtual agents, telepresence experiences, and multimodal content analysis tools. We also develop a suite of models that utilize the dataset to generate dyadic motion gestures and facial expressions aligned with human speech. These models can take as input both the speech and visual behavior of their interlocutors. We present a variant with speech from an LLM model and integrations with 2D and 3D rendering methods, bringing us closer to interactive virtual agents. Additionally, we describe controllable variants of our motion models that can adapt emotional responses and expressivity levels, as well as generating more semantically-relevant gestures. Finally, we discuss methods for assessing the quality of these dyadic motion models, which are demonstrating the potential for more intuitive and responsive human-AI interactions.

Keywords

Cite

@article{arxiv.2506.22554,
  title  = {Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset},
  author = {Vasu Agrawal and Akinniyi Akinyemi and Kathryn Alvero and Morteza Behrooz and Julia Buffalini and Fabio Maria Carlucci and Joy Chen and Junming Chen and Zhang Chen and Shiyang Cheng and Praveen Chowdary and Joe Chuang and Antony D'Avirro and Jon Daly and Ning Dong and Mark Duppenthaler and Cynthia Gao and Jeff Girard and Martin Gleize and Sahir Gomez and Hongyu Gong and Srivathsan Govindarajan and Brandon Han and Sen He and Denise Hernandez and Yordan Hristov and Rongjie Huang and Hirofumi Inaguma and Somya Jain and Raj Janardhan and Qingyao Jia and Christopher Klaiber and Dejan Kovachev and Moneish Kumar and Hang Li and Yilei Li and Pavel Litvin and Wei Liu and Guangyao Ma and Jing Ma and Martin Ma and Xutai Ma and Lucas Mantovani and Sagar Miglani and Sreyas Mohan and Louis-Philippe Morency and Evonne Ng and Kam-Woh Ng and Tu Anh Nguyen and Amia Oberai and Benjamin Peloquin and Juan Pino and Jovan Popovic and Omid Poursaeed and Fabian Prada and Alice Rakotoarison and Rakesh Ranjan and Alexander Richard and Christophe Ropers and Safiyyah Saleem and Vasu Sharma and Alex Shcherbyna and Jia Shen and Jie Shen and Anastasis Stathopoulos and Anna Sun and Paden Tomasello and Tuan Tran and Arina Turkatenko and Bo Wan and Chao Wang and Jeff Wang and Mary Williamson and Carleigh Wood and Tao Xiang and Yilin Yang and Julien Yao and Chen Zhang and Jiemin Zhang and Xinyue Zhang and Jason Zheng and Pavlo Zhyzheria and Jan Zikes and Michael Zollhoefer},
  journal= {arXiv preprint arXiv:2506.22554},
  year   = {2025}
}
R2 v1 2026-07-01T03:37:10.698Z