English

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

Robotics 2026-07-28 v1

Abstract

Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and π0\pi_0, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.

Cite

@article{arxiv.2607.26047,
  title  = {S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information},
  author = {Kaneyoshi Hiratsuka and Benjamin Yen and Ryosuke Kojima},
  journal= {arXiv preprint arXiv:2607.26047},
  year   = {2026}
}

Comments

Project page: https://azuma413.github.io/projects/s2a2