English
Related papers

Related papers: CAVER: Curious Audiovisual Exploring Robot

200 papers

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step…

Robotics · Computer Science 2024-04-16 Roberto Bigazzi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

Recent advances in multimodal learning have largely relied on pairwise contrastive objectives to align different modalities, such as text, video, and audio, in a shared embedding space. While effective in bi-modal setups, these approaches…

Artificial Intelligence · Computer Science 2025-08-19 Haochen You , Baojing Liu

Robots learn as they interact with humans. Consider a human teleoperating an assistive robot arm: as the human guides and corrects the arm's motion, the robot gathers information about the human's desired task. But how does the human know…

Robotics · Computer Science 2024-04-16 James F. Mullen , Josh Mosier , Sounak Chakrabarti , Anqi Chen , Tyler White , Dylan P. Losey

Exploration of unknown space with an autonomous mobile robot is a well-studied problem. In this work we broaden the scope of exploration, moving beyond the pure geometric goal of uncovering as much free space as possible. We believe that…

User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Zhengxian Yang , Shi Pan , Shengqi Wang , Haoxiang Wang , Li Lin , Guanjun Li , Zhengqi Wen , Borong Lin , Jianhua Tao , Tao Yu

Recently, researchers have gradually realized that in some cases, the self-supervised pre-training on large-scale Internet data is better than that of high-quality/manually labeled data sets, and multimodal/large models are better than…

Sound · Computer Science 2023-08-08 Sen Fang , Yangjian Wu , Bowen Gao , Jingwen Cai , Teik Toe Teoh

Vision-based bird's-eye-view (BEV) 3D object detection has advanced significantly in autonomous driving by offering cost-effectiveness and rich contextual information. However, existing methods often construct BEV representations by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jicheng Yuan , Manh Nguyen Duc , Qian Liu , Manfred Hauswirth , Danh Le Phuoc

We introduce SonicSense, a holistic design of hardware and software to enable rich robot object perception through in-hand acoustic vibration sensing. While previous studies have shown promising results with acoustic sensing for object…

Robotics · Computer Science 2024-10-04 Jiaxun Liu , Boyuan Chen

Interactive audio spatialization technology previously developed for video game authoring and rendering has evolved into an essential component of platforms enabling shared immersive virtual experiences for future co-presence, remote…

Sound · Computer Science 2021-09-28 Jean-Marc Jot , Rémi Audfray , Mark Hertensteiner , Brian Schmidt

Cross-platform verification, a critical undertaking in the realm of early-stage quantum computing, endeavors to characterize the similarity of two imperfect quantum devices executing identical algorithms, utilizing minimal measurements.…

Quantum Physics · Physics 2023-11-08 Yang Qian , Yuxuan Du , Zhenliang He , Min-hsiu Hsieh , Dacheng Tao

Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Changan Chen , Unnat Jain , Carl Schissler , Sebastia Vicenc Amengual Gari , Ziad Al-Halah , Vamsi Krishna Ithapu , Philip Robinson , Kristen Grauman

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Xiulong Liu , Sudipta Paul , Moitreya Chatterjee , Anoop Cherian

Understanding the physical world requires perceptual models grounded in physical laws rather than mere statistical correlations. However, existing multimodal learning frameworks, focused on vision and language, lack physical consistency and…

Artificial Intelligence · Computer Science 2025-11-26 Bo Pang , Chenxi Xu , Jierui Ren , Guoping Wang , Sheng Li

Audio-visual video segmentation~(AVVS) aims to generate pixel-level maps of sound-producing objects within image frames and ensure the maps faithfully adhere to the given audio, such as identifying and segmenting a singing person in a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Kexin Li , Zongxin Yang , Lei Chen , Yi Yang , Jun Xiao

Interactive perception enables robots to manipulate the environment and objects to bring them into states that benefit the perception process. Deformable objects pose challenges to this due to significant manipulation difficulty and…

Robots in shared spaces often move in ways that are difficult for people to interpret, placing the burden on humans to adapt. High-DoF robots exhibit motion that people read as expressive, intentionally or not, making it important to…

Robotics · Computer Science 2026-04-07 Jonathan Albert Cohen , Kye Shimizu , Allen Song , Vishnu Bharath , Kent Larson , Pattie Maes

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , Ji-Rong Wen

In this paper, we develop an online active mapping system to enable a quadruped robot to autonomously survey large physical structures. We describe the perception, planning and control modules needed to scan and reconstruct an object of…

Robotics · Computer Science 2020-02-25 Yiduo Wang , Milad Ramezani , Maurice Fallon

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Xuyang Shen , Dong Li , Jinxing Zhou , Zhen Qin , Bowen He , Xiaodong Han , Aixuan Li , Yuchao Dai , Lingpeng Kong , Meng Wang , Yu Qiao , Yiran Zhong

Perception is essential for the active interaction of physical agents with the external environment. The integration of multiple sensory modalities, such as touch and vision, enhances this perceptual process, creating a more comprehensive…

Robotics · Computer Science 2025-02-10 Enrico Donato , Egidio Falotico , Thomas George Thuruthel