English
Related papers

Related papers: CAVER: Curious Audiovisual Exploring Robot

200 papers

The representation of the knowledge needed by a robot to perform complex tasks is restricted by the limitations of perception. One possible way of overcoming this situation and designing "knowledgeable" robots is to rely on the interaction…

Artificial Intelligence · Computer Science 2013-08-02 Emanuele Bastianelli , Domenico Bloisi , Roberto Capobianco , Guglielmo Gemignani , Luca Iocchi , Daniele Nardi

We explore new aspects of assistive living on smart human-robot interaction (HRI) that involve automatic recognition and online validation of speech and gestures in a natural interface, providing social features for HRI. We introduce a…

Multimedia · Computer Science 2017-11-07 A. Zlatintsi , I. Rodomagoulakis , P. Koutras , A. C. Dometios , V. Pitsikalis , C. S. Tzafestas , P. Maragos

Learning high-quality video representation has shown significant applications in computer vision and remains challenging. Previous work based on mask autoencoders such as ImageMAE and VideoMAE has proven the effectiveness of learning…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Xingjian Diao , Ming Cheng , Shitong Cheng

Tactile and visual perception are both crucial for humans to perform fine-grained interactions with their environment. Developing similar multi-modal sensing capabilities for robots can significantly enhance and expand their manipulation…

Robotics · Computer Science 2025-01-08 Binghao Huang , Yixuan Wang , Xinyi Yang , Yiyue Luo , Yunzhu Li

An immersive acoustic experience enabled by spatial audio is just as crucial as the visual aspect in creating realistic virtual environments. However, existing methods for room impulse response estimation rely either on data-demanding…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Derong Jin , Ruohan Gao

Recent years have seen immense progress in 3D computer vision and computer graphics, with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However, alongside immersive visual…

Sound · Computer Science 2024-06-12 Mason Wang , Ryosuke Sawata , Samuel Clarke , Ruohan Gao , Shangzhe Wu , Jiajun Wu

We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverberation is a well-studied problem, our approach incorporates…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Sanjoy Chowdhury , Sreyan Ghosh , Subhrajyoti Dasgupta , Anton Ratnarajah , Utkarsh Tyagi , Dinesh Manocha

This paper introduces CognitiveDog, a pioneering development of quadruped robot with Large Multi-modal Model (LMM) that is capable of not only communicating with humans verbally but also physically interacting with the environment through…

Autonomous systems face the intricate challenge of navigating unpredictable environments and interacting with external objects. The successful integration of robotic agents into real-world situations hinges on their perception capabilities,…

Robotics · Computer Science 2025-02-10 Enrico Donato , Thomas George Thuruthel , Egidio Falotico

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tingle Li , Baihe Huang , Xiaobin Zhuang , Dongya Jia , Jiawei Chen , Yuping Wang , Zhuo Chen , Gopala Anumanchipalli , Yuxuan Wang

As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal expressions are…

Machine Learning · Computer Science 2022-05-17 Anirudh Sundar , Larry Heck

While significant progress has been made on understanding hand-object interactions in computer vision, it is still very challenging for robots to perform complex dexterous manipulation. In this paper, we propose a new platform and pipeline…

Machine Learning · Computer Science 2022-07-07 Yuzhe Qin , Yueh-Hua Wu , Shaowei Liu , Hanwen Jiang , Ruihan Yang , Yang Fu , Xiaolong Wang

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

Sound · Computer Science 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However,…

Robotics · Computer Science 2024-03-04 Amandine Brunetto , Sascha Hornauer , Stella X. Yu , Fabien Moutarde

Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Hugo Bohy , Minh Tran , Kevin El Haddad , Thierry Dutoit , Mohammad Soleymani

Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness in human-machine…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Tan-Hanh Pham , Hoang-Nam Le , Phu-Vinh Nguyen , Chris Ngo , Truong-Son Hy

The current approach to exploring and monitoring complex underwater ecosystems, such as coral reefs, is to conduct surveys using diver-held or static cameras, or deploying sensor buoys. These approaches often fail to capture the full…

The goal of the audio-visual segmentation (AVS) task is to segment the sounding objects in the video frames using audio cues. However, current fusion-based methods have the performance limitations due to the small receptive field of…

Sound · Computer Science 2023-07-26 Jinxiang Liu , Chen Ju , Chaofan Ma , Yanfeng Wang , Yu Wang , Ya Zhang

Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categories. Existing methods leverage the rich semantic knowledge of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Yongqi Wang , Xinxiao Wu , Shuo Yang

\textbf{BEAVR} is an open-source, bimanual, multi-embodiment Virtual Reality (VR) teleoperation system for robots, designed to unify real-time control, data recording, and policy learning across heterogeneous robotic platforms. BEAVR…

Robotics · Computer Science 2025-08-14 Alejandro Posadas-Nava , Alejandro Carrasco , Richard Linares
‹ Prev 1 8 9 10 Next ›