中文
相关论文

相关论文: Learning Visual Affordance from Audio

200 篇论文

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tingle Li , Baihe Huang , Xiaobin Zhuang , Dongya Jia , Jiawei Chen , Yuping Wang , Zhuo Chen , Gopala Anumanchipalli , Yuxuan Wang

Automated Audio Captioning is a cross-modal task, generating natural language descriptions to summarize the audio clips' sound events. However, grounding the actual sound events in the given audio based on its corresponding caption has not…

声音 · 计算机科学 2021-02-24 Xuenan Xu , Heinrich Dinkel , Mengyue Wu , Kai Yu

Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from…

Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed audio. To address…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Minjae Kang , Martim Brandão

A core problem of Embodied AI is to learn object manipulation from observation, as humans do. To achieve this, it is important to localize 3D object affordance areas through observation such as images (3D affordance grounding) and…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Xinhang Wan , Dongqiang Gou , Xinwang Liu , En Zhu , Xuming He

We present SAGA, a versatile and adaptive framework for visuomotor control that can generalize across various environments, task objectives, and user specifications. To efficiently learn such capability, our key idea is to disentangle…

机器人学 · 计算机科学 2025-12-16 Kuan Fang , Yuxin Chen , Xinghao Zhu , Farzad Niroui , Lingfeng Sun , Jiuguang Wang

We address the problem of affordance reasoning in diverse scenes that appear in the real world. Affordances relate the agent's actions to their effects when taken on the surrounding objects. In our work, we take the egocentric view of the…

计算机视觉与模式识别 · 计算机科学 2018-06-18 Ching-Yao Chuang , Jiaman Li , Antonio Torralba , Sanja Fidler

Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on exploring downstream…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Muyi Sun , Yixuan Wang , Hong Wang , Chen Su , Man Zhang , Xingqun Qi , Qi Li , Zhenan Sun

Visual affordance learning is a key component for robots to understand how to interact with objects. Conventional approaches in this field rely on pre-defined objects and actions, falling short of capturing diverse interactions in realworld…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Audio-visual automatic speech recognition (AV-ASR) introduces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the speaker's mouth. The use of the video signal requires…

计算机视觉与模式识别 · 计算机科学 2021-09-21 Dmitriy Serdyuk , Otavio Braga , Olivier Siohan

In recent years, advancements in representation learning and language models have propelled Automated Captioning (AC) to new heights, enabling the generation of human-level descriptions. Leveraging these advancements, we propose AVCap, an…

音频与语音处理 · 电气工程与系统科学 2024-07-12 Jongsuk Kim , Jiwon Shin , Junmo Kim

Visual affordance segmentation identifies image regions of an object an agent can interact with. Existing methods re-use and adapt learning-based architectures for semantic segmentation to the affordance segmentation task and evaluate on…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Tommaso Apicella , Alessio Xompero , Paolo Gastaldo , Andrea Cavallaro

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Short-Term object-interaction Anticipation consists of detecting the location of the next-active objects, the noun and verb categories of the interaction, and the time to contact from the observation of egocentric video. This ability is…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Lorenzo Mur-Labadia , Ruben Martinez-Cantin , Josechu Guerrero , Giovanni Maria Farinella , Antonino Furnari

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduces a novel…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Yiyang Su , Ali Vosoughi , Shijian Deng , Yapeng Tian , Chenliang Xu

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform additive fusion…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Jinbae Seo , Hyeongjun Kwon , Kwonyoung Kim , Jiyoung Lee , Kwanghoon Sohn

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

计算与语言 · 计算机科学 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

With the recent advancements in Artificial Intelligence (AI), Intelligent Virtual Assistants (IVA) such as Alexa, Google Home, etc., have become a ubiquitous part of many homes. Currently, such IVAs are mostly audio-based, but going…

多媒体 · 计算机科学 2019-12-27 Shachi H Kumar , Eda Okur , Saurav Sahay , Jonathan Huang , Lama Nachman

Affordance grounding aims to localize the interaction regions for the manipulated objects in the scene image according to given instructions. A critical challenge in affordance grounding is that the embodied agent should understand human…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Changmao Chen , Yuren Cong , Zhen Kan