中文
相关论文

相关论文: Towards Generating Ambisonics Using Audio-Visual C…

200 篇论文

Visual-audio navigation (VAN) is attracting more and more attention from the robotic community due to its broad applications, \emph{e.g.}, household robots and rescue robots. In this task, an embodied agent must search for and navigate to…

机器人学 · 计算机科学 2023-06-22 Hongcheng Wang , Yuxuan Wang , Fangwei Zhong , Mingdong Wu , Jianwei Zhang , Yizhou Wang , Hao Dong

To open up new possibilities to assess the multimodal perceptual quality of omnidirectional media formats, we proposed a novel open source 360 audiovisual (AV) quality dataset. The dataset consists of high-quality 360 video clips in…

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The…

Accurately localizing 3D sound sources and estimating their semantic labels -- where the sources may not be visible, but are assumed to lie on the physical surface of objects in the scene -- have many real applications, including detecting…

声音 · 计算机科学 2024-12-31 Yuhang He , Sangyun Shin , Anoop Cherian , Niki Trigoni , Andrew Markham

The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches struggle to infer massive amounts of missing geometry in large…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Dehui Wang , Congsheng Xu , Rong Wei , Yue Shi , Shoufa Chen , Dingxiang Luo , Tianshuo Yang , Xiaokang Yang , Wei Sui , Yusen Qin , Rui Tang , Yao Mu

In this paper, we propose a dense depth estimation pipeline for multiview 360{\deg} images. The proposed pipeline leverages a spherical camera model that compensates for radial distortion in 360{\deg} images. The key contribution of this…

计算机视觉与模式识别 · 计算机科学 2022-03-11 Seongyeop Yang , Kunhee Kim , Yeejin Lee

Immersive audio-visual perception relies on the spatial integration of both auditory and visual information which are heterogeneous sensing modalities with different fields of reception and spatial resolution. This study investigates the…

音频与语音处理 · 电气工程与系统科学 2020-03-17 Davide Berghi , Hanne Stenzel , Marco Volino , Adrian Hilton , Philip J. B. Jackson

In today's tech-driven world, significant advancements in artificial intelligence and virtual reality have emerged. These developments drive research into exploring their intersection in the realm of soundscape. Not only do these…

声音 · 计算机科学 2025-04-11 Rima Ayoubi , Laurent Lescop , Sang Bum Park

Our brains combine vision and hearing to create a more elaborate interpretation of the world. When the visual input is insufficient, a rich panoply of sounds can be used to describe our surroundings. Since more than 1,000 hours of videos…

计算机视觉与模式识别 · 计算机科学 2019-12-30 Rohan Mahadev , Hongyu Lu

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Chao Huang , Ruohan Gao , J. M. F. Tsang , Jan Kurcius , Cagdas Bilen , Chenliang Xu , Anurag Kumar , Sanjeel Parekh

This paper presents a new technique for the virtual reality (VR) visu-alization of complex volume images obtained from computer tomography (CT) and Magnetic Resonance Imaging (MRI) by combining three-dimensional (3D) mesh processing and…

多媒体 · 计算机科学 2023-05-02 Iva Vasic , Roberto Pierdicca , Emanuele Frontoni , Bata Vasic

The commercialization of Virtual Reality (VR) headsets has made immersive and 360-degree video streaming the subject of intense interest in the industry and research communities. While the basic principles of video streaming are the same,…

多媒体 · 计算机科学 2021-02-17 Federico Chiariotti

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Daniel Adebi , Sagnik Majumder , Kristen Grauman

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the…

计算机视觉与模式识别 · 计算机科学 2020-06-15 Karren Yang , Bryan Russell , Justin Salamon

360{\deg} images can provide an omnidirectional field of view which is important for stable and long-term scene perception. In this paper, we explore 360{\deg} images for visual object tracking and perceive new challenges caused by large…

计算机视觉与模式识别 · 计算机科学 2023-07-28 Huajian Huang , Yinzhe Xu , Yingshu Chen , Sai-Kit Yeung

Constructing simulation scenes that are both visually and physically realistic is a problem of practical interest in domains ranging from robotics to computer vision. This problem has become even more relevant as researchers wielding large…

机器人学 · 计算机科学 2024-06-03 Zoey Chen , Aaron Walsman , Marius Memmel , Kaichun Mo , Alex Fang , Karthikeya Vemuri , Alan Wu , Dieter Fox , Abhishek Gupta

Visual events are usually accompanied by sounds in our daily lives. However, can the machines learn to correlate the visual scene and sound, as well as localize the sound source only by observing them like humans? To investigate its…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Arda Senocak , Tae-Hyun Oh , Junsik Kim , Ming-Hsuan Yang , In So Kweon

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR)…