中文
相关论文

相关论文: SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spat…

200 篇论文

We present an end-to-end binaural audio rendering approach (Listen2Scene) for virtual reality (VR) and augmented reality (AR) applications. We propose a novel neural-network-based binaural sound propagation method to generate acoustic…

音频与语音处理 · 电气工程与系统科学 2024-02-09 Anton Ratnarajah , Dinesh Manocha

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accomplish this, we build…

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST)…

多媒体 · 计算机科学 2025-07-15 Wenjie Tian , Xinfa Zhu , Haohe Liu , Zhixian Zhao , Zihao Chen , Chaofan Ding , Xinhan Di , Junjie Zheng , Lei Xie

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman

Ambisonics i.e., a full-sphere surround sound, is quintessential with 360-degree visual content to provide a realistic virtual reality (VR) experience. While 360-degree visual content capture gained a tremendous boost recently, the…

声音 · 计算机科学 2019-08-20 Aakanksha Rana , Cagri Ozcinar , Aljoscha Smolic

Despite recent breakthroughs in reinforcement learning (RL) and imitation learning (IL), existing algorithms fail to generalize beyond the training environments. In reality, humans can adapt to new tasks quickly by leveraging prior…

机器学习 · 计算机科学 2023-04-18 Tianshi Cao , Jingkang Wang , Yining Zhang , Sivabalan Manivasagam

People easily recognize new visual categories that are new combinations of known components. This compositional generalization capacity is critical for learning in real-world domains like vision and language because the long tail of new…

计算机视觉与模式识别 · 计算机科学 2020-11-03 Yuval Atzmon , Felix Kreuk , Uri Shalit , Gal Chechik

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we…

声音 · 计算机科学 2025-09-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

Sound field reconstruction aims to estimate pressure fields in areas lacking direct measurements. Existing techniques often rely on strong assumptions or face challenges related to data availability or the explicit modeling of physical…

音频与语音处理 · 电气工程与系统科学 2024-12-25 Stefano Damiano , Federico Miotello , Mirco Pezzoli , Alberto Bernardini , Fabio Antonacci , Augusto Sarti , Toon van Waterschoot

Generative AI has been transforming the way we interact with technology and consume content. In the next decade, AI technology will reshape how we create audio content in various media, including music, theater, films, games, podcasts, and…

声音 · 计算机科学 2024-11-25 Hao-Wen Dong

Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot learning methods focused on single-label classification…

声音 · 计算机科学 2024-09-04 Duygu Dogan , Huang Xie , Toni Heittola , Tuomas Virtanen

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model…

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous studies have…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Taegyeong Lee , Jeonghun Kang , Hyeonyu Kim , Taehwan Kim

In this paper we consider a version of the zero-shot learning problem where seen class source and target domain data are provided. The goal during test-time is to accurately predict the class label of an unseen target domain instance based…

计算机视觉与模式识别 · 计算机科学 2015-09-29 Ziming Zhang , Venkatesh Saligrama

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repetitive beat gestures…

人机交互 · 计算机科学 2025-10-21 Hendric Voss , Stefan Kopp

Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visual actions that are aligned with them, otherwise unnatural…

声音 · 计算机科学 2024-07-16 Santiago Pascual , Chunghsin Yeh , Ioannis Tsiamas , Joan Serrà

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Cheng Yuan , Jian Jiang , Kunyi Yang , Lv Wu , Rui Wang , Zi Meng , Haonan Ping , Ziyu Xu , Yifan Zhou , Wanli Song , Hesheng Wang , Yueming Jin , Qi Dou , Yutong Ban

We study universal zero-shot segmentation in this work to achieve panoptic, instance, and semantic segmentation for novel categories without any training samples. Such zero-shot segmentation ability relies on inter-class relationships in…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Shuting He , Henghui Ding , Wei Jiang