中文
相关论文

相关论文: Synchformer: Efficient Synchronization from Sparse…

200 篇论文

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Reuben Tan , Arijit Ray , Andrea Burns , Bryan A. Plummer , Justin Salamon , Oriol Nieto , Bryan Russell , Kate Saenko

We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify…

声音 · 计算机科学 2022-07-22 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

Speaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propose a self-supervised…

音频与语音处理 · 电气工程与系统科学 2020-02-14 Yifan Ding , Yong Xu , Shi-Xiong Zhang , Yahuan Cong , Liqiang Wang

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Chao Huang , Ruohan Gao , J. M. F. Tsang , Jan Kurcius , Cagdas Bilen , Chenliang Xu , Anurag Kumar , Sanjeel Parekh

Music can be represented in multiple forms, such as in the audio form as a recording of a performance, in the symbolic form as a computer readable score, or in the image form as a scan of the sheet music. Music synchronisation provides a…

声音 · 计算机科学 2022-06-02 Ruchit Agrawal

Lack of audio-video synchronization is a common problem during television broadcasts and video conferencing, leading to an unsatisfactory viewing experience. A widely accepted paradigm is to create an error detection mechanism that…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Akash Gupta , Rohun Tripathi , Wondong Jang

Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant challenges due to…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Fei Yu , Jun Wang

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker's view of the…

声音 · 计算机科学 2023-05-04 Roshan Sharma , Weipeng He , Ju Lin , Egor Lakomkin , Yang Liu , Kaustubh Kalgaonkar

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to…

音频与语音处理 · 电气工程与系统科学 2025-12-08 Kaidi Wang , Yi He , Wenhao Guan , Weijie Wu , Hongwu Ding , Xiong Zhang , Di Wu , Meng Meng , Jian Luan , Lin Li , Qingyang Hong

This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we model clean speech and…

音频与语音处理 · 电气工程与系统科学 2026-02-03 Yochai Yemini , Yoav Ellinson , Rami Ben-Ari , Sharon Gannot , Ethan Fetaya

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

声音 · 计算机科学 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal localization in video, whilst also requiring fine-grained…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Xingyi Zhou , Anurag Arnab , Chen Sun , Cordelia Schmid

Ultrasound tongue imaging is used to visualise the intra-oral articulators during speech production. It is utilised in a range of applications, including speech and language therapy and phonetics research. Ultrasound and speech audio are…

音频与语音处理 · 电气工程与系统科学 2021-06-01 Aciel Eshky , Joanne Cleland , Manuel Sam Ribeiro , Eleanor Sugden , Korin Richmond , Steve Renals

This study addresses the task of performing robust and reliable time-delay estimation in signals in noisy and reverberating environments. In contrast to the popular signal processing based methods, this paper proposes to transform the input…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Akshay Raina , Vipul Arora

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representation using…

计算机视觉与模式识别 · 计算机科学 2016-10-31 Yusuf Aytar , Carl Vondrick , Antonio Torralba

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa

Fingerspelling in sign language has been the means of communicating technical terms and proper nouns when they do not have dedicated sign language gestures. Automatic recognition of fingerspelling can help resolve communication barriers…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Kamala Gajurel , Cuncong Zhong , Guanghui Wang

Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules,…

声音 · 计算机科学 2025-10-13 Zhao Guo , Ziqian Ning , Guobin Ma , Lei Xie