中文
相关论文

相关论文: Synchformer: Efficient Synchronization from Sparse…

200 篇论文

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources together. We propose to…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Ruohan Gao , Rogerio Feris , Kristen Grauman

With the exponential growth of video content, the need for automated video highlight detection to extract key moments or highlights from lengthy videos has become increasingly pressing. This technology has the potential to enhance user…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Zahidul Islam , Sujoy Paul , Mrigank Rochan

In this study, we present an efficient and effective approach for achieving temporally consistent synthetic-to-real video translation in videos of varying lengths. Our method leverages off-the-shelf conditional image diffusion models,…

计算机视觉与模式识别 · 计算机科学 2023-05-31 Ernie Chu , Shuo-Yen Lin , Jun-Cheng Chen

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the…

计算机视觉与模式识别 · 计算机科学 2020-06-15 Karren Yang , Bryan Russell , Justin Salamon

Automatic transcriptions of consumer-generated multi-media content such as "Youtube" videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap…

计算与语言 · 计算机科学 2017-12-08 Abhinav Gupta , Yajie Miao , Leonardo Neves , Florian Metze

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Daniel Adebi , Sagnik Majumder , Kristen Grauman

Building efficient architecture in neural speech processing is paramount to success in keyword spotting deployment. However, it is very challenging for lightweight models to achieve noise robustness with concise neural operations. In a…

声音 · 计算机科学 2023-05-09 Dianwen Ng , Yunqi Chen , Biao Tian , Qiang Fu , Eng Siong Chng

Temporal video alignment aims to synchronize the key events like object interactions or action phase transitions in two videos. Such methods could benefit various video editing, processing, and understanding tasks. However, existing…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Ishan Rajendrakumar Dave , Fabian Caba Heilbron , Mubarak Shah , Simon Jenni

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data?…

声音 · 计算机科学 2021-03-03 Michael Neumann , Ngoc Thang Vu

In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Andrés F. Pérez , Valentina Sanguineti , Pietro Morerio , Vittorio Murino

This paper presents a task of audio-visual scene classification (SC) where input videos are classified into one of five real-life crowded scenes: 'Riot', 'Noise-Street', 'Firework-Event', 'Music-Event', and 'Sport-Atmosphere'. To this end,…

计算机视觉与模式识别 · 计算机科学 2021-12-20 Lam Pham , Dat Ngo , Phu X. Nguyen , Truong Hoang , Alexander Schindler

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

计算机视觉与模式识别 · 计算机科学 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang

Visual events are usually accompanied by sounds in our daily lives. However, can the machines learn to correlate the visual scene and sound, as well as localize the sound source only by observing them like humans? To investigate its…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Arda Senocak , Tae-Hyun Oh , Junsik Kim , Ming-Hsuan Yang , In So Kweon

Audio-visual source localization is a challenging task that aims to predict the location of visual sound sources in a video. Since collecting ground-truth annotations of sounding objects can be costly, a plethora of weakly-supervised…

声音 · 计算机科学 2022-09-21 Shentong Mo , Pedro Morgado

Video harmonization aims to adjust the foreground of a composite video to make it compatible with the background. So far, video harmonization has only received limited attention and there is no public dataset for video harmonization. In…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Xinyuan Lu , Shengyuan Huang , Li Niu , Wenyan Cong , Liqing Zhang

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environment make sounds…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Changan Chen , Ziad Al-Halah , Kristen Grauman

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

声音 · 计算机科学 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

This paper presents the results of a study conducted on the perceptual acceptability of audio-video desynchronization for sports videos. The study was conducted with 45 videos generated by applying 8 audio-video offsets on 5 source…

图像与视频处理 · 电气工程与系统科学 2022-12-06 Joshua Peter Ebenezer

We aim to bridge the gap between typical human and machine-learning environments by extending the standard framework of few-shot learning to an online, continual setting. In this setting, episodes do not have separate training and testing…

机器学习 · 计算机科学 2021-04-26 Mengye Ren , Michael L. Iuzzolino , Michael C. Mozer , Richard S. Zemel