中文
相关论文

相关论文: CASP-Net: Rethinking Video Saliency Prediction fro…

200 篇论文

Understanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essential for saliency…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Yuxin Zhu , Huiyu Duan , Kaiwei Zhang , Yucheng Zhu , Xilei Zhu , Long Teng , Xiongkuo Min , Guangtao Zhai

In many computer vision tasks, the relevant information to solve the problem at hand is mixed to irrelevant, distracting information. This has motivated researchers to design attentional models that can dynamically focus on parts of images…

计算机视觉与模式识别 · 计算机科学 2017-02-14 Loris Bazzani , Hugo Larochelle , Lorenzo Torresani

Patient trajectories from electronic health records are widely used to estimate conditional average potential outcomes (CAPOs) of treatments over time, which then allows to personalize care. Yet, existing neural methods for this purpose…

机器学习 · 计算机科学 2025-02-19 Konstantin Hess , Stefan Feuerriegel

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

计算与语言 · 计算机科学 2025-04-11 Lakshmipathi Balaji , Karan Singla

Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect human preferences. This task is inherently difficult due to strong annotation subjectivity…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Omer Tariq , Syed Muhammad Raza , Jeongbae Son

Given a question-image input, the Visual Commonsense Reasoning (VCR) model can predict an answer with the corresponding rationale, which requires inference ability from the real world. The VCR task, which calls for exploiting the…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Xuejiao Tang , Wenbin Zhang

Contexts play an important role in the saliency detection task. However, given a context region, not all contextual information is helpful for the final task. In this paper, we propose a novel pixel-wise contextual attention network, i.e.,…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Nian Liu , Junwei Han , Ming-Hsuan Yang

Deep saliency prediction algorithms complement the object recognition features, they typically rely on additional information, such as scene context, semantic relationships, gaze direction, and object dissimilarity. However, none of these…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Bahar Aydemir , Ludo Hoffstetter , Tong Zhang , Mathieu Salzmann , Sabine Süsstrunk

Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and…

We propose a novel deep training algorithm for joint representation of audio and visual information which consists of a single stream network (SSNet) coupled with a novel loss function to learn a shared deep latent space representation of…

计算机视觉与模式识别 · 计算机科学 2019-09-20 Shah Nawaz , Muhammad Kamran Janjua , Ignazio Gallo , Arif Mahmood , Alessandro Calefati

Since the early 2000s, computational visual saliency has been a very active research area. Each year, more and more new models are published in the main computer vision conferences. Nowadays, one of the big challenges is to find a way to…

计算机视觉与模式识别 · 计算机科学 2013-07-23 Nicolas Riche , Matthieu Duvinage , Matei Mancas , Bernard Gosselin , Thierry Dutoit

This paper reviews the Challenge on Video Saliency Prediction at AIM 2024. The goal of the participants was to develop a method for predicting accurate saliency maps for the provided set of video sequences. Saliency maps are widely…

With their increase in performance, neural network architectures also become more complex, necessitating explainability. Therefore, many new and improved methods are currently emerging, which often generate so-called saliency maps in order…

机器学习 · 计算机科学 2024-12-24 Leonid Schwenke , Martin Atzmueller

We present AlignNet, a model that synchronizes videos with reference audios under non-uniform and irregular misalignments. AlignNet learns the end-to-end dense correspondence between each frame of a video and an audio. Our method is…

计算机视觉与模式识别 · 计算机科学 2020-02-13 Jianren Wang , Zhaoyuan Fang , Hang Zhao

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior…

声音 · 计算机科学 2023-08-02 Chen Liu , Peike Li , Xingqun Qi , Hu Zhang , Lincheng Li , Dadong Wang , Xin Yu

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

多媒体 · 计算机科学 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

Various linguistic and non-linguistic clues, such as excessive emphasis on a word, a shift in the tone of voice, or an awkward expression, frequently convey sarcasm. The computer vision problem of sarcasm recognition in conversation aims to…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Ananya Pandey , Dinesh Kumar Vishwakarma

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao