English
Related papers

Related papers: CASP-Net: Rethinking Video Saliency Prediction fro…

200 papers

The framework of visually-guided sound source separation generally consists of three parts: visual feature extraction, multimodal feature fusion, and sound signal processing. An ongoing trend in this field has been to tailor involved visual…

Sound · Computer Science 2023-06-21 Zengjie Song , Zhaoxiang Zhang

The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Yu Zeng , Yunzhi Zhuge , Huchuan Lu , Lihe Zhang , Mingyang Qian , Yizhou Yu

Efficient video recognition is a hot-spot research topic with the explosive growth of multimedia data on the Internet and mobile devices. Most existing methods select the salient frames without awareness of the class-specific saliency…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Boyang Xia , Zhihao Wang , Wenhao Wu , Haoran Wang , Jungong Han

Recent advances in deep learning have markedly improved the quality of visual-attention modelling. In this work we apply these advances to video compression. We propose a compression method that uses a saliency model to adaptively compress…

Computer Vision and Pattern Recognition · Computer Science 2019-07-25 Vitaliy Lyudvichenko , Mikhail Erofeev , Alexander Ploshkin , Dmitriy Vatolin

In this paper, we address the problem of separating individual speech signals from videos using audio-visual neural processing. Most conventional approaches utilize frame-wise matching criteria to extract shared information between…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Jiyoung Lee , Soo-Whan Chung , Sunok Kim , Hong-Goo Kang , Kwanghoon Sohn

Explainability in time series forecasting is essential for improving model transparency and supporting informed decision-making. In this work, we present CrossScaleNet, an innovative architecture that combines a patch-based cross-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ibrahim Delibasoglu , Fredrik Heintz

Acoustic Scene Classification (ASC) is a challenging task, as a single scene may involve multiple events that contain complex sound patterns. For example, a cooking scene may contain several sound sources including silverware clinking,…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-20 Weimin Wang , Weiran Wang , Ming Sun , Chao Wang

This paper presents a new way of getting high-quality saliency maps for video, using a cheaper alternative to eye-tracking data. We designed a mouse-contingent video viewing system which simulates the viewers' peripheral vision based on the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Vitaliy Lyudvichenko , Dmitriy Vatolin

Visual saliency detection tries to mimic human vision psychology which concentrates on sparse, important areas in natural image. Saliency prediction research has been traditionally based on low level features such as contrast, edge, etc.…

Computer Vision and Pattern Recognition · Computer Science 2016-05-05 Avisek Lahiri , Sourya Roy , Anirban Santara , Pabitra Mitra , Prabir Kumar Biswas

Data-driven saliency has recently gained a lot of attention thanks to the use of Convolutional Neural Networks for predicting gaze fixations. In this paper we go beyond standard approaches to saliency prediction, in which gaze maps are…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Marcella Cornia , Lorenzo Baraldi , Giuseppe Serra , Rita Cucchiara

Video classification is productive in many practical applications, and the recent deep learning has greatly improved its accuracy. However, existing works often model video frames indiscriminately, but from the view of motion, video frames…

Computer Vision and Pattern Recognition · Computer Science 2017-03-28 Yunzhen Zhao , Yuxin Peng

We propose a novel neural network architecture for visual saliency detections, which utilizes neurophysiologically plausible mechanisms for extraction of salient regions. The model has been significantly inspired by recent findings from…

Computer Vision and Pattern Recognition · Computer Science 2015-04-13 Natalia Efremova , Sergey Tarasenko

This paper presents an approach for top-down saliency detection guided by visual classification tasks. We first learn how to compute visual saliency when a specific visual task has to be accomplished, as opposed to most state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Francesca Murabito , Concetto Spampinato , Simone Palazzo , Konstantin Pogorelov , Michael Riegler

This paper proposes a novel saliency detection method by developing a deeply-supervised recurrent convolutional neural network (DSRCNN), which performs a full image-to-image saliency prediction. For saliency detection, the local, global,…

Computer Vision and Pattern Recognition · Computer Science 2016-08-19 Youbao Tang , Xiangqian Wu , Wei Bu

In saliency detection, every pixel needs contextual information to make saliency prediction. Previous models usually incorporate contexts holistically. However, for each pixel, usually only part of its context region is useful and…

Computer Vision and Pattern Recognition · Computer Science 2018-12-18 Nian Liu , Junwei Han , Ming-Hsuan Yang

We address the issue of visual saliency from three perspectives. First, we consider saliency detection as a frequency domain analysis problem. Second, we achieve this by employing the concept of {\it non-saliency}. Third, we simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2016-05-09 Jian Li , Martin Levine , Xiangjing An , Xin Xu , Hangen He

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-13 Otavio Braga , Olivier Siohan

Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, which include both…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Tianyu Yang , Yiyang Nan , Lisen Dai , Zhenwen Liang , Yapeng Tian , Xiangliang Zhang

This paper focuses on the problem of visual saliency prediction, predicting regions of an image that tend to attract human visual attention, under a constrained computational budget. We modify and test various recent efficient convolutional…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Feiyan Hu , Kevin McGuinness

Computational modeling of visual saliency has become an important research problem in recent years, with applications in video quality estimation, video compression, object tracking, retargeting, summarization, and so on. While most visual…

Multimedia · Computer Science 2016-04-26 Sayed Hossein Khatoonabadi , Ivan V. Bajic , Yufeng Shan
‹ Prev 1 3 4 5 6 7 10 Next ›