English
Related papers

Related papers: ViSAGE @ NTIRE 2026 Challenge on Video Saliency Pr…

200 papers

Integrating human perceptual priors into the training of neural networks has been shown to raise model generalization, serve as an effective regularizer, and align models with human expertise for applications in high-risk domains. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Colton R. Crum , Christopher Sweet , Adam Czajka

We introduce SalGAN, a deep convolutional neural network for visual saliency prediction trained with adversarial examples. The first stage of the network consists of a generator model whose weights are learned by back-propagation computed…

Computer Vision and Pattern Recognition · Computer Science 2018-07-03 Junting Pan , Cristian Canton Ferrer , Kevin McGuinness , Noel E. O'Connor , Jordi Torres , Elisa Sayrol , Xavier Giro-i-Nieto

Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 MinJu Jeon , Si-Woo Kim , Ye-Chan Kim , HyunGee Kim , Dong-Jin Kim

We propose a novel image retrieval framework for visual saliency detection using information about salient objects contained within bounding box annotations for similar images. For each test image, we train a customized SVM from similar…

Computer Vision and Pattern Recognition · Computer Science 2017-09-26 Shuang Li , Peter Mathews

This paper introduces a novel dataset for video enhancement and studies the state-of-the-art methods of the NTIRE 2021 challenge on quality enhancement of compressed video. The challenge is the first NTIRE challenge in this direction, with…

Image and Video Processing · Electrical Eng. & Systems 2021-05-04 Ren Yang , Radu Timofte

Video classification is productive in many practical applications, and the recent deep learning has greatly improved its accuracy. However, existing works often model video frames indiscriminately, but from the view of motion, video frames…

Computer Vision and Pattern Recognition · Computer Science 2017-03-28 Yunzhen Zhao , Yuxin Peng

Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect, while,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Chenglizhao Chen , Mengke Song , Wenfeng Song , Li Guo , Muwei Jian

Substantial research has been done in saliency modeling to develop intelligent machines that can perceive and interpret their surroundings. But existing models treat videos as merely image sequences excluding any audio information, unable…

Image and Video Processing · Electrical Eng. & Systems 2023-02-27 Maryam Qamar Butt , Anis Ur Rahman

In this paper we introduce ViSiL, a Video Similarity Learning architecture that considers fine-grained Spatio-Temporal relations between pairs of videos -- such relations are typically lost in previous video retrieval approaches that embed…

Computer Vision and Pattern Recognition · Computer Science 2019-08-21 Giorgos Kordopatis-Zilos , Symeon Papadopoulos , Ioannis Patras , Ioannis Kompatsiaris

The digital media landscape has seen a pervasive shift toward short-form video advertising on TV, social media and e-commerce platforms. The present study focuses on deep saliency prediction for short-form video advertising. Deep saliency…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Jianping Ye , Michel Wedel

Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic scenes. While many approaches have crafted task-specific…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Junwen Xiong , Peng Zhang , Chuanyue Li , Wei Huang , Yufei Zha , Tao You

The spherical domain representation of 360 video/image presents many challenges related to the storage, processing, transmission and rendering of omnidirectional videos (ODV). Models of human visual attention can be used so that only a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Yasser Dahou , Marouane Tliba , Kevin McGuinness , Noel O'Connor

In this report, we present the winning solution that achieved the 1st place in the Complex Video Reasoning & Robustness Evaluation Challenge 2025. This challenge evaluates the ability to generate accurate natural language answers to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Umihiro Kamoto , Tatsuya Ishibashi , Noriyuki Kugo

This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built upon a plain transformer architecture without task-specific architectural modifications,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Zhu Yu , Jingnan Gao , Runmin Zhang , Lingteng Qiu , Zhengyi Zhao , Rui Peng , Yichao Yan , Kejie Qiu , Siyu Zhu , Si-Yuan Cao , Hui-Liang Shen

Recently, video streams have occupied a large proportion of Internet traffic, most of which contain human faces. Hence, it is necessary to predict saliency on multiple-face videos, which can provide attention cues for many content based…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Yufan Liu , Minglang Qiao , Mai Xu , Bing Li , Weiming Hu , Ali Borji

TASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Kyle Min , Jason J. Corso

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Samyak Jain , Pradeep Yarlagadda , Shreyank Jyoti , Shyamgopal Karthik , Ramanathan Subramanian , Vineet Gandhi

The popularity of immersive videos has prompted extensive research into neural adaptive tile-based streaming to optimize video transmission over networks with limited bandwidth. However, the diversity of users' viewing patterns and Quality…

Networking and Internet Architecture · Computer Science 2024-10-28 Duo Wu , Panlong Wu , Miao Zhang , Fangxin Wang

This paper focuses on the problem of visual saliency prediction, predicting regions of an image that tend to attract human visual attention, under a constrained computational budget. We modify and test various recent efficient convolutional…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Feiyan Hu , Kevin McGuinness

There has been profound progress in visual saliency thanks to the deep learning architectures, however, there still exist three major challenges that hinder the detection performance for scenes with complex compositions, multiple salient…

Computer Vision and Pattern Recognition · Computer Science 2017-08-16 Jing Zhang , Yuchao Dai , Fatih Porikli , Mingyi He