English
Related papers

Related papers: Cube Padding for Weakly-Supervised Saliency Predic…

200 papers

Saliency prediction refers to the computational task of modeling overt attention. Social cues greatly influence our attention, consequently altering our eye movements and behavior. To emphasize the efficacy of such features, we present a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Fares Abawi , Tom Weber , Stefan Wermter

We present a system for converting a fully panoramic ($360^\circ$) video into a normal field-of-view (NFOV) hyperlapse for an optimal viewing experience. Our system exploits visual saliency and semantics to non-uniformly sample in space and…

Computer Vision and Pattern Recognition · Computer Science 2017-10-11 Wei-Sheng Lai , Yujia Huang , Neel Joshi , Chris Buehler , Ming-Hsuan Yang , Sing Bing Kang

Fully supervised salient object detection (SOD) methods have made considerable progress in performance, yet these models rely heavily on expensive pixel-wise labels. Recently, to achieve a trade-off between labeling burden and performance,…

Computer Vision and Pattern Recognition · Computer Science 2023-06-12 Binwei Xu , Haoran Liang , Weihua Gong , Ronghua Liang , Peng Chen

Due to difficulties in acquiring ground truth depth of equirectangular (360) images, the quality and quantity of equirectangular depth data today is insufficient to represent the various scenes in the world. Therefore, 360 depth estimation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Ilwi Yun , Hyuk-Jae Lee , Chae Eun Rhee

For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global…

Computer Vision and Pattern Recognition · Computer Science 2021-11-09 Zilin Gao , Qilong Wang , Bingbing Zhang , Qinghua Hu , Peihua Li

As the complexity of 3D digital content grows exponentially, understanding human visual attention is critical for optimizing rendering and processing resources. Therefore, reliable 3D mesh saliency ground truth (GT) is essential for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Guoquan Zheng , Jie Hao , Huiyu Duan , Long Tang , Shuo Yang , Yucheng Zhu , Yongming Han , Liang Yuan , Patrick Le Callet , Guangtao Zhai

Significant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain. To relieve the burden of data…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Wangbo Zhao , Jing Zhang , Long Li , Nick Barnes , Nian Liu , Junwei Han

In this work, we introduce VQA 360, a novel task of visual question answering on 360 images. Unlike a normal field-of-view image, a 360 image captures the entire visual content around the optical center of a camera, demanding more…

Computer Vision and Pattern Recognition · Computer Science 2020-01-13 Shih-Han Chou , Wei-Lun Chao , Wei-Sheng Lai , Min Sun , Ming-Hsuan Yang

360-degree visual content is widely shared on platforms such as YouTube and plays a central role in virtual reality, robotics, and autonomous navigation. However, consumer-grade dual-fisheye systems consistently yield imperfect panoramas…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Changha Shin , Woong Oh Cho , Seon Joo Kim

We introduce MVSplat360, a feed-forward approach for 360{\deg} novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Yuedong Chen , Chuanxia Zheng , Haofei Xu , Bohan Zhuang , Andrea Vedaldi , Tat-Jen Cham , Jianfei Cai

Gaze is an essential prompt for analyzing human behavior and attention. Recently, there has been an increasing interest in determining gaze direction from facial videos. However, video gaze estimation faces significant challenges, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Swati Jindal , Mohit Yadav , Roberto Manduchi

Although equirectangular projection (ERP) is a convenient form to store omnidirectional images (also known as 360-degree images), it is neither equal-area nor conformal, thus not friendly to subsequent visual communication. In the context…

Image and Video Processing · Electrical Eng. & Systems 2021-12-28 Mu Li , Kede Ma , Jinxing Li , David Zhang

Ultra-high-resolution 360-degree video streaming is severely constrained by the massive bandwidth required to deliver immersive experiences. Current viewport prediction techniques predominately rely on kinematics or low-level visual…

Multimedia · Computer Science 2026-01-12 Arman Nik Khah , Arvin Bahreini , Ravi Prakash

Monocular 360 depth estimation is challenging due to the inherent distortion of the equirectangular projection (ERP). This distortion causes a problem: spherical adjacent points are separated after being projected to the ERP plane,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Zidong Cao , Lin Wang

Current video representations heavily rely on learning from manually annotated video datasets which are time-consuming and expensive to acquire. We observe videos are naturally accompanied by abundant text information such as YouTube titles…

Computer Vision and Pattern Recognition · Computer Science 2021-01-29 Tianhao Li , Limin Wang

Current virtual reality (VR) headsets encounter a trade-off between high processing power and affordability. Consequently, offloading 3D rendering to remote servers helps reduce costs, battery usage, and headset weight. Maintaining network…

Systems and Control · Electrical Eng. & Systems 2024-10-04 Ali Majlesi Kopaee , Seyed Amir Hajseyedtaghia , Hossein Chitsaz

Depth estimation is an important step in many computer vision problems such as 3D reconstruction, novel view synthesis, and computational photography. Most existing work focuses on depth estimation from single frames. When applied to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Numair Khan , Eric Penner , Douglas Lanman , Lei Xiao

The objective of this work is human pose estimation in videos, where multiple frames are available. We investigate a ConvNet architecture that is able to benefit from temporal context by combining information across the multiple frames…

Computer Vision and Pattern Recognition · Computer Science 2015-11-10 Tomas Pfister , James Charles , Andrew Zisserman

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Rakshit Kothari , Shalini De Mello , Umar Iqbal , Wonmin Byeon , Seonwook Park , Jan Kautz

A well-known challenge in applying deep-learning methods to omnidirectional images is spherical distortion. In dense regression tasks such as depth estimation, where structural details are required, using a vanilla CNN layer on the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Yuyan Li , Yuliang Guo , Zhixin Yan , Xinyu Huang , Ye Duan , Liu Ren