English
Related papers

Related papers: Audio-visual Saliency for Omnidirectional Videos

200 papers

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask pairs in supervised learning fashion. This limits their…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Xiatian Zhu

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-15 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Hoirin Kim , Se-Young Yun

In recent years, the deep learning techniques have been applied to the estimation of saliency maps, which represent probability density functions of fixations when people look at the images. Although the methods of saliency-map estimation…

Computer Vision and Pattern Recognition · Computer Science 2018-07-18 Tatsuya Suzuki , Takao Yamanaka

Natural environment and our interaction with it is essentially multisensory, where we may deploy visual, tactile and/or auditory senses to perceive, learn and interact with our environment. Our objective in this study is to develop a scene…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-17 Sudarshan Ramenahalli

Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Junwen Xiong , Peng Zhang , Tao You , Chuanyue Li , Wei Huang , Yufei Zha

Omnidirectional video enables spherical stimuli with the $360 \times 180^ \circ$ viewing range. Meanwhile, only the viewport region of omnidirectional video can be seen by the observer through head movement (HM), and an even smaller region…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Chen Li , Mai Xu , Xinzhe Du , Zulin Wang

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Sihan Chen , Handong Li , Qunbo Wang , Zijia Zhao , Mingzhen Sun , Xinxin Zhu , Jing Liu

Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360{\deg} scene. With the rapid advancements in virtual/augmented reality, metaverse, and generative artificial intelligence, the demand for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Hongyu An , Xinfeng Zhang , Shijie Zhao , Li Zhang , Ruiqin Xiong

As the most fundamental scene understanding tasks, object detection and segmentation have made tremendous progress in deep learning era. Due to the expensive manual labeling cost, the annotated categories in existing datasets are often…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Chaoyang Zhu , Long Chen

Optical flow is the motion of a pixel between at least two consecutive video frames and can be estimated through an end-to-end trainable convolutional neural network. To this end, large training datasets are required to improve the accuracy…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Roman Seidel , André Apitzsch , Gangolf Hirtz

Audio-visual saliency prediction aims to mimic human visual attention by identifying salient regions in videos through the integration of both visual and auditory information. Although visual-only approaches have significantly advanced,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Kiana Hooshanfar , Alireza Hosseini , Ahmad Kalhor , Babak Nadjar Araabi

In this paper, we present a novel methodology we call MDS-ViTNet (Multi Decoder Saliency by Vision Transformer Network) for enhancing visual saliency prediction or eye-tracking. This approach holds significant potential for diverse fields,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Polezhaev Ignat , Goncharenko Igor , Iurina Natalya

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the decoder infers a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Samyak Jain , Pradeep Yarlagadda , Shreyank Jyoti , Shyamgopal Karthik , Ramanathan Subramanian , Vineet Gandhi

Dynamic emotion recognition in the wild remains challenging due to the transient nature of emotional expressions and temporal misalignment of multi-modal cues. Traditional approaches predict valence and arousal and often overlook the…

Machine Learning · Computer Science 2025-05-05 Vrushank Ahire , Kunal Shah , Mudasir Nazir Khan , Nikhil Pakhale , Lownish Rai Sookha , M. A. Ganaie , Abhinav Dhall

Omnidirectional videos that capture the entire surroundings are employed in a variety of fields such as VR applications and remote sensing. However, their wide field of view often causes unwanted objects to appear in the videos. This…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Ryosuke Seshimo , Mariko Isogawa

In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Ruohao Guo , Xianghua Ying , Yaru Chen , Dantong Niu , Guangyao Li , Liao Qu , Yanyu Qi , Jinxing Zhou , Bowei Xing , Wenzhen Yue , Ji Shi , Qixun Wang , Peiliang Zhang , Buwen Liang

We propose Unified Model of Saliency and Scanpaths (UMSS) -- a model that learns to predict visual saliency and scanpaths (i.e. sequences of eye fixations) on information visualisations. Although scanpaths provide rich information about the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Yao Wang , Mihai Bâce , Andreas Bulling

Computational modeling of visual saliency has become an important research problem in recent years, with applications in video quality estimation, video compression, object tracking, retargeting, summarization, and so on. While most visual…

Multimedia · Computer Science 2016-04-26 Sayed Hossein Khatoonabadi , Ivan V. Bajic , Yufeng Shan

Human visual attention is a complex phenomenon. A computational modeling of this phenomenon must take into account where people look in order to evaluate which are the salient locations (spatial distribution of the fixations), when they…

Computer Vision and Pattern Recognition · Computer Science 2020-05-08 Dario Zanca , Stefano Melacci , Marco Gori

Omnidirectional (or 360-degree) images and videos are emergent signals in many areas such as robotics and virtual/augmented reality. In particular, for virtual reality, they allow an immersive experience in which the user is provided with a…