English
Related papers

Related papers: Supervising Neural Attention Models for Video Capt…

200 papers

State-of-the-art approaches for image captioning require supervised training data consisting of captions with paired image data. These methods are typically unable to use unsupervised data such as textual data with no corresponding images,…

Computer Vision and Pattern Recognition · Computer Science 2017-06-27 Wenhu Chen , Aurelien Lucchi , Thomas Hofmann

Automatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Nayyer Aafaq , Naveed Akhtar , Wei Liu , Syed Zulqarnain Gilani , Ajmal Mian

Representations learned by convolutional neural networks (CNNs) exhibit a remarkable resemblance to information processing patterns observed in the primate visual system on large neuroimaging datasets collected under diverse, naturalistic…

Neurons and Cognition · Quantitative Biology 2026-03-16 Dora Gozukara , Nasir Ahmad , Katja Seeliger , Djamari Oetringer , Linda Geerligs

Many vision and language models suffer from poor visual grounding - often falling back on easy-to-learn language priors rather than basing their decisions on visual concepts in the image. In this work, we propose a generic approach called…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ramprasaath R. Selvaraju , Stefan Lee , Yilin Shen , Hongxia Jin , Shalini Ghosh , Larry Heck , Dhruv Batra , Devi Parikh

Humans acquire semantic object representations from egocentric visual streams with minimal supervision, but the underlying mechanisms remain unclear. Importantly, the visual system only processes the center of its field of view with high…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Timothy Schaumlöffel , Arthur Aubret , Gemma Roig , Jochen Triesch

Neonatal resuscitations demand an exceptional level of attentiveness from providers, who must process multiple streams of information simultaneously. Gaze strongly influences decision making; thus, understanding where a provider is looking…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Felipe Parodi , Jordan Matelsky , Alejandra Regla-Vargas , Elizabeth Foglia , Charis Lim , Danielle Weinberg , Konrad Kording , Heidi Herrick , Michael Platt

With the escalated demand of human-machine interfaces for intelligent systems, development of gaze controlled system have become a necessity. Gaze, being the non-intrusive form of human interaction, is one of the best suited approach.…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Somsukla Maiti , Akshansh Gupta

There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Minkuk Kim , Hyeon Bae Kim , Jinyoung Moon , Jinwoo Choi , Seong Tae Kim

Dense event captioning aims to detect and describe all events of interest contained in a video. Despite the advanced development in this area, existing methods tackle this task by making use of dense temporal annotations, which is…

Computer Vision and Pattern Recognition · Computer Science 2018-12-11 Xuguang Duan , Wenbing Huang , Chuang Gan , Jingdong Wang , Wenwu Zhu , Junzhou Huang

The attention mechanism provides a sequential prediction framework for learning spatial models with enhanced implicit temporal consistency. In this work, we show a systematic design (from 2D to 3D) for how conventional networks and other…

Computer Vision and Pattern Recognition · Computer Science 2021-03-05 Ruixu Liu , Ju Shen , He Wang , Chen Chen , Sen-ching Cheung , Vijayan K. Asari

This PhD. Thesis concerns the study and development of hierarchical representations for spatio-temporal visual attention modeling and understanding in video sequences. More specifically, we propose two computational models for visual…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Miguel-Ángel Fernández-Torres

Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision models tested, none…

Neurons and Cognition · Quantitative Biology 2026-05-14 Kathy Garcia , Leyla Isik

Visual perception is critically influenced by the focus of attention. Due to limited resources, it is well known that neural representations are biased in favor of attended locations. Using concurrent eye-tracking and functional Magnetic…

Computer Vision and Pattern Recognition · Computer Science 2020-10-02 Meenakshi Khosla , Gia H. Ngo , Keith Jamison , Amy Kuceyeski , Mert R. Sabuncu

When deep neural network (DNN) was first introduced to the medical image analysis community, researchers were impressed by its performance. However, it is evident now that a large number of manually labeled data is often a must to train a…

Image and Video Processing · Electrical Eng. & Systems 2022-05-16 Sheng Wang , Xi Ouyang , Tianming Liu , Qian Wang , Dinggang Shen

While exploring visual scenes, humans' scanpaths are driven by their underlying attention processes. Understanding visual scanpaths is essential for various applications. Traditional scanpath models predict the where and when of gaze shifts…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Xianyu Chen , Ming Jiang , Qi Zhao

We discuss an attentional model for simultaneous object tracking and recognition that is driven by gaze data. Motivated by theories of perception, the model consists of two interacting pathways: identity and control, intended to mirror the…

Artificial Intelligence · Computer Science 2011-09-20 Misha Denil , Loris Bazzani , Hugo Larochelle , Nando de Freitas

Deep-fake videos, generated through AI face-swapping techniques, have gained significant attention due to their potential for impactful impersonation attacks. While most research focuses on real vs. fake detection, attributing a deep-fake…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Wasim Ahmad , Yan-Tsung Peng , Yuan-Hao Chang , Gaddisa Olani Ganfure , Sarwar Khan

While neural networks with attention mechanisms have achieved superior performance on many natural language processing tasks, it remains unclear to which extent learned attention resembles human visual attention. In this paper, we propose a…

Computation and Language · Computer Science 2020-10-28 Ekta Sood , Simon Tannert , Diego Frassinelli , Andreas Bulling , Ngoc Thang Vu

Image captioning is a significant field across computer vision and natural language processing. We propose and present AIC-AB NET, a novel Attribute-Information-Combined Attention-Based Network that combines spatial attention architecture…

Computer Vision and Pattern Recognition · Computer Science 2023-07-17 Guoyun Tu , Ying Liu , Vladimir Vlassov

Nuanced understanding and the generation of detailed descriptive content for (bimanual) manipulation actions in videos is important for disciplines such as robotics, human-computer interaction, and video content analysis. This study…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Fatemeh Ziaeetabar , Reza Safabakhsh , Saeedeh Momtazi , Minija Tamosiunaite , Florentin Wörgötter