English
Related papers

Related papers: Which Experimental Design is Better Suited for VQA…

200 papers

In today's society, our cognition is constantly influenced by information intake, attention switching, and task interruptions. This increases the difficulty of a given task, adding to the existing workload and leading to compromised…

Human-Computer Interaction · Computer Science 2020-10-22 Thomas Kosch

Top-down attention allows neural networks, both artificial and biological, to focus on the information most relevant for a given task. This is known to enhance performance in visual perception. But it remains unclear how attention brings…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Freddie Bickford Smith , Brett D Roads , Xiaoliang Luo , Bradley C Love

Eye movements can provide informative cues to understand human visual scan/search behavior and cognitive load during varying tasks. Visualizations of real-time gaze measures during tasks, provide an understanding of human behavior as the…

Human-Computer Interaction · Computer Science 2024-09-11 Gavindya Jayawardena , Vikas Ashok , Sampath Jayarathna

Visual question answering (VQA) has witnessed great progress since May, 2015 as a classic problem unifying visual and textual data into a system. Many enlightening VQA works explore deep into the image and question encodings and fusing…

Computer Vision and Pattern Recognition · Computer Science 2017-02-23 Yuetan Lin , Zhangyang Pang , Donghui Wang , Yueting Zhuang

Understanding how people allocate visual attention is central to Human-Computer Interaction (HCI), yet existing computational models of attention are often either descriptive, task-specific, or difficult to interpret. My dissertation…

Human-Computer Interaction · Computer Science 2026-03-03 Yunpeng Bai

To further advance driver monitoring and assistance systems, it is important to understand how drivers allocate their attention, in other words, where do they tend to look and why. Traditionally, factors affecting human visual attention…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Iuliia Kotseruba , John K. Tsotsos

Visual Question Answering (VQA) is an increasingly popular topic in deep learning research, requiring coordination of natural language processing and computer vision modules into a single architecture. We build upon the model which placed…

Computation and Language · Computer Science 2018-03-22 Jasdeep Singh , Vincent Ying , Alex Nutkiewicz

We present an empirical study of active learning for Visual Question Answering, where a deep VQA model selects informative question-image pairs from a pool and queries an oracle for answers to maximally improve its performance under a…

Computer Vision and Pattern Recognition · Computer Science 2017-11-07 Xiao Lin , Devi Parikh

Humans navigate and understand complex visual environments by subconsciously quantifying what they see, a process known as visual enumeration. However, traditional studies using flat screens fail to capture the cognitive dynamics of this…

Human-Computer Interaction · Computer Science 2025-10-08 B. Sankar , Devottama Sen , Dibakar Sen

Two prominent strategies that the human visual system uses to reduce incoming information are spatial integration and selective attention. Although spatial integration summarizes and combines information over the visual field, selective…

Neurons and Cognition · Quantitative Biology 2019-06-28 Alessandro Grillini , Remco J. Renken , Frans W. Cornelissen

Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering strategies. We explore a question decomposition strategy…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Zaid Khan , Vijay Kumar BG , Samuel Schulter , Manmohan Chandraker , Yun Fu

The potential of using gaze as an input modality in the mobile context is growing. While users often encumber themselves by carrying objects and using mobile devices while walking, the impact of encumbrance on gaze input performance remains…

Human-Computer Interaction · Computer Science 2025-12-19 Omar Namnakani , Yasmeen Abdrabou , John H. Williamson , Mohamed Khamis

Virtual Reality (VR) is increasingly used for training and demonstration purposes including a variety of applications ranging from robot learning to rehabilitation. However, the choice of input device and its visualization might influence…

Human-Computer Interaction · Computer Science 2026-02-12 Robin Beierling , Manuel Scheibl , Jonas Dech , Abhijit Vyas , Anna-Lisa Vollmer

In multi-modal reasoning tasks, such as visual question answering (VQA), there have been many modeling and training paradigms tested. Previous models propose different methods for the vision and language tasks, but which ones perform the…

Machine Learning · Computer Science 2021-03-23 Karan Samel , Zelin Zhao , Binghong Chen , Kuan Wang , Robin Luo , Le Song

In this paper, we describe our study on how humans allocate their attention during visual crowd counting. Using an eye tracker, we collect gaze behavior of human participants who are tasked with counting the number of people in crowd…

Computer Vision and Pattern Recognition · Computer Science 2020-09-29 Raji Annadi , Yupei Chen , Viresh Ranjan , Dimitris Samaras , Gregory Zelinsky , Minh Hoai

In this paper, we present an analysis of eye gaze patterns pertaining to visual cues in augmented reality (AR) for head-mounted displays (HMDs). We conducted an experimental study involving a picking and assembly task, which was guided by…

Human-Computer Interaction · Computer Science 2021-11-09 Arne Seeliger , Gerrit Merz , Christian Holz , Stefan Feuerriegel

Visual question answering (VQA) is an interesting learning setting for evaluating the abilities and shortcomings of current systems for image understanding. Many of the recently proposed VQA systems include attention or memory mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2016-11-24 Allan Jabri , Armand Joulin , Laurens van der Maaten

Prompting and steering techniques are well established in general-purpose generative AI, yet assistive visual question answering (VQA) tools for blind users still follow rigid interaction patterns with limited opportunities for…

Human-Computer Interaction · Computer Science 2026-02-20 Farnaz Zamiri Zeraati , Yang Trista Cao , Yuehan Qiao , Hal Daumé , Hernisa Kacorri

Designing public transportation cabins that effectively engage passengers and encourage more sustainable mobility options requires a deep understanding of how users from different backgrounds, visually interact with these environments. The…

Human-Computer Interaction · Computer Science 2025-01-07 Yasaman Hakiminejad , Elizabeth Pantesco , Arash Tavakoli

Modeling users' cognitive states (e.g., cognitive load and decision confidence) is essential for building adaptive AI in high-stakes decision-making. While eye tracking provides non-invasive behavioral signals correlated with cognitive…

Human-Computer Interaction · Computer Science 2026-04-03 Xin Sun , Shu Wei , Ting Pan , Yajing Wang , Jos A. Bosch , Isao Echizen , Abdallah El Ali , Saku Sugawara