中文
相关论文

相关论文: Which Experimental Design is Better Suited for VQA…

200 篇论文

In today's society, our cognition is constantly influenced by information intake, attention switching, and task interruptions. This increases the difficulty of a given task, adding to the existing workload and leading to compromised…

人机交互 · 计算机科学 2020-10-22 Thomas Kosch

Top-down attention allows neural networks, both artificial and biological, to focus on the information most relevant for a given task. This is known to enhance performance in visual perception. But it remains unclear how attention brings…

计算机视觉与模式识别 · 计算机科学 2021-06-23 Freddie Bickford Smith , Brett D Roads , Xiaoliang Luo , Bradley C Love

Eye movements can provide informative cues to understand human visual scan/search behavior and cognitive load during varying tasks. Visualizations of real-time gaze measures during tasks, provide an understanding of human behavior as the…

人机交互 · 计算机科学 2024-09-11 Gavindya Jayawardena , Vikas Ashok , Sampath Jayarathna

Visual question answering (VQA) has witnessed great progress since May, 2015 as a classic problem unifying visual and textual data into a system. Many enlightening VQA works explore deep into the image and question encodings and fusing…

计算机视觉与模式识别 · 计算机科学 2017-02-23 Yuetan Lin , Zhangyang Pang , Donghui Wang , Yueting Zhuang

Understanding how people allocate visual attention is central to Human-Computer Interaction (HCI), yet existing computational models of attention are often either descriptive, task-specific, or difficult to interpret. My dissertation…

人机交互 · 计算机科学 2026-03-03 Yunpeng Bai

To further advance driver monitoring and assistance systems, it is important to understand how drivers allocate their attention, in other words, where do they tend to look and why. Traditionally, factors affecting human visual attention…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Iuliia Kotseruba , John K. Tsotsos

Visual Question Answering (VQA) is an increasingly popular topic in deep learning research, requiring coordination of natural language processing and computer vision modules into a single architecture. We build upon the model which placed…

计算与语言 · 计算机科学 2018-03-22 Jasdeep Singh , Vincent Ying , Alex Nutkiewicz

We present an empirical study of active learning for Visual Question Answering, where a deep VQA model selects informative question-image pairs from a pool and queries an oracle for answers to maximally improve its performance under a…

计算机视觉与模式识别 · 计算机科学 2017-11-07 Xiao Lin , Devi Parikh

Humans navigate and understand complex visual environments by subconsciously quantifying what they see, a process known as visual enumeration. However, traditional studies using flat screens fail to capture the cognitive dynamics of this…

人机交互 · 计算机科学 2025-10-08 B. Sankar , Devottama Sen , Dibakar Sen

Two prominent strategies that the human visual system uses to reduce incoming information are spatial integration and selective attention. Although spatial integration summarizes and combines information over the visual field, selective…

神经元与认知 · 定量生物学 2019-06-28 Alessandro Grillini , Remco J. Renken , Frans W. Cornelissen

Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering strategies. We explore a question decomposition strategy…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Zaid Khan , Vijay Kumar BG , Samuel Schulter , Manmohan Chandraker , Yun Fu

The potential of using gaze as an input modality in the mobile context is growing. While users often encumber themselves by carrying objects and using mobile devices while walking, the impact of encumbrance on gaze input performance remains…

人机交互 · 计算机科学 2025-12-19 Omar Namnakani , Yasmeen Abdrabou , John H. Williamson , Mohamed Khamis

Virtual Reality (VR) is increasingly used for training and demonstration purposes including a variety of applications ranging from robot learning to rehabilitation. However, the choice of input device and its visualization might influence…

人机交互 · 计算机科学 2026-02-12 Robin Beierling , Manuel Scheibl , Jonas Dech , Abhijit Vyas , Anna-Lisa Vollmer

In multi-modal reasoning tasks, such as visual question answering (VQA), there have been many modeling and training paradigms tested. Previous models propose different methods for the vision and language tasks, but which ones perform the…

机器学习 · 计算机科学 2021-03-23 Karan Samel , Zelin Zhao , Binghong Chen , Kuan Wang , Robin Luo , Le Song

In this paper, we describe our study on how humans allocate their attention during visual crowd counting. Using an eye tracker, we collect gaze behavior of human participants who are tasked with counting the number of people in crowd…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Raji Annadi , Yupei Chen , Viresh Ranjan , Dimitris Samaras , Gregory Zelinsky , Minh Hoai

In this paper, we present an analysis of eye gaze patterns pertaining to visual cues in augmented reality (AR) for head-mounted displays (HMDs). We conducted an experimental study involving a picking and assembly task, which was guided by…

人机交互 · 计算机科学 2021-11-09 Arne Seeliger , Gerrit Merz , Christian Holz , Stefan Feuerriegel

Visual question answering (VQA) is an interesting learning setting for evaluating the abilities and shortcomings of current systems for image understanding. Many of the recently proposed VQA systems include attention or memory mechanisms…

计算机视觉与模式识别 · 计算机科学 2016-11-24 Allan Jabri , Armand Joulin , Laurens van der Maaten

Prompting and steering techniques are well established in general-purpose generative AI, yet assistive visual question answering (VQA) tools for blind users still follow rigid interaction patterns with limited opportunities for…

人机交互 · 计算机科学 2026-02-20 Farnaz Zamiri Zeraati , Yang Trista Cao , Yuehan Qiao , Hal Daumé , Hernisa Kacorri

Designing public transportation cabins that effectively engage passengers and encourage more sustainable mobility options requires a deep understanding of how users from different backgrounds, visually interact with these environments. The…

人机交互 · 计算机科学 2025-01-07 Yasaman Hakiminejad , Elizabeth Pantesco , Arash Tavakoli

Modeling users' cognitive states (e.g., cognitive load and decision confidence) is essential for building adaptive AI in high-stakes decision-making. While eye tracking provides non-invasive behavioral signals correlated with cognitive…

人机交互 · 计算机科学 2026-04-03 Xin Sun , Shu Wei , Ting Pan , Yajing Wang , Jos A. Bosch , Isao Echizen , Abdallah El Ali , Saku Sugawara