English
Related papers

Related papers: BARISTA: A Multi-Task Egocentric Benchmark for Com…

200 papers

Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first person perspective. In this paper we address the challenge of egocentric human action recognition by…

Computer Vision and Pattern Recognition · Computer Science 2019-05-03 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas P. J. J. Noldus , Remco C. Veltkamp

Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic scenes. While many approaches have crafted task-specific…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Junwen Xiong , Peng Zhang , Chuanyue Li , Wei Huang , Yufei Zha , Tao You

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we…

Artificial Intelligence · Computer Science 2022-03-02 Tayfun Ates , M. Samil Atesoglu , Cagatay Yigit , Ilker Kesen , Mert Kobas , Erkut Erdem , Aykut Erdem , Tilbe Goksun , Deniz Yuret

Pre-trained vision-language models provide a robust foundation for efficient transfer learning across various downstream tasks. In the field of video action recognition, mainstream approaches often introduce additional modules to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Haoxing Chen , Zizheng Huang , Yan Hong , Yanshuo Wang , Zhongcai Lyu , Zhuoer Xu , Jun Lan , Zhangxuan Gu

We propose a computational model of visual search that incorporates Bayesian interpretations of the neural mechanisms that underlie categorical perception and saccade planning. To enable meaningful comparisons between simulated and human…

Computer Vision and Pattern Recognition · Computer Science 2020-06-08 Maell Cullen , Jonathan Monney , M. Berk Mirza , Rosalyn Moran

A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over possible actions in context. However, most existing benchmarks for multimodal foundation models…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Chuhan Li , Ruilin Han , Joy Hsu , Yongyuan Liang , Rajiv Dhawan , Jiajun Wu , Ming-Hsuan Yang , Xin Eric Wang

To build a fashion recommendation system, we need to help users retrieve fashionable items that are visually similar to a particular query, for reasons ranging from searching alternatives (i.e., substitutes), to generating stylish outfits…

Information Retrieval · Computer Science 2016-04-04 Ruining He , Chunbin Lin , Julian McAuley

Recent studies have shown that the environment where people eat can affect their nutritional behaviour. In this work, we provide automatic tools for a personalised analysis of a person's health habits by the examination of daily recorded…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Estefania Talavera , Maria Leyva-Vallina , Md. Mostafa Kamal Sarker , Domenec Puig , Nicolai Petkov , Petia Radeva

The goal of building a benchmark (suite of datasets) is to provide a unified protocol for fair evaluation and thus facilitate the evolution of a specific area. Nonetheless, we point out that existing protocols of action recognition could…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Andong Deng , Taojiannan Yang , Chen Chen

Systems based on bag-of-words models from image features collected at maxima of sparse interest point operators have been used successfully for both computer visual object and action recognition tasks. While the sparse, interest-point based…

Computer Vision and Pattern Recognition · Computer Science 2013-12-31 Stefan Mathe , Cristian Sminchisescu

The inherent complexity of video understanding makes it difficult to attribute whether performance gains stem from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Geuntaek Lim , Minho Shim , Sungjune Park , Jaeyun Lee , Inwoong Lee , Taeoh Kim , Dongyoon Wee , Yukyung Choi

Real-world multimodal agents solve multi-step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhaochen Su , Jincheng Gao , Hangyu Guo , Zhenhua Liu , Lueyang Zhang , Xinyu Geng , Shijue Huang , Peng Xia , Guanyu Jiang , Cheng Wang , Yue Zhang , Yi R. Fung , Junxian He

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Junsheng Huang , Shengyu Hao , Bocheng Hu , Hongwei Wang , Gaoang Wang

We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video synthesis and dynamic scene understanding. The dataset contains nearly 1,000 short egocentric videos captured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jae Yong Lee , Daniel Scharstein , Akash Bapat , Hao Hu , Andrew Fu , Haoru Zhao , Paul Sammut , Xiang Li , Stephen Jeapes , Anik Gupta , Lior David , Saketh Madhuvarasu , Jay Girish Joshi , Jason Wither

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hindered by the lack…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Sonia Joseph , Praneet Suresh , Lorenz Hufe , Edward Stevinson , Robert Graham , Yash Vadi , Danilo Bzdok , Sebastian Lapuschkin , Lee Sharkey , Blake Aaron Richards

We introduce EASG-Bench, a question-answering benchmark for egocentric videos where the question-answering pairs are created from spatio-temporally grounded dynamic scene graphs capturing intricate relationships among actors, actions, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Ivan Rodin , Tz-Ying Wu , Kyle Min , Sharath Nittur Sridhar , Antonino Furnari , Subarna Tripathi , Giovanni Maria Farinella

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data…

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Yuping He , Yifei Huang , Guo Chen , Lidong Lu , Baoqi Pei , Jilan Xu , Tong Lu , Yoichi Sato

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Xinyu Chen , Yunxin Li , Haoyuan Shi , Baotian Hu , Wenhan Luo , Yaowei Wang , Min Zhang