中文
相关论文

相关论文: Minerva-Ego: Spatiotemporal Hints for Egocentric V…

200 篇论文

Modern video understanding systems excel at tasks such as scene classification, object detection, and short video retrieval. However, as video analysis becomes increasingly central to real-world applications, there is a growing need for…

人工智能 · 计算机科学 2025-05-21 Sahil Shah , Harsh Goel , Sai Shankar Narasimhan , Minkyu Choi , S P Sharan , Oguzhan Akcin , Sandeep Chinchali

Understanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders. There is currently a lack of research in this area, and…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Minghui Wu , Chenxu Zhao , Anyang Su , Donglin Di , Tianyu Fu , Da An , Min He , Ya Gao , Meng Ma , Kun Yan , Ping Wang

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous,…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Aniket Rege , Arka Sadhu , Yuliang Li , Kejie Li , Ramya Korlakai Vinayak , Yuning Chai , Yong Jae Lee , Hyo Jin Kim

This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC VQA challenge. HD-EPIC evaluates whether a vision-language model can reason over realistic…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Zhiwei Chen , Yupeng Hu , Zixu Li , Zhiheng Fu , Guozhi Qiu , Weili Guan , Liqiang Nie

Despite exciting recent results showing vision-language systems' capacity to reason about images using natural language, their capacity for video reasoning remains under-explored. We motivate framing video reasoning as the sequential…

AI assistants that support humans in daily life are becoming increasingly feasible, driven by the rapid advancements in multimodal language models. A key challenge lies in overcoming the generic nature of these models to deliver…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Soroush Seifi , Simon Gardier , Vaggelis Dorovatas , Daniel Olmeda Reino , Rahaf Aljundi

We introduce ReXTime, a benchmark designed to rigorously test AI models' ability to perform temporal reasoning within video events. Specifically, ReXTime focuses on reasoning across time, i.e. human-like understanding when the question and…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Jr-Jen Chen , Yu-Chien Liao , Hsi-Che Lin , Yu-Chu Yu , Yen-Chun Chen , Yu-Chiang Frank Wang

Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advances, egocentric motion generation remains largely…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Ruibing Hou , Mingyue Zhou , Yuwei Gui , Mingshuang Luo , Bingpeng Ma , Hong Chang , Shiguang Shan , Xilin Chen

With the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Xilei Zhu , Huiyu Duan , Liu Yang , Yucheng Zhu , Xiongkuo Min , Guangtao Zhai , Patrick Le Callet

Video question answering (Video QA) presents a powerful testbed for human-like intelligent behaviors. The task demands new capabilities to integrate video processing, language understanding, binding abstract linguistic concepts to concrete…

计算机视觉与模式识别 · 计算机科学 2021-07-12 Long Hoang Dang , Thao Minh Le , Vuong Le , Truyen Tran

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Binjie Zhang , Mike Zheng Shou

Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation of emotional reasoning with audiovisual cues remains limited.…

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Zuyao Lin , Jianhui Zhang , Peidong Jia , Xiaoguang Zhao , Shanghang Zhang , Xingyu Chen

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Deheng Zhang , Yuqian Fu , Runyi Yang , Yang Miao , Tianwen Qian , Xu Zheng , Guolei Sun , Ajad Chhatkuli , Xuanjing Huang , Yu-Gang Jiang , Luc Van Gool , Danda Pani Paudel

Video generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models often lack robust understanding, limiting their ability to…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Xiaowei Chi , Chun-Kai Fan , Hengyuan Zhang , Xingqun Qi , Rongyu Zhang , Anthony Chen , Chi-min Chan , Wei Xue , Qifeng Liu , Shanghang Zhang , Yike Guo

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three key components: a…

Video reasoning constitutes a comprehensive assessment of a model's capabilities, as it demands robust perceptual and interpretive skills, thereby serving as a means to explore the boundaries of model performance. While recent research has…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Yudi Shi , Shangzhe Di , Qirui Chen , Qinian Wang , Jiayin Cai , Xiaolong Jiang , Yao Hu , Weidi Xie

Understanding affect is central to anticipating human behavior, yet current egocentric vision benchmarks largely ignore the person's emotional states that shape their decisions and actions. Existing tasks in egocentric perception focus on…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Matthias Jammot , Björn Braun , Paul Streli , Rafael Wampfler , Christian Holz

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Yang Luo , Xuanlei Zhao , Baijiong Lin , Lingting Zhu , Liyao Tang , Yuqi Liu , Ying-Cong Chen , Shengju Qian , Xin Wang , Yang You

We propose MIRA, a new benchmark designed to evaluate models in scenarios where generating intermediate visual images is essential for successful reasoning. Unlike traditional CoT methods that rely solely on text, tasks in MIRA require…

计算机视觉与模式识别 · 计算机科学 2025-11-05 Yiyang Zhou , Haoqin Tu , Zijun Wang , Zeyu Wang , Niklas Muennighoff , Fan Nie , Yejin Choi , James Zou , Chaorui Deng , Shen Yan , Haoqi Fan , Cihang Xie , Huaxiu Yao , Qinghao Ye