English
Related papers

Related papers: VISTA: A Generative Egocentric Video Framework for…

200 papers

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Aniket Rege , Arka Sadhu , Yuliang Li , Kejie Li , Ramya Korlakai Vinayak , Yuning Chai , Yong Jae Lee , Hyo Jin Kim

Long-form video understanding represents a significant challenge within computer vision, demanding a model capable of reasoning over long multi-modal sequences. Motivated by the human cognitive process for long-form video understanding, we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Xiaohan Wang , Yuhui Zhang , Orr Zohar , Serena Yeung-Levy

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Vishakha Lall , Yisi Liu

As general intelligent agents are poised for widespread deployment in diverse households, evaluation tailored to each unique unseen 3D environment has become a critical prerequisite. However, existing benchmarks suffer from severe data…

Artificial Intelligence · Computer Science 2026-02-06 Xinyi He , Ying Yang , Chuanjian Fu , Sihan Guo , Songchun Zhu , Lifeng Fan , Zhenliang Zhang , Yujia Peng

As large language models (LLMs) become increasingly integrated into daily life, there is growing demand for AI assistants that are not only reactive but also proactive and personalized. While recent advances have pushed forward proactivity…

Computation and Language · Computer Science 2026-02-24 Jiho Kim , Junseong Choi , Woosog Chay , Daeun Kyung , Yeonsu Kwon , Yohan Jo , Edward Choi

Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential modeling techniques, particularly transformer-like…

Information Retrieval · Computer Science 2026-03-27 Zhimin Chen , Chenyu Zhao , Ka Chun Mo , Yunjiang Jiang , Jane H. Lee , Khushhall Chandra Mahajan , Ning Jiang , Kai Ren , Jinhui Li , Wen-Yun Yang

To build Video Question Answering (VideoQA) systems capable of assisting humans in daily activities, seeking answers from long-form videos with diverse and complex events is a must. Existing multi-modal VQA models achieve promising…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Difei Gao , Luowei Zhou , Lei Ji , Linchao Zhu , Yi Yang , Mike Zheng Shou

Achieving generalizable manipulation in unconstrained environments requires the robot to proactively resolve information uncertainty, i.e., the capability of active perception. However, existing methods are often confined in limited types…

Robotics · Computer Science 2026-02-05 Jialiang Li , Yi Qiao , Yunhan Guo , Changwen Chen , Wenzhao Lian

Current Graphical User Interface (GUI) agents operate primarily under a reactive paradigm: a user must provide an explicit instruction for the agent to execute a task. However, an intelligent AI assistant should be proactive, which is…

Artificial Intelligence · Computer Science 2026-03-10 Yuxiang Chai , Shunye Tang , Han Xiao , Rui Liu , Hongsheng Li

Multimodal AI Agents are AI models that have the capability of interactively and cooperatively assisting human users to solve day-to-day tasks. Augmented Reality (AR) head worn devices can uniquely improve the user experience of solving…

Artificial Intelligence · Computer Science 2025-01-17 Saptarashmi Bandyopadhyay , Vikas Bahirwani , Lavisha Aggarwal , Bhanu Guda , Lin Li , Andrea Colaco

Designing and evaluating personalized and proactive assistant agents remains challenging due to the time, cost, and ethical concerns associated with human-in-the-loop experimentation. Existing Human-Computer Interaction (HCI) methods often…

Human-Computer Interaction · Computer Science 2025-11-25 Ziyi Xuan , Yiwen Wu , Xuhai Xu , Vinod Namboodiri , Mooi Choo Chuah , Yu Yang

The analysis and use of egocentric videos for robotic tasks is made challenging by occlusion due to the hand and the visual mismatch between the human hand and a robot end-effector. In this sense, the human hand presents a nuisance.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Matthew Chang , Aditya Prakash , Saurabh Gupta

Large Vision-Language Models (LVLMs) have shown significant progress in video understanding, yet they face substantial challenges in tasks requiring precise spatiotemporal localization at the instance level. Existing methods primarily rely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Yiming Zhao , Yu Zeng , Wenxuan Huang , Zhen Fang , Qing Miao , Qisheng Su , Jiawei Zhao , Jiayin Cai , Lin Chen , Zehui Chen , Yukun Qi , Yao Hu , Xiaolong Jiang , Feng Zhao

In robotic surgery, surgeons fully engage their hands and visual attention in procedures, making it difficult to access and manipulate multimodal patient data without interrupting the workflow. To overcome this problem, we propose a…

Computation and Language · Computer Science 2025-12-19 Hyeryun Park , Byung Mo Gu , Jun Hee Lee , Byeong Hyeon Choi , Sekeun Kim , Hyun Koo Kim , Kyungsang Kim

Generative AI is reshaping education, but it also raises concerns about instability and overreliance. In programming classrooms, we aim to leverage its feedback capabilities while reinforcing the educator's role in guiding student-AI…

Human-Computer Interaction · Computer Science 2026-02-09 Gefei Zhang , Guodao Sun , Meng Xia , Ronghua Liang

Vision-based robotic policies often struggle with even minor viewpoint changes, underscoring the need for view-invariant visual representations. This challenge becomes more pronounced in real-world settings, where viewpoint variability is…

Robotics · Computer Science 2026-01-07 Youngjoon Jeong , Junha Chun , Taesup Kim

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often leading to…

Robotics · Computer Science 2025-09-10 Qi Lv , Weijie Kong , Hao Li , Jia Zeng , Zherui Qiu , Delin Qu , Haoming Song , Qizhi Chen , Xiang Deng , Jiangmiao Pang

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

Artificial Intelligence · Computer Science 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

Current vision and language tasks usually take complete visual data (e.g., raw images or videos) as input, however, practical scenarios may often consist the situations where part of the visual information becomes inaccessible due to…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Ye Zhu , Yu Wu , Yi Yang , Yan Yan