English
Related papers

Related papers: AssistQ: Affordance-centric Question-driven Task C…

200 papers

In this report, we present our champion solution for Ego4D EgoSchema Challenge in CVPR 2024. To deeply integrate the powerful egocentric captioning model and question reasoning model, we propose a novel Hierarchical Comprehension scheme for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Haoyu Zhang , Yuquan Xie , Yisen Feng , Zaijing Li , Meng Liu , Liqiang Nie

Detecting assistance from artificial intelligence is increasingly important as they become ubiquitous across complex tasks such as text generation, medical diagnosis, and autonomous driving. Aid detection is challenging for humans,…

Artificial Intelligence · Computer Science 2025-07-16 Tyler King , Nikolos Gurney , John H. Miller , Volkan Ustun

Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from…

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Katsuyuki Nakamura , Hiroki Ohashi , Mitsuhiro Okada

Acquiring knowledge about object interactions and affordances can facilitate scene understanding and human-robot collaboration tasks. As humans tend to use objects in many different ways depending on the scene and the objects' availability,…

Artificial Intelligence · Computer Science 2023-04-13 Alexia Toumpa , Anthony G. Cohn

We present AliMe Assist, an intelligent assistant designed for creating an innovative online shopping experience in E-commerce. Based on question answering (QA), AliMe Assist offers assistance service, customer service, and chatting…

Computation and Language · Computer Science 2018-01-17 Feng-Lin Li , Minghui Qiu , Haiqing Chen , Xiongwei Wang , Xing Gao , Jun Huang , Juwei Ren , Zhongzhou Zhao , Weipeng Zhao , Lei Wang , Guwei Jin , Wei Chu

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

Computer Vision and Pattern Recognition · Computer Science 2021-08-13 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

This work introduces a novel task, location-aware visual question generation (LocaVQG), which aims to generate engaging questions from data relevant to a particular geographical location. Specifically, we represent such location-aware…

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Sanjoy Chowdhury , Subrata Biswas , Sayan Nag , Tushar Nagarajan , Calvin Murdock , Ishwarya Ananthabhotla , Yijun Qian , Vamsi Krishna Ithapu , Dinesh Manocha , Ruohan Gao

Although First Person Vision systems can sense the environment from the user's perspective, they are generally unable to predict his intentions and goals. Since human activities can be decomposed in terms of atomic actions and interactions…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Antonino Furnari , Sebastiano Battiato , Kristen Grauman , Giovanni Maria Farinella

The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Pha Nguyen , Sailik Sengupta , Girik Malik , Arshit Gupta , Bonan Min

Augmented-reality (AR) glasses that will have access to onboard sensors and an ability to display relevant information to the user present an opportunity to provide user assistance in quotidian tasks. Many such tasks can be characterized as…

Human-Computer Interaction · Computer Science 2020-10-16 Benjamin Newman , Kevin Carlberg , Ruta Desai

Conversational Task Assistants (CTAs) guide users in performing a multitude of activities, such as making recipes. However, ensuring that interactions remain engaging, interesting, and enjoyable for CTA users is not trivial, especially for…

Computation and Language · Computer Science 2024-04-11 Nikhita Vedula , Giuseppe Castellucci , Eugene Agichtein , Oleg Rokhlenko , Shervin Malmasi

This technical report describes the EgoTask Translation approach that explores relations among a set of egocentric video tasks in the Ego4D challenge. To improve the primary task of interest, we propose to leverage existing models developed…

Computer Vision and Pattern Recognition · Computer Science 2023-02-06 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Procedural activities are sequences of key-steps aimed at achieving specific goals. They are crucial to build intelligent agents able to assist users effectively. In this context, task graphs have emerged as a human-understandable…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Luigi Seminara , Giovanni Maria Farinella , Antonino Furnari

UI task automation enables efficient task execution by simulating human interactions with graphical user interfaces (GUIs), without modifying the existing application code. However, its broader adoption is constrained by the need for…

Human-Computer Interaction · Computer Science 2025-03-19 Tian Huang , Chun Yu , Weinan Shi , Zijian Peng , David Yang , Weiqi Sun , Yuanchun Shi

This paper introduces a challenging object grasping task and proposes a self-supervised learning approach. The goal of the task is to grasp an object which is not feasible with a single parallel gripper, but only with harnessing environment…

Robotics · Computer Science 2021-04-06 Hengyue Liang , Xibai Lou , Yang Yang , Changhyun Choi

Automatic generation of textual video descriptions that are time-aligned with video content is a long-standing goal in computer vision. The task is challenging due to the difficulty of bridging the semantic gap between the visual and…

Computer Vision and Pattern Recognition · Computer Science 2018-09-25 Meera Hahn , Nataniel Ruiz , Jean-Baptiste Alayrac , Ivan Laptev , James M. Rehg

High-precision control tasks present substantial challenges for reinforcement learning (RL) algorithms, frequently resulting in suboptimal performance attributed to network approximation inaccuracies and inadequate sample quality.These…

Machine Learning · Computer Science 2025-02-05 Donghe Chen , Yubin Peng , Tengjie Zheng , Han Wang , Chaoran Qu , Lin Cheng