English
Related papers

Related papers: AssistQ: Affordance-centric Question-driven Task C…

200 papers

Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Towards this vision, we focus on the innovative task where, given…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yulin Zhang , Cheng Shi , Yang Wang , Sibei Yang

Live streaming commerce has become a prominent form of broadcasting in the modern era. To facilitate more efficient and convenient product promotions for streamers, we present Click-to-Ask, an AI-driven assistant for live streaming commerce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Ruizhi Yu , Keyang Zhong , Peng Liu , Qi Wu , Haoran Zhang , Yanhao Zhang , Chen Chen , Haonan Lu

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Pengzhan Sun , Junbin Xiao , Tze Ho Elden Tse , Yicong Li , Arjun Akula , Angela Yao

Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentric action anticipation task, which predicts future action…

Computer Vision and Pattern Recognition · Computer Science 2021-01-20 Yu Wu , Linchao Zhu , Xiaohan Wang , Yi Yang , Fei Wu

In this paper, we present a novel approach for learning bimanual manipulation actions from human demonstration by extracting spatial constraints between affordance regions, termed affordance constraints, of the objects involved. Affordance…

Robotics · Computer Science 2024-11-19 Björn S. Plonka , Christian Dreher , Andre Meixner , Rainer Kartmann , Tamim Asfour

Equal access to digital technologies is critical for education, employment, and social participation. However, mainstream interfaces are visually oriented, creating steep learning curves and frequent obstacles for screen reader users, and…

Human-Computer Interaction · Computer Science 2026-01-27 Nan Chen , Jing Lu , Zilong Wang , Luna K. Qiu , Siming Chen , Yuqing Yang

The concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that…

Computer Use Agents (CUAs) are designed to autonomously operate digital interfaces, yet they often fail to reliably determine whether a given task has been completed. We present an autonomous evaluation and feedback framework that uses…

Artificial Intelligence · Computer Science 2025-11-26 Marta Sumyk , Oleksandr Kosovan

Everyday devices like light bulbs and kitchen appliances are now embedded with so many features and automated behaviors that they have become complicated to actually use. While such "smart" capabilities can better support users' goals, the…

Human-Computer Interaction · Computer Science 2024-05-08 Evan King , Haoxiang Yu , Sahil Vartak , Jenna Jacob , Sangsu Lee , Christine Julien

Wheelchair-mounted robotic arms (and other assistive robots) should help their users perform everyday tasks. One way robots can provide this assistance is shared autonomy. Within shared autonomy, both the human and robot maintain control…

Robotics · Computer Science 2021-09-29 Ananth Jonnavittula , Dylan P. Losey

Assistive robot arms enable people with disabilities to conduct everyday tasks on their own. These arms are dexterous and high-dimensional; however, the interfaces people must use to control their robots are low-dimensional. Consider…

Today, users ask Large language models (LLMs) as assistants to answer queries that require external knowledge; they ask about the weather in a specific city, about stock prices, and even about where specific locations are within their…

Software Engineering · Computer Science 2023-10-06 Jieyu Zhang , Ranjay Krishna , Ahmed H. Awadallah , Chi Wang

Planning in realistic environments requires searching in large planning spaces. Affordances are a powerful concept to simplify this search, because they model what actions can be successful in a given situation. However, the classical…

Robotics · Computer Science 2021-06-24 Danfei Xu , Ajay Mandlekar , Roberto Martín-Martín , Yuke Zhu , Silvio Savarese , Li Fei-Fei

Assistive robotic arms enable users with physical disabilities to perform everyday tasks without relying on a caregiver. Unfortunately, the very dexterity that makes these arms useful also makes them challenging to teleoperate: the robot…

Robotics · Computer Science 2019-12-10 Dylan P. Losey , Krishnan Srinivasan , Ajay Mandlekar , Animesh Garg , Dorsa Sadigh

In this paper we describe and evaluate a mixed reality system that aims to augment users in task guidance applications by combining automated and unsupervised information collection with minimally invasive video guides. The result is a…

Human-Computer Interaction · Computer Science 2017-01-11 Teesid Leelasawassuk , Dima Damen , Walterio Mayol-Cuevas

Interactive object understanding, or what we can do to objects and how is a long-standing goal of computer vision. In this paper, we tackle this problem through observation of human hands in in-the-wild egocentric videos. We demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-04-11 Mohit Goyal , Sahil Modi , Rishabh Goyal , Saurabh Gupta

Due to burdensome data requirements, learning from demonstration often falls short of its promise to allow users to quickly and naturally program robots. Demonstrations are inherently ambiguous and incomplete, making correct generalization…

Machine Learning · Computer Science 2019-04-29 Wonjoon Goo , Scott Niekum

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Tushar Nagarajan , Yanghao Li , Christoph Feichtenhofer , Kristen Grauman

Developing robots that can assist humans efficiently, safely, and adaptively is crucial for real-world applications such as healthcare. While previous work often assumes a centralized system for co-optimizing human-robot interactions, we…

Robotics · Computer Science 2024-12-30 Jason Qin , Shikun Ban , Wentao Zhu , Yizhou Wang , Dimitris Samaras

We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our…

‹ Prev 1 8 9 10 Next ›