中文
相关论文

相关论文: AssistQ: Affordance-centric Question-driven Task C…

200 篇论文

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid egomotion and…

The end-to-end learning ability of self-driving vehicles has achieved significant milestones over the last decade owing to rapid advances in deep learning and computer vision algorithms. However, as autonomous driving technology is a…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Shahin Atakishiyev , Mohammad Salameh , Housam Babiker , Randy Goebel

The great amount of datasets generated by various data sources have posed the challenge to machine learning algorithm selection and hyperparameter configuration. For a specific machine learning task, it usually takes domain experts plenty…

机器学习 · 计算机科学 2020-07-08 Tianyu Mu , Hongzhi Wang , Chunnan Wang , Zheng Liang

Recent technological advances have made lightweight, head mounted cameras both practical and affordable and products like Google Glass show first approaches to introduce the idea of egocentric (first-person) video to the mainstream.…

计算机视觉与模式识别 · 计算机科学 2015-01-14 Sven Bambach

Real-time conversational assistants for procedural tasks often depend on video input, which can be computationally expensive and compromise user privacy. For the first time, we propose a real-time conversational assistant that provides…

多媒体 · 计算机科学 2026-02-18 Rehana Mahfuz , Yinyi Guo , Erik Visser , Phanidhar Chinchili

One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent. To measure progress towards that goal, we argue for the…

Short Term object-interaction Anticipation consists in detecting the location of the next active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric video. This ability…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Lorenzo Mur Labadia , Ruben Martinez-Cantin , Jose J. Guerrero , Giovanni M. Farinella , Antonino Furnari

Egocentric videos can bring a lot of information about how humans perceive the world and interact with the environment, which can be beneficial for the analysis of human behaviour. The research in egocentric video analysis is developing…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Ivan Rodin , Antonino Furnari , Dimitrios Mavroedis , Giovanni Maria Farinella

Egocentric action anticipation consists in predicting a future action the camera wearer will perform from egocentric video. While the task has recently attracted the attention of the research community, current approaches assume that the…

计算机视觉与模式识别 · 计算机科学 2022-02-10 Ivan Rodin , Antonino Furnari , Dimitrios Mavroeidis , Giovanni Maria Farinella

Despite the significant demand for assistive technology among vulnerable groups (e.g., the elderly, children, and the disabled) in daily tasks, research into advanced AI-driven assistive solutions that genuinely accommodate their diverse…

机器人学 · 计算机科学 2024-04-16 Zhihao Cao , Zidong Wang , Siwen Xie , Anji Liu , Lifeng Fan

Affordances are key attributes of what must be perceived by an autonomous robotic agent in order to effectively interact with novel objects. Historically, the concept derives from the literature in psychology and cognitive science, where…

机器人学 · 计算机科学 2020-04-17 Paola Ardón , Èric Pairet , Katrin S. Lohan , Subramanian Ramamoorthy , Ronald P. A. Petrick

Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence. Due to the diversity of interactive affordance, the uniqueness of different individuals leads to diverse…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Hongchen Luo , Wei Zhai , Jing Zhang , Yang Cao , Dacheng Tao

Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question answering is a specific scenario of such AI-human…

计算与语言 · 计算机科学 2019-08-01 Guan-Lin Chao , Abhinav Rastogi , Semih Yavuz , Dilek Hakkani-Tür , Jindong Chen , Ian Lane

We are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Yiming Li , Ziang Cao , Andrew Liang , Benjamin Liang , Luoyao Chen , Hang Zhao , Chen Feng

Task-oriented queries (e.g., one-shot queries to play videos, order food, or call a taxi) are crucial for assessing the quality of virtual assistants, chatbots, and other large language model (LLM)-based services. However, a standard…

信息检索 · 计算机科学 2024-06-06 Keun Soo Yim

Human drivers produce a vast amount of data which could, in principle, be used to improve autonomous driving systems. Unfortunately, seemingly straightforward approaches for creating end-to-end driving models that map sensor data directly…

计算机视觉与模式识别 · 计算机科学 2020-11-10 Yi Xiao , Felipe Codevilla , Christopher Pal , Antonio M. Lopez

Lifelong language learning aims to stream learning NLP tasks while retaining knowledge of previous tasks. Previous works based on the language model and following data-free constraint approaches have explored formatting all data as "begin…

计算与语言 · 计算机科学 2022-12-27 Han Wang , Ruiliu Fu , Xuejun Zhang , Jun Zhou , Qingwei Zhao

Data wrangling tasks such as obtaining and linking data from various sources, transforming data formats, and correcting erroneous records, can constitute up to 80% of typical data engineering work. Despite the rise of machine learning and…

Embodied Question Answering (EQA) is an essential yet challenging task for robot assistants. Large vision-language models (VLMs) have shown promise for EQA, but existing approaches either treat it as static video question answering without…

机器人学 · 计算机科学 2025-08-12 Kai Cheng , Zhengyuan Li , Xingpeng Sun , Byung-Cheol Min , Amrit Singh Bedi , Aniket Bera

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. Despite these…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Joungbin An , Yunsu Park , Hyolim Kang , Seon Joo Kim