中文
相关论文

相关论文: GRILLBot: An Assistant for Real-World Tasks with N…

200 篇论文

Human intelligence has the remarkable ability to adapt to new tasks and environments quickly. Starting from a very young age, humans acquire new skills and learn how to solve new tasks either by imitating the behavior of others or by…

Large language models (LLMs) like ChatGPT, exhibit powerful zero-shot and instruction-following capabilities, have catalyzed a revolutionary transformation across diverse fields, especially for open-ended tasks. While the idea is less…

人工智能 · 计算机科学 2024-02-29 Mengmei Zhang , Mingwei Sun , Peng Wang , Shen Fan , Yanhu Mo , Xiaoxiao Xu , Hong Liu , Cheng Yang , Chuan Shi

Large Language Models~(LLMs) have demonstrated capabilities across various applications but face challenges such as hallucination, limited reasoning abilities, and factual inconsistencies, especially when tackling complex, domain-specific…

This paper describes Team Delft's robot, which won the Amazon Picking Challenge 2016, including both the Picking and the Stowing competitions. The goal of the challenge is to automate pick and place operations in unstructured environments,…

The growing complexity of power system operations has created an urgent need for intelligent, automated tools to support reliable and efficient grid management. Conventional analysis tools often require significant domain expertise and…

系统与控制 · 电气工程与系统科学 2025-12-25 Yihan , Wen , Xin Chen

Intelligent and reliable task planning is a core capability for generalized robotics, requiring a descriptive domain representation that sufficiently models all object and state information for the scene. We present CLIMB, a continual…

机器人学 · 计算机科学 2024-10-18 Walker Byrnes , Miroslav Bogdanovic , Avi Balakirsky , Stephen Balakirsky , Animesh Garg

The task of predicting time and location from images is challenging and requires complex human-like puzzle-solving ability over different clues. In this work, we formalize this ability into core skills and implement them using different…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Hammad Ayyubi , Xuande Feng , Junzhang Liu , Xudong Lin , Zhecan Wang , Shih-Fu Chang

For effective human-robot interaction, it is important that a robotic assistant can forecast the next action a human will consider in a given task. Unfortunately, real-world tasks are often very long, complex, and repetitive; as a result…

计算机视觉与模式识别 · 计算机科学 2017-09-20 Tengda Han , Jue Wang , Anoop Cherian , Stephen Gould

Student commitment towards a learning recommendation is not separable from their understanding of the reasons it was recommended to them; and their ability to modify it based on that understanding. Among explainability approaches, chatbots…

人工智能 · 计算机科学 2024-01-25 Hasan Abu-Rasheed , Mohamad Hussam Abdulsalam , Christian Weber , Madjid Fathi

This paper presents BURG-Toolkit, a set of open-source tools for Benchmarking and Understanding Robotic Grasping. Our tools allow researchers to: (1) create virtual scenes for generating training data and performing grasping in simulation;…

机器人学 · 计算机科学 2022-05-30 Martin Rudorfer , Markus Suchi , Mohan Sridharan , Markus Vincze , Aleš Leonardis

We present Chirpy Cardinal, an open-domain dialogue agent, as a research platform for the 2019 Alexa Prize competition. Building an open-domain socialbot that talks to real people is challenging - such a system must meet multiple user…

Predicting the success of Conversational Task Assistants (CTA) can be critical to understand user behavior and act accordingly. In this paper, we propose TB-Rater, a Transformer model which combines conversational-flow features with user…

计算与语言 · 计算机科学 2023-09-21 Rafael Ferreira , David Semedo , João Magalhães

Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily…

人工智能 · 计算机科学 2026-05-07 Lilin Xu , Bufang Yang , Siyang Jiang , Kaiwei Liu , Kaiyuan Hou , Yuang Fan , Hongkai Chen , Zhenyu Yan , Xiaofan Jiang

The rapid adoption of Generative AI, including LLM-based chatbots like ChatGPT, has highlighted the need for accessible ways to support public understanding and AI literacy. To address this need, we introduce a game-based, interactive…

计算与语言 · 计算机科学 2026-05-21 Francesca Padovani , Malvina Nissim

We present PROGRESSOR, a novel framework that learns a task-agnostic reward function from videos, enabling policy training through goal-conditioned reinforcement learning (RL) without manual supervision. Underlying this reward is an…

机器人学 · 计算机科学 2024-11-28 Tewodros Ayalew , Xiao Zhang , Kevin Yuanbo Wu , Tianchong Jiang , Michael Maire , Matthew R. Walter

The well-known artificial intelligence-based chatbot ChatGPT-4 has become able to process image data as input in October 2023. We investigated its performance on the Test of Understanding Graphs in Kinematics to inform the physics education…

物理教育 · 物理学 2025-02-10 Giulia Polverini , Bor Gregorcic

Generalization to unseen tasks is an important ability for few-shot learners to achieve better zero-/few-shot performance on diverse tasks. However, such generalization to vision-language tasks including grounding and generation tasks has…

Manipulation tasks, like loading a dishwasher, can be seen as a sequence of spatial constraints and relationships between different objects. We aim to discover these rules from demonstrations by posing manipulation as a classification…

机器人学 · 计算机科学 2022-01-13 Yixin Lin , Austin S. Wang , Eric Undersander , Akshara Rai

A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave for 1 minute?". However, there is still no clear task…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Benita Wong , Joya Chen , You Wu , Stan Weixian Lei , Dongxing Mao , Difei Gao , Mike Zheng Shou

We present GR-2, a state-of-the-art generalist robot agent for versatile and generalizable robot manipulation. GR-2 is first pre-trained on a vast number of Internet videos to capture the dynamics of the world. This large-scale…