中文
相关论文

相关论文: SuFIA: Language-Guided Augmented Dexterity for Rob…

200 篇论文

Virtual reality simulation has become a popular approach for training and assessing medical students. It offers diverse scenarios, realistic visuals, and quantitative performance metrics for objective evaluation. However, creating these…

软件工程 · 计算机科学 2023-11-27 Vladimir Poliakov , Dzmitry Tsetserukou , Emmanuel Vander Poorten

This survey organizes the intricate literature on the design and optimization of emerging structures around post-trained LMs. We refer to this overarching structure as scaffolded LMs and focus on LMs that are integrated into multi-step…

计算与语言 · 计算机科学 2025-11-05 Matthieu Lin , Jenny Sheng , Andrew Zhao , Shenzhi Wang , Yang Yue , Victor Shea Jay Huang , Huan Liu , Jun Liu , Gao Huang , Yong-Jin Liu

Ultrasonography has revolutionized non-invasive diagnostic methodologies, significantly enhancing patient outcomes across various medical domains. Despite its advancements, integrating ultrasound technology with robotic systems for…

机器人学 · 计算机科学 2024-06-19 Huan Xu , Jinlin Wu , Guanglin Cao , Zhen Chen , Zhen Lei , Hongbin Liu

How can we train an assistive human-machine interface (e.g., an electromyography-based limb prosthesis) to translate a user's raw command signals into the actions of a robot or computer when there is no prior mapping, we cannot ask the user…

机器学习 · 计算机科学 2022-09-16 Siddharth Reddy , Sergey Levine , Anca D. Dragan

The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods have advanced task-specific action alignment, they often…

机器人学 · 计算机科学 2026-03-16 Wuding Weng , Tongshu Wu , Liucheng Chen , Siyu Xie , Zheng Wang , Xing Xu , Jingkuan Song , Heng Tao Shen

In this paper, we introduce SUTRA, multilingual Large Language Model architecture capable of understanding, reasoning, and generating text in over 50 languages. SUTRA's design uniquely decouples core conceptual understanding from…

计算与语言 · 计算机科学 2024-05-14 Abhijit Bendale , Michael Sapienza , Steven Ripplinger , Simon Gibbs , Jaewon Lee , Pranav Mistry

Conversation agents powered by large language models are revolutionizing the way we interact with visual data. Recently, large vision-language models (LVLMs) have been extensively studied for both images and videos. However, these studies…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Juseong Jin , Chang Wook Jeong

In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two…

声音 · 计算机科学 2025-02-04 Dominik Wagner , Alexander Churchill , Siddharth Sigtia , Erik Marchi

The rapid progress of vision--language models (VLMs) has sparked growing interest in robotic control, where natural language can express the operation goals while visual feedback links perception to action. However, directly deploying…

机器人学 · 计算机科学 2025-11-04 Sarthak Mishra , Rishabh Dev Yadav , Avirup Das , Saksham Gupta , Wei Pan , Spandan Roy

Gestures serve as a fundamental and significant mode of non-verbal communication among humans. Deictic gestures (such as pointing towards an object), in particular, offer valuable means of efficiently expressing intent in situations where…

机器人学 · 计算机科学 2023-09-08 Li-Heng Lin , Yuchen Cui , Yilun Hao , Fei Xia , Dorsa Sadigh

The safe deployment of autonomous systems in safety-critical settings requires a paradigm that combines human expertise with AI-driven analysis, especially when anomalies are unforeseen. We introduce AURA (Autonomous Resilience Agent), a…

机器人学 · 计算机科学 2025-11-06 Markus Buchholz , Ignacio Carlucho , Yvan R. Petillot

In recent years, the rapid development of Large Language Models (LLMs) has significantly enhanced natural language understanding and human-computer interaction, creating new opportunities in the field of robotics. However, the integration…

机器人学 · 计算机科学 2026-01-06 Shenqi Lu , Liangwei Zhang

The dominant paradigm for end-to-end robot learning focuses on optimizing task-specific objectives that solve a single robotic problem such as picking up an object or reaching a target position. However, recent work on high-capacity models…

机器人学 · 计算机科学 2024-01-02 Samuel Schmidgall , Ji Woong Kim , Alan Kuntz , Ahmed Ezzat Ghazi , Axel Krieger

We present an assistance system that reasons about a human's intended actions during robot teleoperation in order to provide appropriate corrections for unintended behavior. We model the human's physical interaction with a control interface…

机器人学 · 计算机科学 2020-11-09 Deepak Gopinath , Mahdieh Nejati Javaremi , Brenna D. Argall

Ultrasound robots are increasingly used in medical diagnostics and early disease screening. However, current ultrasound robots lack the intelligence to understand human intentions and instructions, hindering autonomous ultrasound scanning.…

机器人学 · 计算机科学 2024-06-20 Huan Xu , Jinlin Wu , Guanglin Cao , Zhen Lei , Zhen Chen , Hongbin Liu

In shared autonomy, user input is combined with semi-autonomous control to achieve a common goal. The goal is often unknown ex-ante, so prior work enables agents to infer the goal from user input and assist with the task. Such methods tend…

机器学习 · 计算机科学 2018-05-24 Siddharth Reddy , Anca D. Dragan , Sergey Levine

The previous advancements in pathology image understanding primarily involved developing models tailored to specific tasks. Recent studies has demonstrated that the large vision-language model can enhance the performance of various…

人工智能 · 计算机科学 2024-08-20 Dawei Dai , Yuanhui Zhang , Long Xu , Qianlan Yang , Xiaojing Shen , Shuyin Xia , Guoyin Wang

Interpreting human intent accurately is a central challenge in human-robot interaction (HRI) and a key requirement for achieving more natural and intuitive collaboration between humans and machines. This work presents a novel multimodal HRI…

机器人学 · 计算机科学 2026-02-25 Guanting Shen , Zi Tian

Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still…

计算机视觉与模式识别 · 计算机科学 2024-08-16 Jiajie Li , Garrett Skinner , Gene Yang , Brian R Quaranto , Steven D Schwaitzberg , Peter C W Kim , Jinjun Xiong

Telementoring surgeons as they perform surgery can be essential in the treatment of patients when in situ expertise is not available. Nonetheless, expert mentors are often unavailable to provide trainees with real-time medical guidance.…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Edgar Rojas-Muñoz , Kyle Couperus , Juan Wachs