English
Related papers

Related papers: Can You Move These Over There? An LLM-based VR Mov…

200 papers

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment the input…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Kevin Qu , Haozhe Qi , Mihai Dusmanu , Mahdi Rad , Rui Wang , Marc Pollefeys

It has been an ambition of many to control a robot for a complex task using natural language (NL). The rise of large language models (LLMs) makes it closer to coming true. However, an LLM-powered system still suffers from the ambiguity…

Robotics · Computer Science 2025-03-10 Teun van de Laar , Zengjie Zhang , Shuhao Qi , Sofie Haesaert , Zhiyong Sun

Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical…

Artificial Intelligence · Computer Science 2025-02-28 Lei Li , Sen Jia , Jianhao Wang , Zhaochong An , Jiaang Li , Jenq-Neng Hwang , Serge Belongie

Physically Assistive Robots (PARs) require personalized behaviors to ensure user safety and comfort. However, traditional preference learning methods, like exhaustive pairwise comparisons, cause severe physical and cognitive fatigue for…

Robotics · Computer Science 2026-04-03 Keshav Shankar , Dan Ding , Wei Gao

Modeling 3D objects in domains like Computer Aided Design (CAD) is time-consuming and comes with a steep learning curve needed to master the design process as well as tool complexities. In order to simplify the modeling process, we designed…

Human-Computer Interaction · Computer Science 2020-11-19 Markus Friedrich , Stefan Langer , Fabian Frey

Large language models (LLMs) have emerged as powerful and general solutions to many natural language tasks. However, many of the most important applications of language generation are interactive, where an agent has to talk to a person to…

Machine Learning · Computer Science 2023-11-10 Joey Hong , Sergey Levine , Anca Dragan

Vision-language models (VLMs) have shown powerful capabilities in visual question answering and reasoning tasks by combining visual representations with the abstract skill set large language models (LLMs) learn during pretraining. Vision,…

Artificial Intelligence · Computer Science 2023-09-01 Riley Tavassoli , Mani Amani , Reza Akhavian

This study examines the potential of utilizing Vision Language Models (VLMs) to improve the perceptual capabilities of semi-autonomous prosthetic hands. We introduce a unified benchmark for end-to-end perception and grasp inference,…

Robotics · Computer Science 2025-09-18 Ozan Karaali , Hossam Farag , Strahinja Dosen , Cedomir Stefanovic

Large language models (LLMs) pre-trained on vast internet-scale data have showcased remarkable capabilities across diverse domains. Recently, there has been escalating interest in deploying LLMs for robotics, aiming to harness the power of…

Robotics · Computer Science 2024-10-16 Yen-Jen Wang , Bike Zhang , Jianyu Chen , Koushil Sreenath

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interaction with social…

Robotics · Computer Science 2025-04-09 Hassan Ali , Philipp Allgeuer , Stefan Wermter

As more applications of large language models (LLMs) for 3D content for immersive environments emerge, it is crucial to study user behaviour to identify interaction patterns and potential barriers to guide the future design of immersive…

Human-Computer Interaction · Computer Science 2026-04-09 Junlong Chen , Jens Grubert , Per Ola Kristensson

Advances in 3D generative AI have enabled the creation of physical objects from text prompts, but challenges remain in creating objects involving multiple component types. We present a pipeline that integrates 3D generative AI with…

Comparative to conventional 2D interaction methods, virtual reality (VR) demonstrates an opportunity for unique interface and interaction design decisions. Currently, this poses a challenge when developing an accessible VR experience as…

Human-Computer Interaction · Computer Science 2024-05-14 Dr Corrie Green , Dr Yang Jiang , Dr John Isaacs , Dr Michael Heron

Virtual reality (VR) is an important new technology that is fun-damentally changing the way people experience entertainment and education content. Due to the fact that most currently available VR products are one size fits all, the…

Human-Computer Interaction · Computer Science 2019-04-18 Zhijiong Huang , Yu Zhang , Kathryn C. Quigley , Ramya Sankar , Clemence Wormser , Xinxin Mo , Allen Y. Yang

Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate for a novel framework…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Lei Yao , Yong Chen , Yuejiao Su , Yi Wang , Moyun Liu , Lap-Pui Chau

Lifestyle support through robotics is an increasingly promising field, with expectations for robots to take over or assist with chores like floor cleaning, table setting and clearing, and fetching items. The growth of AI, particularly…

Robotics · Computer Science 2024-10-23 Haru Nakajima , Jun Miura

In this paper, we extended the method proposed in [21] to enable humans to interact naturally with autonomous agents through vocal and textual conversations. Our extended method exploits the inherent capabilities of pre-trained large…

Robotics · Computer Science 2024-12-31 Linus Nwankwo , Elmar Rueckert

We address the problem of teleoperating an industrial robot manipulator via a commercially available Virtual Reality (VR) interface. Previous works on VR teleoperation for robot manipulators focus primarily on collaborative or research…

Robotics · Computer Science 2023-05-19 Eric Rosen , Devesh K. Jha

We propose a CompliantVLA-adaptor that augments the state-of-the-art Vision-Language-Action (VLA) models with vision-language model (VLM)-informed context-aware variable impedance control (VIC) to improve the safety and effectiveness of…

In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Keqin Chen , Zhao Zhang , Weili Zeng , Richong Zhang , Feng Zhu , Rui Zhao