English
Related papers

Related papers: Incantation: Natural Language as the Action Interf…

200 papers

Recent advances in explainable recommendations have explored the integration of language models to analyze natural language rationales for user-item interactions. Despite their potential, existing methods often rely on ID-based…

Machine Learning · Computer Science 2025-12-18 Xinshun Feng , Mingzhe Liu , Yi Qiao , Tongyu Zhu , Leilei Sun , Shuai Wang

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Lingdong Kong , Dongyue Lu , Ao Liang , Rong Li , Yuhao Dong , Tianshuai Hu , Lai Xing Ng , Wei Tsang Ooi , Benoit R. Cottereau

One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal…

Robotics · Computer Science 2025-09-17 Anna Deichler , Siyang Wang , Simon Alexanderson , Jonas Beskow

Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic…

Robotics · Computer Science 2026-04-22 Boyu Chen , Yi Chen , Lu Qiu , Jerry Bai , Yuying Ge , Yixiao Ge

This paper introduces a framework, called EMOTION, for generating expressive motion sequences in humanoid robots, enhancing their ability to engage in humanlike non-verbal communication. Non-verbal cues such as facial expressions, gestures,…

Robotics · Computer Science 2024-10-31 Peide Huang , Yuhan Hu , Nataliya Nechyporenko , Daehwa Kim , Walter Talbott , Jian Zhang

Eye-hand coordinated interaction is becoming a mainstream interaction modality in Virtual Reality (VR) user interfaces.Current paradigms for this multimodal interaction require users to learn predefined gestures and memorize multiple…

Human-Computer Interaction · Computer Science 2026-03-03 Zhimin Wang , Chenyu Gu , Feng Lu

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

This paper proposes FABG (Facial Affective Behavior Generation), an end-to-end imitation learning system for human-robot interaction, designed to generate natural and fluid facial affective behaviors. In interaction, effectively obtaining…

Robotics · Computer Science 2025-03-05 Yanghai Zhang , Changyi Liu , Keting Fu , Wenbin Zhou , Qingdu Li , Jianwei Zhang

Natural language is the most intuitive medium for us to interact with other people when expressing commands and instructions. However, using language is seldom an easy task when humans need to express their intent towards robots, since most…

Robotics · Computer Science 2022-03-28 Arthur Bucker , Luis Figueredo , Sami Haddadin , Ashish Kapoor , Shuang Ma , Rogerio Bonatti

Image Captioning is a traditional vision-and-language task that aims to generate the language description of an image. Recent studies focus on scaling up the model size and the number of training data, which significantly increase the cost…

Computation and Language · Computer Science 2023-03-14 Ziyang Luo , Zhipeng Hu , Yadong Xi , Rongsheng Zhang , Jing Ma

We provide a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives. Interactive narratives -- or text-adventure games -- are partially observable environments…

Computation and Language · Computer Science 2021-06-18 Prithviraj Ammanabrolu , Mark O. Riedl

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Bohong Chen , Yumeng Li , Yinglin Xu , Youyi Zheng , Yanlin Weng , Kun Zhou

A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through natural language. Here we study how to design artificial…

We introduce WorldGen, a system that enables the automatic creation of large-scale, interactive 3D worlds directly from text prompts. Our approach transforms natural language descriptions into traversable, fully textured environments that…

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described…

Our goal is to generate realistic human motion from natural language. Modern methods often face a trade-off between model expressiveness and text-to-motion alignment. Some align text and motion latent spaces but sacrifice expressiveness;…

Computer Vision and Pattern Recognition · Computer Science 2024-10-21 Nefeli Andreou , Xi Wang , Victoria Fernández Abrevaya , Marie-Paule Cani , Yiorgos Chrysanthou , Vicky Kalogeiton

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yixuan Ye , Xuanyu Lu , Yuxin Jiang , Yuchao Gu , Rui Zhao , Qiwei Liang , Jiachun Pan , Fengda Zhang , Weijia Wu , Alex Jinpeng Wang

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Foundation models have received much attention due to their effectiveness across a broad range of downstream applications. Though there is a big convergence in terms of architecture, most pretrained models are typically still developed for…

Computation and Language · Computer Science 2022-06-14 Yaru Hao , Haoyu Song , Li Dong , Shaohan Huang , Zewen Chi , Wenhui Wang , Shuming Ma , Furu Wei
‹ Prev 1 8 9 10 Next ›