中文
相关论文

相关论文: Understanding Co-speech Gestures in-the-wild

200 篇论文

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Ruohan Gao , Changan Chen , Ziad Al-Halah , Carl Schissler , Kristen Grauman

While most machine translation systems to date are trained on large parallel corpora, humans learn language in a different way: by being grounded in an environment and interacting with other humans. In this work, we propose a communication…

计算与语言 · 计算机科学 2018-04-12 Jason Lee , Kyunghyun Cho , Jason Weston , Douwe Kiela

Systems for multimodal emotion recognition (ER) are commonly trained to extract features from different modalities (e.g., visual, audio, and textual) that are combined to predict individual basic emotions. However, compound emotions often…

Embodied conversational agents (ECA) are often designed to produce nonverbal behavior to complement or enhance their verbal communication. One such form of nonverbal behavior is co-speech gesturing, which involves movements that the agent…

人机交互 · 计算机科学 2022-03-02 Pieter Wolfert , Nicole Robinson , Tony Belpaeme

Large language models have given social robots the ability to autonomously engage in open-domain conversations. However, they are still missing a fundamental social skill: making use of the multiple modalities that carry social…

机器人学 · 计算机科学 2025-08-19 Ruben Janssens , Tony Belpaeme

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Zexu Pan , Xinyuan Qian , Shengkui Zhao , Kun Zhou , Bin Ma

In order for robots to operate effectively in homes and workplaces, they must be able to manipulate the articulated objects common within environments built for and by humans. Previous work learns kinematic models that prescribe this…

机器人学 · 计算机科学 2016-07-04 Zhengyang Wu , Mohit Bansal , Matthew R. Walter

Human-Robot Interaction (HRI) has become increasingly important as robots are being integrated into various aspects of daily life. One key aspect of HRI is gesture recognition, which allows robots to interpret and respond to human gestures…

人机交互 · 计算机科学 2024-01-10 Sandeep Reddy Sabbella , Sara Kaszuba , Francesco Leotta , Pascal Serrarens , Daniele Nardi

Investigating cooperativity of interlocutors is central in studying pragmatics of dialogue. Models of conversation that only assume cooperative agents fail to explain the dynamics of strategic conversations. Thus, we investigate the ability…

计算与语言 · 计算机科学 2022-07-18 Anthony Sicilia , Tristan Maidment , Pat Healy , Malihe Alikhani

We introduce a new multi-modal task for computer systems, posed as a combined vision-language comprehension challenge: identifying the most suitable text describing a scene, given several similar options. Accomplishing the task entails…

计算与语言 · 计算机科学 2016-12-26 Nan Ding , Sebastian Goodman , Fei Sha , Radu Soricut

Humans use multiple senses to comprehend the environment. Vision and language are two of the most vital senses since they allow us to easily communicate our thoughts and perceive the world around us. There has been a lot of interest in…

计算与语言 · 计算机科学 2026-05-13 Thong Nguyen , Yi Bin , Junbin Xiao , Leigang Qu , Yicong Li , Jay Zhangjie Wu , Cong-Duy Nguyen , See-Kiong Ng , Luu Anh Tuan

Grounded language acquisition -- learning how language-based interactions refer to the world around them -- is amajor area of research in robotics, NLP, and HCI. In practice the data used for learning consists almost entirely of textual…

For effective human-robot collaboration, it is crucial for robots to understand requests from users and ask reasonable follow-up questions when there are ambiguities. While comprehending the users' object descriptions in the requests,…

机器人学 · 计算机科学 2021-07-13 Fethiye Irmak Dogan , Gaspar I. Melsion , Iolanda Leite

This paper strives for motion-focused video-language representations. Existing methods to learn video-language representations use spatial-focused data, where identifying the objects and scene is often enough to distinguish the relevant…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Hazel Doughty , Fida Mohammad Thoker , Cees G. M. Snoek

Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Xu Cao , Pranav Virupaksha , Wenqi Jia , Bolin Lai , Fiona Ryan , Sangmin Lee , James M. Rehg

Recent works have shown Generative Adversarial Networks (GANs) to be particularly effective in image-to-image translations. However, in tasks such as body pose and hand gesture translation, existing methods usually require precise…

计算机视觉与模式识别 · 计算机科学 2019-11-11 Yahui Liu , Marco De Nadai , Gloria Zen , Nicu Sebe , Bruno Lepri

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repetitive beat gestures…

人机交互 · 计算机科学 2025-10-21 Hendric Voss , Stefan Kopp

While there has been significant progress towards modelling coherence in written discourse, the work in modelling spoken discourse coherence has been quite limited. Unlike the coherence in text, coherence in spoken discourse is also…

计算与语言 · 计算机科学 2021-01-05 Rajaswa Patil , Yaman Kumar Singla , Rajiv Ratn Shah , Mika Hama , Roger Zimmermann

We propose direct multimodal few-shot models that learn a shared embedding space of spoken words and images from only a few paired examples. Imagine an agent is shown an image along with a spoken word describing the object in the picture,…

计算与语言 · 计算机科学 2021-07-30 Leanne Nortje , Herman Kamper