English
Related papers

Related papers: Gaze-Enhanced Multimodal Turn-Taking Prediction in…

200 papers

Humans sense of distance depends on the integration of multi sensory cues. The incoming visual luminance, auditory pitch and tactile vibration could all contribute to the ability of distance judgement. This ability can be enhanced if the…

Human-Computer Interaction · Computer Science 2020-02-18 Feng Feng , Tony Stockman

Multi-party multi-turn dialogue comprehension brings unprecedented challenges on handling the complicated scenarios from multiple speakers and criss-crossed discourse relationship among speaker-aware utterances. Most existing methods deal…

Computation and Language · Computer Science 2021-09-10 Xinbei Ma , Zhuosheng Zhang , Hai Zhao

Speech-to-speech models handle turn-taking naturally but offer limited support for tool-calling or complex reasoning, while production ASR-LLM-TTS voice pipelines offer these capabilities but rely on silence timeouts, which lead to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Shangeth Rajaa

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we propose three new tasks and benchmarks to evaluate a model's…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Sindhu B Hegde , K R Prajwal , Taein Kwon , Andrew Zisserman

As human eyes serve as conduits of rich information, unveiling emotions, intentions, and even aspects of an individual's health and overall well-being, gaze tracking also enables various human-computer interaction applications, as well as…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Sikai Yang , Wan Du

For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it…

Graphics · Computer Science 2020-09-07 Youngwoo Yoon , Bok Cha , Joo-Haeng Lee , Minsu Jang , Jaeyeon Lee , Jaehong Kim , Geehyuk Lee

Gaze estimation involves predicting where the person is looking at within an image or video. Technically, the gaze information can be inferred from two different magnification levels: face orientation and eye orientation. The inference is…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Ashesh , Chu-Song Chen , Hsuan-Tien Lin

In recent years we have witnessed an increasing number of interactive systems on handheld mobile devices which utilise gaze as a single or complementary interaction modality. This trend is driven by the enhanced computational power of these…

Human-Computer Interaction · Computer Science 2023-07-04 Yaxiong Lei , Shijing He , Mohamed Khamis , Juan Ye

Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture labels is available…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Rui Hong , Jana Kosecka

Different from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the…

Computation and Language · Computer Science 2019-10-25 Zheng Lian , Jianhua Tao , Bin Liu , Jian Huang

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not…

Computation and Language · Computer Science 2025-06-27 Anne Wu , Laurent Mazaré , Neil Zeghidour , Alexandre Défossez

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio,…

The attention mechanisms in deep neural networks are inspired by human's attention that sequentially focuses on the most relevant parts of the information over time to generate prediction output. The attention parameters in those models are…

Computer Vision and Pattern Recognition · Computer Science 2017-07-20 Youngjae Yu , Jongwook Choi , Yeonhwa Kim , Kyung Yoo , Sang-Hun Lee , Gunhee Kim

A person's demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fails to capture fine-grained contextual cues such as intent,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Gabriel Sarch , Balasaravanan Thoravi Kumaravel , Sahithya Ravi , Vibhav Vineet , Andrew D. Wilson

In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details. TridentSE maintains T-F bin level representation to capture details, and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Dacheng Yin , Zhiyuan Zhao , Chuanxin Tang , Zhiwei Xiong , Chong Luo

A robot needs contextual awareness, effective speech production and complementing non-verbal gestures for successful communication in society. In this paper, we present our end-to-end system that tries to enhance the effectiveness of…

Robotics · Computer Science 2024-10-01 Bishal Ghosh , Abhinav Dhall , Ekta Singla

We introduce a novel approach to transformers that learns hierarchical representations in multiparty dialogue. First, three language modeling tasks are used to pre-train the transformers, token- and utterance-level language modeling and…

Computation and Language · Computer Science 2020-06-01 Changmao Li , Jinho D. Choi

We study gaze estimation on tablets, our key design goal is uncalibrated gaze estimation using the front-facing camera during natural use of tablets, where the posture and method of holding the tablet is not constrained. We collected the…

Computer Vision and Pattern Recognition · Computer Science 2016-07-19 Qiong Huang , Ashok Veeraraghavan , Ashutosh Sabharwal

We study the problem of imposing conversational goals/keywords on open-domain conversational agents, where the agent is required to lead the conversation to a target keyword smoothly and fast. Solving this problem enables the application of…

Computation and Language · Computer Science 2021-03-02 Peixiang Zhong , Yong Liu , Hao Wang , Chunyan Miao
‹ Prev 1 8 9 10 Next ›