English
Related papers

Related papers: Character design for soccer commmentary

200 papers

Advancing human-robot communication is crucial for autonomous systems operating in dynamic environments, where accurate real-time interpretation of human signals is essential. RoboCup provides a compelling scenario for testing these…

We describe our approach to create and deliver a custom voice for a conversational AI use-case. More specifically, we provide a voice for a Digital Einstein character, to enable human-computer interaction within the digital conversation…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-23 Joanna Rownicka , Kilian Sprenkamp , Antonio Tripiana , Volodymyr Gromoglasov , Timo P Kunz

The domain of 3D talking head generation has witnessed significant progress in recent years. A notable challenge in this field consists in blending speech-related motions with expression dynamics, which is primarily caused by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Federico Nocentini , Claudio Ferrari , Stefano Berretti

Communication in both human-human and human-robot interac-tion (HRI) contexts consists of verbal (speech-based) and non-verbal(facial expressions, eye gaze, gesture, body pose, etc.) components.The verbal component contains semantic and…

Robotics · Computer Science 2021-03-05 Micol Spitale , Maja J Matarić

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-23 Sefik Emre Eskimez , You Zhang , Zhiyao Duan

In the pursuit of natural language understanding, there has been a long standing interest in tracking state changes throughout narratives. Impressive progress has been made in modeling the state of transaction-centric dialogues and…

Computation and Language · Computer Science 2021-06-04 Ruochen Zhang , Carsten Eickhoff

As the phonetic and acoustic manifestations of laughter in conversation are highly diverse, laughter synthesis should be capable of accommodating such diversity while maintaining high controllability. This paper proposes a generative model…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-01 Hiroki Mori , Shunya Kimura

In this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from multi-modal sources…

Computer Vision and Pattern Recognition · Computer Science 2022-04-20 Zipeng Ye , Mengfei Xia , Ran Yi , Juyong Zhang , Yu-Kun Lai , Xuwei Huang , Guoxin Zhang , Yong-jin Liu

Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Akshay Anand , Chenxu Guo , Cheol Jun Cho , Jiachen Lian , Gopala Anumanchipalli

To improve the experiences of face-to-face conversation with avatar, this paper presents a novel conversation system. It is composed of two sequence-to-sequence models respectively for listening and speaking and a Generative Adversarial…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Zezhou Chen , Zhaoxiang Liu , Huan Hu , Jinqiang Bai , Shiguo Lian , Fuyuan Shi , Kai Wang

Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectional conversational…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-25 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that "listening" and "eye contact" play crucial roles in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li

This paper proposes a method for generating bullet comments for live-streaming games based on highlights (i.e., the exciting parts of video clips) extracted from the game content and evaluate the effect of mental health promotion. Game live…

Multimedia · Computer Science 2021-08-19 Junjie H. Xu , Yulin Cai , Zhou Fang , Pujana Paliyawan

Today, as seen in smart speakers, spoken dialogue technology is rapidly advancing to enable human-like interaction. However, current dialogue systems cannot pay attention not only to the content of speech, but also to the way of speaking…

Robotics · Computer Science 2022-10-20 Koki Inoue , Shuichiro Ogake , Hayato Kawamura , Naoki Igo

A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these components often requires extensive domain expertise and may contain…

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Determining the head orientation of a talker is not only beneficial for various speech signal processing applications, such as source localization or speech enhancement, but also facilitates intuitive voice control and interaction with…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-10 Kaspar Müller , Bilgesu Çakmak , Paul Didier , Simon Doclo , Jan Østergaard , Tobias Wolff

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interlocutor which is known as entrainment. To produce human-like…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-22 Jian Cong , Shan Yang , Na Hu , Guangzhi Li , Lei Xie , Dan Su

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi
‹ Prev 1 3 4 5 6 7 10 Next ›