中文
相关论文

相关论文: PolySLGen: Online Multimodal Speaking-Listening Re…

200 篇论文

Interactional synchrony refers to how the speech or behavior of two or more people involved in a conversation become more finely synchronized with each other, and they can appear to behave almost in direct response to one another. Studies…

社会与信息网络 · 计算机科学 2018-07-18 Nicholas Watkins , Ifeoma Nwogu

Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing approaches usually consider extracting 3D Morphable Model…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Zesheng Wang , Alexandre Bruckert , Patrick Le Callet , Guangtao Zhai

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent methods need handcrafted features that are time-consuming to…

音频与语音处理 · 电气工程与系统科学 2020-06-23 Huirong Huang , Zhiyong Wu , Shiyin Kang , Dongyang Dai , Jia Jia , Tianxiao Fu , Deyi Tuo , Guangzhi Lei , Peng Liu , Dan Su , Dong Yu , Helen Meng

Text-to-image (T2I) generation models have significantly advanced in recent years. However, effective interaction with these models is challenging for average users due to the need for specialized prompt engineering knowledge and the…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Minbin Huang , Yanxin Long , Xinchi Deng , Ruihang Chu , Jiangfeng Xiong , Xiaodan Liang , Hong Cheng , Qinglin Lu , Wei Liu

Embodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this…

人工智能 · 计算机科学 2026-04-30 Carson Yu Liu , Gelareh Mohammadi , Yang Song , Wafa Johal

In this work, we propose a joint system combining a talking face generation system with a text-to-speech system that can generate multilingual talking face videos from only the text input. Our system can synthesize natural multilingual…

计算机视觉与模式识别 · 计算机科学 2022-05-16 Hyoung-Kyu Song , Sang Hoon Woo , Junhyeok Lee , Seungmin Yang , Hyunjae Cho , Youseong Lee , Dongho Choi , Kang-wook Kim

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

By combining voice and touch interactions, multimodal interfaces can surpass the efficiency of either modality alone. Traditional multimodal frameworks require laborious developer work to support rich multimodal commands where the user's…

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating a spoken response…

计算与语言 · 计算机科学 2024-10-04 Kentaro Mitsui , Koh Mitsuda , Toshiaki Wakatsuki , Yukiya Hono , Kei Sawada

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system…

计算机视觉与模式识别 · 计算机科学 2024-08-05 Se Jin Park , Chae Won Kim , Hyeongseop Rha , Minsu Kim , Joanna Hong , Jeong Hun Yeo , Yong Man Ro

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Xi Liu , Ying Guo , Cheng Zhen , Tong Li , Yingying Ao , Pengfei Yan

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

The popularity of image sharing on social media and the engagement it creates between users reflects the important role that visual context plays in everyday conversations. We present a novel task, Image-Grounded Conversations (IGC), in…

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

计算机视觉与模式识别 · 计算机科学 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generates expressive,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Payam Jome Yazdian , Zoe Stanley , Angelica Lim

According to the Stimulus Organism Response (SOR) theory, all human behavioral reactions are stimulated by context, where people will process the received stimulus and produce an appropriate reaction. This implies that in a specific context…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Siyang Song , Micol Spitale , Yiming Luo , Batuhan Bal , Hatice Gunes

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Spoken dialogue generation is crucial for applications like podcasts, dynamic commentary, and entertainment content, but poses significant challenges compared to single-utterance text-to-speech (TTS). Key requirements include accurate…

In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based…

声音 · 计算机科学 2025-07-08 Kaung Myat Kyaw , Jonathan Hoyin Chan