English
Related papers

Related papers: CleanS2S: Single-file Framework for Proactive Spee…

200 papers

We present a textless speech-to-speech translation (S2ST) system that can translate speech from one language into another language and can be built without the need of any text data. Different from existing work in the literature, we tackle…

Computation and Language · Computer Science 2022-05-06 Ann Lee , Hongyu Gong , Paul-Ambroise Duquenne , Holger Schwenk , Peng-Jen Chen , Changhan Wang , Sravya Popuri , Yossi Adi , Juan Pino , Jiatao Gu , Wei-Ning Hsu

A Spoken dialogue system for an unseen language is referred to as Zero resource speech. It is especially beneficial for developing applications for languages that have low digital resources. Zero resource speech synthesis is the task of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-11 Karthik Pandia D S , Anusha Prakash , Mano Ranjith Kumar , Hema A Murthy

Recent work in speech-to-speech translation (S2ST) has focused primarily on offline settings, where the full input utterance is available before any output is given. This, however, is not reasonable in many real-world scenarios. In…

Computation and Language · Computer Science 2023-06-05 Liam Dugan , Anshul Wadhawan , Kyle Spence , Chris Callison-Burch , Morgan McGuire , Victor Zordan

Recent trends in neural network based text-to-speech/speech synthesis pipelines have employed recurrent Seq2seq architectures that can synthesize realistic sounding speech directly from text characters. These systems however have complex…

Computation and Language · Computer Science 2019-03-19 Gary Wang

Contemporary conversational systems often present a significant limitation: their responses lack the emotional depth and disfluent characteristic of human interactions. This absence becomes particularly noticeable when users seek more…

Computation and Language · Computer Science 2024-04-03 Rohan Chaudhury , Mihir Godbole , Aakash Garg , Jinsil Hwaryoung Seo

Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication. Beyond accomplishing translation…

Computation and Language · Computer Science 2024-06-06 Shaolei Zhang , Qingkai Fang , Shoutao Guo , Zhengrui Ma , Min Zhang , Yang Feng

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ziqiao Peng , Yanbo Fan , Haoyu Wu , Xuan Wang , Hongyan Liu , Jun He , Zhaoxin Fan

This paper introduces DiFlow-TTS, a novel zero-shot text-to-speech (TTS) system that employs discrete flow matching for generative speech modeling. We position this work as an entry point that may facilitate further advances in this…

Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that "listening" and "eye contact" play crucial roles in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li

Text-to-image generation tasks have driven remarkable advances in diverse media applications, yet most focus on single-turn scenarios and struggle with iterative, multi-turn creative tasks. Recent dialogue-based systems attempt to bridge…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shichao Ma , Yunhe Guo , Jiahao Su , Qihe Huang , Zhengyang Zhou , Yang Wang

In this paper, we propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST). Recently proposed expressive S2ST systems have achieved impressive…

Computation and Language · Computer Science 2024-06-06 Min-Jae Hwang , Ilia Kulikov , Benjamin Peloquin , Hongyu Gong , Peng-Jen Chen , Ann Lee

Every individual carries a unique and personal life story shaped by their memories and experiences. However, these memories are often scattered and difficult to organize into a coherent narrative, a challenge that defines the task of…

Human-Computer Interaction · Computer Science 2025-09-30 Shayan Talaei , Meijin Li , Kanu Grover , James Kent Hippler , Diyi Yang , Amin Saberi

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Se Jin Park , Chae Won Kim , Hyeongseop Rha , Minsu Kim , Joanna Hong , Jeong Hun Yeo , Yong Man Ro

We study response generation for open domain conversation in chatbots. Existing methods assume that words in responses are generated from an identical vocabulary regardless of their inputs, which not only makes them vulnerable to generic…

Computation and Language · Computer Science 2017-12-01 Yu Wu , Wei Wu , Dejian Yang , Can Xu , Zhoujun Li , Ming Zhou

Physician-physician discussions of patient cases represent a rich source of clinical knowledge and reasoning that could feed AI agents to enrich and even participate in subsequent interactions. However, privacy regulations and ethical…

Computation and Language · Computer Science 2026-04-13 Beny Rubinstein , Sergio Matos

Although Sequence-to-Sequence (S2S) architectures have become state-of-the-art in speech synthesis, capable of generating outputs that approach the perceptual quality of natural samples, they are limited by a lack of flexibility when it…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Slava Shechtman , Raul Fernandez , David Haws

We propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional approaches, the proposed approach utilizes phonetic…

Multimedia · Computer Science 2022-04-07 Liyang Chen , Zhiyong Wu , Jun Ling , Runnan Li , Xu Tan , Sheng Zhao

To build open-domain chatbots that are able to use diverse communicative skills, we propose a novel framework BotsTalk, where multiple agents grounded to the specific target skills participate in a conversation to automatically annotate…

Computation and Language · Computer Science 2022-10-25 Minju Kim , Chaehyeong Kim , Yongho Song , Seung-won Hwang , Jinyoung Yeo

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity…

Sound · Computer Science 2026-04-30 Yusheng Dai , Zehua Chen , Yuxuan Jiang , Baolong Gao , Qiuhong Ke , Jianfei Cai , Jun Zhu

While current dialogue systems like ChatGPT have made significant advancements in text-based interactions, they often overlook the potential of other modalities in enhancing the overall user experience. We present FaceChat, a web-based…

Computation and Language · Computer Science 2023-03-14 Deema Alnuhait , Qingyang Wu , Zhou Yu