English
Related papers

Related papers: DubWise: Video-Guided Speech Duration Control in M…

200 papers

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based…

Computation and Language · Computer Science 2025-08-08 Wenqian Cui , Dianzhi Yu , Xiaoqi Jiao , Ziqiao Meng , Guangyan Zhang , Qichao Wang , Yiwen Guo , Irwin King

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing E2E approaches…

The visual dubbing task aims to generate mouth movements synchronized with the driving audio, which has seen significant progress in recent years. However, two critical deficiencies hinder their wide application: (1) Audio-only driving…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Liyang Chen , Tianze Zhou , Xu He , Boshi Tang , Zhiyong Wu , Yang Huang , Yang Wu , Zhongqian Sun , Wei Yang , Helen Meng

In visual speech processing, context modeling capability is one of the most important requirements due to the ambiguous nature of lip movements. For example, homophenes, words that share identical lip movements but produce different sounds,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Jeong Hun Yeo , Seunghee Han , Minsu Kim , Yong Man Ro

Video summarization aims to create short, accurate, and cohesive summaries of longer videos. Despite the existence of various video summarization datasets, a notable limitation is their limited amount of source videos, which hampers the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Hang Hua , Yolo Yunlong Tang , Chenliang Xu , Jiebo Luo

Currently, large language models (LLMs) predominantly focus on the text modality. To enable more natural human-AI interaction, speech LLMs are emerging, but building effective end-to-end speech LLMs remains challenging due to limited data…

Computation and Language · Computer Science 2026-04-14 Yan Zhou , Qingkai Fang , Yun Hong , Yang Feng

Transformer-based text to speech (TTS) model (e.g., Transformer TTS~\cite{li2019neural}, FastSpeech~\cite{ren2019fastspeech}) has shown the advantages of training and inference efficiency over RNN-based model (e.g.,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Mingjian Chen , Xu Tan , Yi Ren , Jin Xu , Hao Sun , Sheng Zhao , Tao Qin , Tie-Yan Liu

User Simulators play a pivotal role in training and evaluating task-oriented dialogue systems. Traditional user simulators typically rely on human-engineered agendas, resulting in generated responses that often lack diversity and…

Computation and Language · Computer Science 2024-05-24 Xiang Luo , Zhiwen Tang , Jin Wang , Xuejie Zhang

Learning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical alternative, a large-scale but uncurated and narrated video…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Dohwan Ko , Joonmyung Choi , Juyeon Ko , Shinyeong Noh , Kyoung-Woon On , Eun-Sol Kim , Hyunwoo J. Kim

Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint speech-text modeling is a promising direction to achieve this. However, the effectiveness of recent speech…

Computation and Language · Computer Science 2026-02-06 Liang-Hsuan Tseng , Yi-Chang Chen , Kuan-Yi Lee , Da-Shan Shiu , Hung-yi Lee

Speech LLM post-training increasingly relies on efficient cross-modal alignment and robust low-resource adaptation, yet collecting large-scale audio-text pairs remains costly. Text-only alignment methods such as TASU reduce this burden by…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-10 Jing Peng , Chenghao Wang , Yi Yang , Lirong Qian , Junjie Li , Yu Xi , Shuai Wang , Kai Yu

Recent advancements in text-to-speech (TTS) technology have increased demand for personalized audio synthesis. Zero-shot voice cloning, a specialized TTS task, aims to synthesize a target speaker's voice using only a single audio sample and…

Sound · Computer Science 2025-06-03 Ming Meng , Ziyi Yang , Jian Yang , Zhenjie Su , Yonggui Zhu , Zhaoxin Fan

The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and contextually relevant dialogues. However, extending these…

Computation and Language · Computer Science 2024-08-23 Yinghao Aaron Li , Xilin Jiang , Jordan Darefsky , Ge Zhu , Nima Mesgarani

In recent years, Text-To-Speech (TTS) has been used as a data augmentation technique for speech recognition to help complement inadequacies in the training data. Correspondingly, we investigate the use of a multi-speaker TTS system to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-25 Yiling Huang , Yutian Chen , Jason Pelecanos , Quan Wang

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both approaches leave models with little understanding of global…

While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public…

Sound · Computer Science 2026-05-19 Qifan Liang , Yuansen Liu , Ruixin Wei , Nan Lu , Junchuan Zhao , Ye Wang

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches…

Sound · Computer Science 2025-12-22 Zhedong Zhang , Liang Li , Gaoxiang Cong , Chunshan Liu , Yuhan Gao , Xiaowan Wang , Tao Gu , Yuankai Qi

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiaxin Ye , Hongming Shan

For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique speaking style or…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Longhao Zhang , Shuang Liang , Zhipeng Ge , Tianshu Hu

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to…

Sound · Computer Science 2026-04-15 Gaoxiang Cong , Liang Li , Jiaxin Ye , Zhedong Zhang , Hongming Shan , Yuankai Qi , Qingming Huang