English
Related papers

Related papers: Step-Audio-EditX Technical Report

200 papers

Existing autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a…

Computation and Language · Computer Science 2025-09-04 Siyi Zhou , Yiquan Zhou , Yi He , Xun Zhou , Jinchao Wang , Wei Deng , Jingchen Shu

Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making…

Sound · Computer Science 2026-03-11 Hezhao Zhang , Huang-Cheng Chou , Shrikanth Narayanan , Thomas Hain

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon persists in audio language models: they consistently…

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel multi-step hierarchical ED prediction module that quantifies…

Sound · Computer Science 2025-07-08 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Large language models (LLMs) have demonstrated impressive performance in mathematical and commonsense reasoning tasks using chain-of-thought (CoT) prompting techniques. But can they perform emotional reasoning by concatenating `Let's think…

Computation and Language · Computer Science 2024-08-12 Ankita Bhaumik , Tomek Strzalkowski

Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis. However, existing foundation models rely on multi-stage processing or complex architectures for predicting…

This paper introduces StyleSpeech, a novel Text-to-Speech~(TTS) system that enhances the naturalness and accuracy of synthesized speech. Building upon existing TTS technologies, StyleSpeech incorporates a unique Style Decorator structure…

Sound · Computer Science 2024-12-31 Haowei Lou , Helen Paik , Wen Hu , Lina Yao

We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Shitian Zhao , Qilong Wu , Xinyue Li , Bo Zhang , Ming Li , Qi Qin , Dongyang Liu , Kaipeng Zhang , Hongsheng Li , Yu Qiao , Peng Gao , Bin Fu , Zhen Li

Large Language Models (LLMs) have significantly advanced natural language processing, demonstrating strong capabilities in tasks such as text generation, summarization, and reasoning. Recently, their potential for automating precise text…

Computation and Language · Computer Science 2026-01-27 Yiming Zeng , Wanhao Yu , Zexin Li , Tao Ren , Yu Ma , Jinghan Cao , Xiyan Chen , Tingting Yu

Current instruction-based editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Yuzhou Huang , Liangbin Xie , Xintao Wang , Ziyang Yuan , Xiaodong Cun , Yixiao Ge , Jiantao Zhou , Chao Dong , Rui Huang , Ruimao Zhang , Ying Shan

The problem of audio-to-audio (A2A) style transfer involves replacing the style features of the source audio with those from the target audio while preserving the content related attributes of the source audio. In this paper, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

As text-based speech editing becomes increasingly prevalent, the demand for unrestricted free-text editing continues to grow. However, existing speech editing techniques encounter significant challenges, particularly in maintaining…

Sound · Computer Science 2024-09-23 Yang Chen , Yuhang Jia , Shiwan Zhao , Ziyue Jiang , Haoran Li , Jiarong Kang , Yong Qin

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At its core is a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-06 Jiyoung Lee , Song Park , Sanghyuk Chun , Soo-Whan Chung

Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting…

Multimedia · Computer Science 2026-01-16 Diqiong Jiang , Kai Zhu , Dan Song , Jian Chang , Chenglizhao Chen , Zhenyu Wu

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haojie Zheng , Yixin Yang , Siqi Yang , Shuchen Weng , Boxin Shi

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (context-aware mask prediction network) is proposed to realize…

Sound · Computer Science 2022-12-21 Tao Wang , Jiangyan Yi , Ruibo Fu , Jianhua Tao , Zhengqi Wen , Chu Yuan Zhang

Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a truly unified framework…

Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite recent progress,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Lixiang Lin , Siyuan Jin , Jinshan Zhang

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to…