English
Related papers

Related papers: SAO-Instruct: Free-form Audio Editing using Natura…

200 papers

Large "instruction-tuned" language models (i.e., finetuned to respond to instructions) have demonstrated a remarkable ability to generalize zero-shot to new tasks. Nevertheless, they depend heavily on human-written instruction data that is…

Computation and Language · Computer Science 2023-05-29 Yizhong Wang , Yeganeh Kordi , Swaroop Mishra , Alisa Liu , Noah A. Smith , Daniel Khashabi , Hannaneh Hajishirzi

Complex narrative contexts often challenge language models' ability to follow instructions, and existing benchmarks fail to capture these difficulties. To address this, we propose Concise-SAE, a training-free framework that improves…

Computation and Language · Computer Science 2025-05-23 Runcong Zhao , Chengyu Cao , Qinglin Zhu , Xiucheng Lv , Shun Shao , Lin Gui , Ruifeng Xu , Yulan He

Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-based methods achieved zero-shot audio editing by using a…

Sound · Computer Science 2023-04-06 Yuancheng Wang , Zeqian Ju , Xu Tan , Lei He , Zhizheng Wu , Jiang Bian , Sheng Zhao

Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a…

Sound · Computer Science 2024-06-10 Manjie Xu , Chenxing Li , Duzhen zhang , Dan Su , Wei Liang , Dong Yu

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are restricted to…

Sound · Computer Science 2025-09-29 Zitong Lan , Yiduo Hao , Mingmin Zhao

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend on explicit labels or reference audio, which limits both…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Yun Chen , Qi Chen , Zheqi Dai , Arshdeep Singh , Philip J. B. Jackson , Mark D. Plumbley

In this paper, we explore audio-editing with non-rigid text edits. We show that the proposed editing pipeline is able to create audio edits that remain faithful to the input audio. We explore text prompts that perform addition, style…

Free-form, text-based audio editing remains a persistent challenge, despite progress in inversion-based neural methods. Current approaches rely on slow inversion procedures, limiting their practicality. We present a virtual-consistency…

Sound · Computer Science 2025-09-23 Matthieu Cervera , Francesco Paissan , Mirco Ravanelli , Cem Subakan

The NLP community has broadly focused on text-only approaches of cognitive state tasks, but audio can provide vital missing cues through prosody. We posit that text-to-speech models learn to track aspects of cognitive state in order to…

Sound · Computer Science 2025-02-12 Adil Soubki , John Murzaku , Peter Zeng , Owen Rambow

We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Bosheng Qin , Juncheng Li , Siliang Tang , Tat-Seng Chua , Yueting Zhuang

Although natural language instructions offer an intuitive way to guide automated image editing, deep-learning models often struggle to achieve high-quality results, largely due to the difficulty of creating large, high-quality training…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Sherry X. Chen , Misha Sra , Pradeep Sen

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-25 Chunyu Qiang , Kang Yin , Xiaopeng Wang , Yuzhe Liang , Jiahui Zhao , Ruibo Fu , Tianrui Wang , Cheng Gong , Chen Zhang , Longbiao Wang , Jianwu Dang

Diffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audios. However, beyond generation, the task of audio editing…

Sound · Computer Science 2024-10-01 Yuhang Jia , Yang Chen , Jinghua Zhao , Shiwan Zhao , Wenjia Zeng , Yong Chen , Yong Qin

Instruction tuning enables pretrained language models to perform new tasks from inference-time natural language descriptions. These approaches rely on vast amounts of human supervision in the form of crowdsourced datasets or user…

Computation and Language · Computer Science 2022-12-20 Or Honovich , Thomas Scialom , Omer Levy , Timo Schick

Instruction tuning plays a pivotal role in Code Large Language Models (Code LLMs) for the task of program synthesis. Presently, two dominant paradigms for collecting tuning data are natural-instruct (human-written) and self-instruct…

Computation and Language · Computer Science 2024-03-04 Xianzhen Luo , Qingfu Zhu , Zhiming Zhang , Xu Wang , Qing Yang , Dongliang Xu , Wanxiang Che

Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text…

Sound · Computer Science 2024-03-06 Xuenan Xu , Zhiling Zhang , Zelin Zhou , Pingyue Zhang , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Speech editing systems aim to naturally modify speech content while preserving acoustic consistency and speaker identity. However, previous studies often struggle to adapt to unseen and diverse acoustic conditions, resulting in degraded…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-04 Taewoo Kim , Uijong Lee , Hayoung Park , Choongsang Cho , Nam In Park , Young Han Lee

Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlighting the need for additional non-text editing prompts. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Hyeonyu Kim , Seokhoon Jeong , Seonghee Han , Chanhyuk Choi , Taehwan Kim

Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally limited. Training-free methods often suffer from signal…

Sound · Computer Science 2026-01-21 Ye Tao , Wen Wu , Chao Zhang , Mengyue Wu , Shuai Wang , Xuenan Xu

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

Sound · Computer Science 2024-02-08 Dan Lyth , Simon King
‹ Prev 1 2 3 10 Next ›