中文
相关论文

相关论文: An Audio-textual Diffusion Model For Converting Sp…

200 篇论文

Speech-to-speech translation (S2ST) aims to convert spoken input in one language to spoken output in another, typically focusing on either language translation or accent adaptation. However, effective cross-cultural communication requires…

计算与语言 · 计算机科学 2025-05-09 Abhishek Mishra , Ritesh Sur Chowdhury , Vartul Bahuguna , Isha Pandey , Ganesh Ramakrishnan

Diffusion models for text-to-image generation, known for their efficiency, accessibility, and quality, have gained popularity. While inference with these systems on consumer-grade GPUs is increasingly feasible, training from scratch…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Bram de Wilde , Anindo Saha , Maarten de Rooij , Henkjan Huisman , Geert Litjens

Text-to-audio (TTA), which generates audio signals from textual descriptions, has received huge attention in recent years. However, recent works focused on text to monaural audio only. As we know, spatial audio provides more immersive…

声音 · 计算机科学 2025-06-09 Lei Zhao , Sizhou Chen , Linfeng Feng , Jichao Zhang , Xiao-Lei Zhang , Chi Zhang , Xuelong Li

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

Properly setting up recording conditions, including microphone type and placement, room acoustics, and ambient noise, is essential to obtaining the desired acoustic characteristics of speech. In this paper, we propose Diff-R-EN-T, a…

声音 · 计算机科学 2024-01-17 Jaekwon Im , Juhan Nam

Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale…

声音 · 计算机科学 2026-04-21 Yuxuan Jiang , Zehua Chen , Zeqian Ju , Yusheng Dai , Weibei Dou , Jun Zhu

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper proposes the…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Rui-Chen Zheng , Yang Ai , Zhen-Hua Ling

Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound…

声音 · 计算机科学 2023-05-01 Dongchao Yang , Jianwei Yu , Helin Wang , Wen Wang , Chao Weng , Yuexian Zou , Dong Yu

We investigate the automatic processing of child speech therapy sessions using ultrasound visual biofeedback, with a specific focus on complementing acoustic features with ultrasound images of the tongue for the tasks of speaker diarization…

音频与语音处理 · 电气工程与系统科学 2019-08-16 Manuel Sam Ribeiro , Aciel Eshky , Korin Richmond , Steve Renals

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models…

Collecting sufficient amount of data that can represent various acoustic environmental attributes is a critical problem for distributed acoustic machine learning. Several audio data augmentation techniques have been introduced to address…

声音 · 计算机科学 2021-01-07 Chunheng Jiang , Jae-wook Ahn , Nirmit Desai

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem. Specifically,…

计算与语言 · 计算机科学 2023-10-27 Yongxin Zhu , Zhujin Gao , Xinyuan Zhou , Zhongyi Ye , Linli Xu

Text-to-image generation (TTI) refers to the usage of models that could process text input and generate high fidelity images based on text descriptions. Text-to-image generation using neural networks could be traced back to the emergence of…

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tingle Li , Baihe Huang , Xiaobin Zhuang , Dongya Jia , Jiawei Chen , Yuping Wang , Zhuo Chen , Gopala Anumanchipalli , Yuxuan Wang

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disordered speech…

音频与语音处理 · 电气工程与系统科学 2022-03-22 Shujie Hu , Shansong Liu , Xurong Xie , Mengzhe Geng , Tianzi Wang , Shoukang Hu , Mingyu Cui , Xunying Liu , Helen Meng

Text-To-Image (TTI) generation is significant for controlled and diverse image generation with broad potential applications. Although current medical TTI methods have made some progress in report-to-Chest-Xray (CXR) generation, their…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Peng Huang , Bowen Guo , Shuyu Liang , Junhu Fu , Yuanyuan Wang , Yi Guo

Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consistency due to limited natural language understanding and data…

The goal of this work is zero-shot text-to-speech synthesis, with speaking styles and voices learnt from facial characteristics. Inspired by the natural fact that people can imagine the voice of someone when they look at his or her face, we…

机器学习 · 计算机科学 2023-02-28 Jiyoung Lee , Joon Son Chung , Soo-Whan Chung

Human infants face a formidable challenge in speech acquisition: mapping extremely variable acoustic inputs into appropriate articulatory movements without explicit instruction. We present a computational model that addresses the…

音频与语音处理 · 电气工程与系统科学 2025-09-16 Marvin Lavechin , Thomas Hueber

Thousands of individuals need surgical removal of their larynx due to critical diseases every year and therefore, require an alternative form of communication to articulate speech sounds after the loss of their voice box. This work…

图像与视频处理 · 电气工程与系统科学 2020-07-01 Pramit Saha , Yadong Liu , Bryan Gick , Sidney Fels