中文
相关论文

相关论文: Text2FX: Harnessing CLAP Embeddings for Text-Guide…

200 篇论文

Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-Language Multirate and Multimodal Representations (CALM), an…

音频与语音处理 · 电气工程与系统科学 2022-02-09 Vin Sachidananda , Shao-Yen Tseng , Erik Marchi , Sachin Kajarekar , Panayiotis Georgiou

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

In this paper, we propose a visual embedding approach to improving embedding aware speech enhancement (EASE) by synchronizing visual lip frames at the phone and place of articulation levels. We first extract visual embedding from lip frames…

声音 · 计算机科学 2020-09-22 Hang Chen , Jun Du , Yu Hu , Li-Rong Dai , Bao-Cai Yin , Chin-Hui Lee

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

声音 · 计算机科学 2021-02-11 Zeqian Li , Jacob Whitehill

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding…

声音 · 计算机科学 2021-10-12 Cheng Gong , Longbiao Wang , Zhenhua Ling , Ju Zhang , Jianwu Dang

Diffusion models have shown superior performance in image generation and manipulation, but the inherent stochasticity presents challenges in preserving and manipulating image content and identity. While previous approaches like DreamBooth…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Inhwa Han , Serin Yang , Taesung Kwon , Jong Chul Ye

This study presents a control framework leveraging vision language models (VLMs) for multiple tasks and robots. Notably, existing control methods using VLMs have achieved high performance in various tasks and robots in the training…

机器人学 · 计算机科学 2024-01-19 Kazuki Shibata , Hideki Deguchi , Shun Taguchi

This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker…

音频与语音处理 · 电气工程与系统科学 2024-06-25 Mingyang Zhang , Yi Zhou , Yi Ren , Chen Zhang , Xiang Yin , Haizhou Li

Automated audio captioning is multi-modal translation task that aim to generate textual descriptions for a given audio clip. In this paper we propose a full Transformer architecture that utilizes Patchout as proposed in [1], significantly…

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest,…

声音 · 计算机科学 2026-04-17 Liting Gao , Yi Yuan , Yaru Chen , Yuelan Cheng , Zhenbo Li , Juan Wen , Shubin Zhang , Wenwu Wang

Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint…

声音 · 计算机科学 2025-10-17 Qixin Deng , Bryan Pardo , Thrasyvoulos N Pappas

Audio captioning is an important research area that aims to generate meaningful descriptions for audio clips. Most of the existing research extracts acoustic features of audio clips as input to encoder-decoder and transformer architectures…

声音 · 计算机科学 2022-04-20 Ayşegül Özkaya Eren , Mustafa Sert

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Text embeddings are fundamental to many natural language processing (NLP) tasks, extensively applied in domains such as recommendation systems and information retrieval (IR). Traditionally, transmitting embeddings instead of raw text has…

计算与语言 · 计算机科学 2025-07-11 Dominykas Seputis , Yongkang Li , Karsten Langerak , Serghei Mihailov

Automatic accent identification (AID) remains a challenging task due to the complex variability of accents, the entanglement of accent cues with speaker traits, and the scarcity of reliable accentlabelled data. To address these challenges,…

信号处理 · 电气工程与系统科学 2026-04-29 Rayane Bakari , Olivier Le Blouch , Nicolas Gengembre , Nicholas Evans

In this paper, we introduce the task of learning unsupervised dialogue embeddings. Trivial approaches such as combining pre-trained word or sentence embeddings and encoding through pre-trained language models (PLMs) have been shown to be…

计算与语言 · 计算机科学 2022-10-28 Che Liu , Rui Wang , Junfeng Jiang , Yongbin Li , Fei Huang

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose…

声音 · 计算机科学 2025-02-14 Kyungsu Kim , Junghyun Koo , Sungho Lee , Haesun Joung , Kyogu Lee

We propose Context-Adaptive Multi-Prompt Embedding, a novel approach to enrich semantic representations in vision-language contrastive learning. Unlike standard CLIP-style models that rely on a single text embedding, our method introduces…

机器学习 · 计算机科学 2025-08-07 Dahun Kim , Anelia Angelova

In this work, we are dedicated to text-guided image generation and propose a novel framework, i.e., CLIP2GAN, by leveraging CLIP model and StyleGAN. The key idea of our CLIP2GAN is to bridge the output feature embedding space of CLIP and…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Yixuan Wang , Wengang Zhou , Jianmin Bao , Weilun Wang , Li Li , Houqiang Li

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Wayne Wu , Chen Change Loy , Xun Cao , Feng Xu