中文
相关论文

相关论文: Leveraging Audio-Only Data for Text-Queried Target…

200 篇论文

Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously,…

声音 · 计算机科学 2025-11-13 Zixuan Li , Xueliang Zhang , Lei Miao , Zhipeng Yan , Ying Sun , Chong Zhu

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their…

A speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track.…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Automatic speech recognition (ASR) has been widely researched with supervised approaches, while many low-resourced languages lack audio-text aligned data, and supervised methods cannot be applied on them. In this work, we propose a…

计算与语言 · 计算机科学 2018-08-14 Yi-Chen Chen , Chia-Hao Shen , Sung-Feng Huang , Hung-yi Lee

This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic teleportation transfers room characteristics between speech…

音频与语音处理 · 电气工程与系统科学 2025-10-16 Philipp Grundhuber , Mhd Modar Halimeh , Emanuël A. P. Habets

Audio tagging has attracted increasing attention since last decade and has various potential applications in many fields. The objective of audio tagging is to predict the labels of an audio clip. Recently deep learning methods have been…

声音 · 计算机科学 2018-08-14 Shengyun Wei , Kele Xu , Dezhi Wang , Feifan Liao , Huaimin Wang , Qiuqiang Kong

Speaker extraction aims to extract target speech signal from a multi-talker environment with interference speakers and surrounding noise, given the target speaker's reference information. Most speaker extraction systems achieve satisfactory…

音频与语音处理 · 电气工程与系统科学 2022-08-12 Chengyun Deng , Shiqian Ma , Yi Zhang , Yongtao Sha , Hui Zhang , Hui Song , Xiangang Li

Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text…

声音 · 计算机科学 2024-03-06 Xuenan Xu , Zhiling Zhang , Zelin Zhou , Pingyue Zhang , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document's stance towards that target. Many works classify…

计算与语言 · 计算机科学 2025-10-28 Ethan Mines , Bonnie Dorr

It's assumed that training data is sufficient in base session of few-shot class-incremental audio classification. However, it's difficult to collect abundant samples for model training in base session in some practical scenarios due to the…

音频与语音处理 · 电气工程与系统科学 2024-06-13 Yongjie Si , Yanxiong Li , Jialong Li , Jiaxin Tan , Qianhua He

Mapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) performance in new domains. However, the length of speech…

声音 · 计算机科学 2023-10-10 Jiaxu Zhu , Weinan Tong , Yaoxun Xu , Changhe Song , Zhiyong Wu , Zhao You , Dan Su , Dong Yu , Helen Meng

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the performance of AV-TSE systems heavily relies on the quality of…

声音 · 计算机科学 2025-07-22 Junjie Li , Wenxuan Wu , Shuai Wang , Zexu Pan , Kong Aik Lee , Helen Meng , Haizhou Li

Recent advances in speech-enabled language models have shown promising results in building intelligent voice assistants. However, most existing approaches rely on large-scale paired speech-text data and extensive computational resources,…

计算与语言 · 计算机科学 2025-06-10 Taesoo Kim , Jong Hwan Ko

Recently, sequence-to-sequence (seq-to-seq) models have been successfully applied in text-to-speech (TTS) to synthesize speech for single-language text. To synthesize speech for multiple languages usually requires multi-lingual speech from…

声音 · 计算机科学 2022-11-18 Haitong Zhang , Yue Lin

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we…

Joint audio-text models are widely used for music retrieval, yet they struggle with semantic phenomena such as negation. Negation is fundamental for distinguishing the absence (or presence) of musical elements (e.g., "with vocals" vs.…

声音 · 计算机科学 2026-01-21 Yannis Vasilakis , Rachel Bittner , Johan Pauwels

Query-by-example (QbE) speech search is the task of matching spoken queries to utterances within a search collection. In low- or zero-resource settings, QbE search is often addressed with approaches based on dynamic time warping (DTW).…

计算与语言 · 计算机科学 2020-11-25 Yushi Hu , Shane Settle , Karen Livescu

We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs a front-end…

音频与语音处理 · 电气工程与系统科学 2023-03-16 Mohamed Elminshawi , Srikanth Raj Chetupalli , Emanuël A. P. Habets

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between…

音频与语音处理 · 电气工程与系统科学 2021-02-11 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Encoder pre-training is promising in end-to-end Speech Translation (ST), given the fact that speech-to-translation data is scarce. But ST encoders are not simple instances of Automatic Speech Recognition (ASR) or Machine Translation (MT)…

计算与语言 · 计算机科学 2021-06-16 Chen Xu , Bojie Hu , Yanyang Li , Yuhao Zhang , shen huang , Qi Ju , Tong Xiao , Jingbo Zhu