中文
相关论文

相关论文: Speaker embeddings by modeling channel-wise correl…

200 篇论文

Speaker verification aims to verify whether an input speech corresponds to the claimed speaker, and conventionally, this kind of system is deployed based on single-stream scenario, wherein the feature extractor operates in full frequency…

声音 · 计算机科学 2025-09-03 Wei Yao , Shen Chen , Jiamin Cui , Yaolin Lou

We address monaural multi-speaker-image separation in reverberant conditions, aiming at separating mixed speakers but preserving the reverberation of each speaker. A straightforward approach for this task is to directly train end-to-end DNN…

音频与语音处理 · 电气工程与系统科学 2025-10-08 Jingqi Sun , Shulin He , Ruizhe Pang , Zhong-Qiu Wang

We tackle the multi-party speech recovery problem through modeling the acoustic of the reverberant chambers. Our approach exploits structured sparsity models to perform room modeling and speech recovery. We propose a scheme for…

机器学习 · 计算机科学 2012-10-26 Afsaneh Asaei , Mohammad Golbabaee , Hervé Bourlard , Volkan Cevher

In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without…

音频与语音处理 · 电气工程与系统科学 2025-12-01 Sara Barahona , Ladislav Mošner , Themos Stafylakis , Oldřich Plchot , Junyi Peng , Lukáš Burget , Jan Černocký

Previous work has encouraged domain-invariance in deep speaker embedding by adversarially classifying the dataset or labelled environment to which the generated features belong. We propose a training strategy which aims to produce features…

声音 · 计算机科学 2020-04-17 Chau Luu , Peter Bell , Steve Renals

In this paper, a novel Convolutional Neural Network architecture has been developed for speaker verification in order to simultaneously capture and discard speaker and non-speaker information, respectively. In training phase, the network is…

音频与语音处理 · 电气工程与系统科学 2018-08-13 Hossein Salehghaffari

Deep speaker embeddings have shown promising results in speaker recognition, as well as in other speaker-related tasks. However, some issues are still under explored, for instance, the information encoded in these representations and their…

音频与语音处理 · 电气工程与系统科学 2022-12-15 Zifeng Zhao , Ding Pan , Junyi Peng , Rongzhi Gu

From the patter of rain to the crunch of snow, the sounds we hear often convey the visual textures that appear within a scene. In this paper, we present a method for learning visual styles from unlabeled audio-visual data. Our model learns…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Tingle Li , Yichen Liu , Andrew Owens , Hang Zhao

In this paper, we propose an online speaker adaptation method for WaveNet-based neural vocoders in order to improve their performance on speaker-independent waveform generation. In this method, a speaker encoder is first constructed using a…

音频与语音处理 · 电气工程与系统科学 2020-08-17 Qiuchen Huang , Yang Ai , Zhenhua Ling

This paper proposes a speech rhythm-based method for speaker embeddings to model phoneme duration using a few utterances by the target speaker. Speech rhythm is one of the essential factors among speaker characteristics, along with acoustic…

声音 · 计算机科学 2024-02-13 Kenichi Fujita , Atsushi Ando , Yusuke Ijima

Modeling virtual agents with behavior style is one factor for personalizing human agent interaction. We propose an efficient yet effective machine learning approach to synthesize gestures driven by prosodic features and text in the style of…

声音 · 计算机科学 2022-08-04 Mireille Fares , Michele Grimaldi , Catherine Pelachaud , Nicolas Obin

Many image enhancement or editing operations, such as forward and inverse tone mapping or color grading, do not have a unique solution, but instead a range of solutions, each representing a different style. Despite this, existing…

计算机视觉与模式识别 · 计算机科学 2022-10-05 Aamir Mustafa , Param Hanji , Rafal K. Mantiuk

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesised speech of a target speaker's timbre. Most previous approaches rely on data with style labels, but manually-annotated labels are…

声音 · 计算机科学 2022-12-14 Chunyu Qiang , Peng Yang , Hao Che , Xiaorui Wang , Zhongyuan Wang

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and…

声音 · 计算机科学 2025-05-29 Zhenghai You , Zhenyu Zhou , Lantian Li , Dong Wang

A primary challenge when deploying speaker recognition systems in real-world applications is performance degradation caused by environmental mismatch. We propose a diffusion-based method that takes speaker embeddings extracted from a…

音频与语音处理 · 电气工程与系统科学 2025-05-23 KiHyun Nam , Jungwoo Heo , Jee-weon Jung , Gangin Park , Chaeyoung Jung , Ha-Jin Yu , Joon Son Chung

Audio-driven talking head animation is a challenging research topic with many real-world applications. Recent works have focused on creating photo-realistic 2D animation, while learning different talking or singing styles remains an open…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Trong-Thang Pham , Nhat Le , Tuong Do , Hung Nguyen , Erman Tjiputra , Quang D. Tran , Anh Nguyen

While many recent any-to-any voice conversion models succeed in transferring some target speech's style information to the converted speech, they still lack the ability to faithfully reproduce the speaking style of the target speaker. In…

音频与语音处理 · 电气工程与系统科学 2023-12-18 Hyungseob Lim , Kyungguen Byun , Sunkuk Moon , Erik Visser

Recently, there have been efforts to encode the linguistic information of speech using a self-supervised framework for speech synthesis. However, predicting representations from surrounding representations can inadvertently entangle speaker…

声音 · 计算机科学 2024-04-02 Injune Hwang , Kyogu Lee

Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, classifying or indexing. Audio Word2Vec previously proposed was…

计算与语言 · 计算机科学 2018-11-08 Sung-Feng Huang , Yi-Chen Chen , Hung-yi Lee , Lin-shan Lee

Human interlocutors tend to engage in adaptive behavior known as entrainment to become more similar to each other. Isolating the effect of consistency, i.e., speakers adhering to their individual styles, is a critical part of the analysis…

计算与语言 · 计算机科学 2020-11-04 Andreas Weise , Rivka Levitan