中文
相关论文

相关论文: Design Choices for X-vector Based Speaker Anonymiz…

200 篇论文

Automatic Speaker Verification systems are gaining popularity these days; spoofing attacks are of prime concern as they make these systems vulnerable. Some spoofing attacks like Replay attacks are easier to implement but are very hard to…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Rahul T P , P R Aravind , Ranjith C , Usamath Nechiyil , Nandakumar Paramparambath

Voice trigger detection is an important task, which enables activating a voice assistant when a target user speaks a keyword phrase. A detector is typically trained on speech data independent of speaker information and used for the voice…

We study the individuality of the human voice with respect to a widely used feature representation of speech utterances, namely, the i-vector model. As a first step toward this goal, we compare and contrast uniqueness measures proposed for…

音频与语音处理 · 电气工程与系统科学 2021-03-04 Erkam Sinan Tandogan , Husrev Taha Sencar

This paper summarizes the applied deep learning practices in the field of speaker recognition, both verification and identification. Speaker recognition has been a widely used field topic of speech technology. Many research works have been…

音频与语音处理 · 电气工程与系统科学 2022-09-27 Dávid Sztahó , György Szaszák , András Beke

We propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification. Typically, speaker embedding extraction includes training a DNN for speaker classification and using…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Amber Afshan , Abeer Alwan

In this paper, we propose a new differentiable neural network alignment mechanism for text-dependent speaker verification which uses alignment models to produce a supervector representation of an utterance. Unlike previous works with…

声音 · 计算机科学 2018-12-27 Victoria Mingote , Antonio Miguel , Alfonso Ortega , Eduardo Lleida

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. This paper proposes a hierarchical network with transformer encoders and memory mechanism to address this problem. The proposed…

声音 · 计算机科学 2020-11-02 Yanpei Shi , Mingjie Chen , Qiang Huang , Thomas Hain

Contrary to i-vectors, speaker embeddings such as x-vectors are incapable of leveraging unlabelled utterances, due to the classification loss over training speakers. In this paper, we explore an alternative training strategy to enable the…

计算机视觉与模式识别 · 计算机科学 2019-04-24 Themos Stafylakis , Johan Rohdin , Oldrich Plchot , Petr Mizera , Lukas Burget

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to…

声音 · 计算机科学 2024-06-17 Tanel Pärnamaa , Ando Saabas

DER is the primary metric to evaluate diarization performance while facing a dilemma: the errors in short utterances or segments tend to be overwhelmed by longer ones. Short segments, e.g., `yes' or `no,' still have semantic information.…

声音 · 计算机科学 2022-11-09 Tao Liu , Kai Yu

We provide the technical report for Ego4D audio-only diarization challenge in ECCV 2022. Speaker diarization takes the audio streams as input and outputs the homogeneous segments according to the speaker's identity. It aims to solve the…

声音 · 计算机科学 2022-11-17 Jiahao Wang , Guo Chen , Yin-Dong Zheng , Tong Lu

Speech synthesis, voice cloning, and voice conversion techniques present severe privacy and security threats to users of voice user interfaces (VUIs). These techniques transform one or more elements of a speech signal, e.g., identity and…

密码学与安全 · 计算机科学 2021-07-23 Ranya Aloufi , Hamed Haddadi , David Boyle

For many decades, research in speech technologies has focused upon improving reliability. With this now meeting user expectations for a range of diverse applications, speech technology is today omni-present. As result, a focus on security…

Speech is easily leaked imperceptibly, such as being recorded by mobile phones in different situations. Private content in speech may be maliciously extracted through speech enhancement technology. Speech enhancement technology has…

声音 · 计算机科学 2022-06-17 Mingyu Dong , Diqun Yan , Rangding Wang

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily…

声音 · 计算机科学 2026-05-19 Jun Xue , Tong Zhang , Zhuolin Yi , Yihuan Huang , Yi Chai , Yiyang Zhang , Yanzhen Ren

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic…

声音 · 计算机科学 2021-04-26 Chau Luu , Peter Bell , Steve Renals

As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately.…

音频与语音处理 · 电气工程与系统科学 2026-01-29 Yuxiang Wang , Hongyu Liu , Dekun Chen , Xueyao Zhang , Zhizheng Wu

This paper presents SpecWav-Attack, an adversarial model for detecting speakers in anonymized speech. It leverages Wav2Vec2 for feature extraction and incorporates spectrogram resizing and incremental training for improved performance.…

声音 · 计算机科学 2025-05-16 Yuqi Li , Yuanzhong Zheng , Zhongtian Guo , Yaoxuan Wang , Jianjun Yin , Haojun Fei

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing…

音频与语音处理 · 电气工程与系统科学 2021-10-20 Sefik Emre Eskimez , Takuya Yoshioka , Huaming Wang , Xiaofei Wang , Zhuo Chen , Xuedong Huang

In this paper, adaptive mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification. First, adaptive convolutional neural networks (ACNNs) are employed in frame-level embedding…

音频与语音处理 · 电气工程与系统科学 2025-12-18 Bin Gu , Wu Guo , Lirong Dai , Jun Du
‹ 上一页 1 8 9 10 下一页 ›