中文
相关论文

相关论文: How phonemes contribute to deep speaker models?

200 篇论文

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining…

声音 · 计算机科学 2024-10-11 Alican Akman , Qiyang Sun , Björn W. Schuller

Many hearables contain an in-ear microphone, which may be used to capture the own voice of its user in noisy environments. Since the in-ear microphone mostly records body-conducted speech due to ear canal occlusion, it suffers from…

音频与语音处理 · 电气工程与系统科学 2024-03-25 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo

Speech contains information that is clinically relevant to some diseases, which has the potential to be used for health assessment. Recent work shows an interest in applying deep learning algorithms, especially pretrained large speech…

声音 · 计算机科学 2024-07-02 Hok-Shing Lau , Mark Huntly , Nathon Morgan , Adesua Iyenoma , Biao Zeng , Tim Bashford

Speaker verification systems are vulnerable to spoofing attacks which presents a major problem in their real-life deployment. To date, most of the proposed synthetic speech detectors (SSDs) have weighted the importance of different segments…

声音 · 计算机科学 2016-10-11 Ali Khodabakhsh , Cenk Demiroglu

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

This study presents a novel zero-shot user-defined keyword spotting model that utilizes the audio-phoneme relationship of the keyword to improve performance. Unlike the previous approach that estimates at utterance level, we use both…

音频与语音处理 · 电气工程与系统科学 2023-09-01 Yong-Hyeok Lee , Namhyun Cho

With recent research advancements, deep learning models are becoming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality…

音频与语音处理 · 电气工程与系统科学 2021-05-20 Sebastian Braun , Hannes Gamper , Chandan K. A. Reddy , Ivan Tashev

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily…

声音 · 计算机科学 2026-05-19 Jun Xue , Tong Zhang , Zhuolin Yi , Yihuan Huang , Yi Chai , Yiyang Zhang , Yanzhen Ren

Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these models often exhibit…

计算与语言 · 计算机科学 2025-05-22 Yi-Cheng Lin , Wei-Chih Chen , Hung-yi Lee

Language models provide a key framework for studying linguistic theories based on prediction, but phonological analysis using large language models (LLMs) is difficult; there are few phonological benchmarks beyond English and the standard…

计算与语言 · 计算机科学 2025-06-13 Zébulon Goriely , Paula Buttery

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

声音 · 计算机科学 2022-06-22 Yuan Gong , Jin Yu , James Glass

Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based LLMs fundamentally limits…

计算与语言 · 计算机科学 2025-12-17 Marianne de Heer Kloots , Paul Boersma , Willem Zuidema

The fundamental frequency (F0) contour of speech is a key aspect to represent speech prosody that finds use in speech and spoken language analysis such as voice conversion and speech synthesis as well as speaker and language identification.…

音频与语音处理 · 电气工程与系统科学 2018-05-09 Akihiro Kato , Tomi Kinnunen

Understanding how speech foundation models capture non-verbal cues is crucial for improving their interpretability and adaptability across diverse tasks. In our work, we analyze several prominent models such as Whisper, Seamless, Wav2Vec,…

计算与语言 · 计算机科学 2024-10-18 Abdul Waheed , Hanin Atwany , Bhiksha Raj , Rita Singh

Cochlear implant (CI) users have considerable difficulty in understanding speech in reverberant listening environments. Time-frequency (T-F) masking is a common technique that aims to improve speech intelligibility by multiplying…

音频与语音处理 · 电气工程与系统科学 2021-06-01 Kevin M. Chu , Leslie M. Collins , Boyla O. Mainsah

This study analyzes changes in the attention mechanisms of large language models (LLMs) when used to understand natural conversations between humans (human-human). We analyze three use cases of LLMs: interactions over web content, code, and…

计算与语言 · 计算机科学 2024-03-11 Toshish Jawale , Chaitanya Animesh , Sekhar Vallath , Kartik Talamadupula , Larry Heck

Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and those are averaged to…

声音 · 计算机科学 2019-07-03 Miquel India , Pooyan Safari , Javier Hernando

Recently deep neural networks (DNNs) have been used to learn speaker features. However, the quality of the learned features is not sufficiently good, so a complex back-end model, either neural or probabilistic, has to be used to address the…

声音 · 计算机科学 2017-05-11 Lantian Li , Yixiang Chen , Ying Shi , Zhiyuan Tang , Dong Wang

We propose an approach for learning critical articulators for phonemes through a machine learning approach. We formulate the learning with three models trained end to end. First, we use Acoustic to Articulatory Inversion (AAI) to predict…

音频与语音处理 · 电气工程与系统科学 2025-05-02 Jesuraj Bandekar , Sathvik Udupa , Prasanta Kumar Ghosh

Recent advances in speech foundation models (SFMs) have enabled the direct processing of spoken language from raw audio, bypassing intermediate textual representations. This capability allows SFMs to be exposed to, and potentially respond…

音频与语音处理 · 电气工程与系统科学 2025-10-30 Harm Lameris , Shree Harsha Bokkahalli Satish , Joakim Gustafson , Éva Székely