中文
相关论文

相关论文: Accommodating Sample Size Effect on Similarity Mea…

200 篇论文

In this paper, we consider the effect of a bandwidth extension of narrow-band speech signals (0.3-3.4 kHz) to 0.3-8 kHz on speaker verification. Using covariance matrix based verification systems together with detection error trade-off…

声音 · 计算机科学 2022-04-06 Marcos Faundez-Zanuy , Mattias Nilsson , W. Bastiaan Kleijn

Most studies on speaker verification systems focus on long-duration utterances, which are composed of sufficient phonetic information. However, the performances of these systems are known to degrade when short-duration utterances are…

音频与语音处理 · 电气工程与系统科学 2020-08-05 Seung-bin Kim , Jee-weon Jung , Hye-jin Shim , Ju-ho Kim , Ha-Jin Yu

Despite growing interest in generating high-fidelity accents, evaluating accent similarity in speech synthesis has been underexplored. We aim to enhance both subjective and objective evaluation methods for accent similarity. Subjectively,…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Jinzuomu Zhong , Suyuan Liu , Dan Wells , Korin Richmond

We propose three regularization-based speaker adaptation approaches to adapt the attention-based encoder-decoder (AED) model with very limited adaptation data from target speakers for end-to-end automatic speech recognition. The first…

计算与语言 · 计算机科学 2019-11-12 Zhong Meng , Yashesh Gaur , Jinyu Li , Yifan Gong

The paper introduces a new kernel-based Maximum Mean Discrepancy (MMD) statistic for measuring the distance between two distributions given finitely-many multivariate samples. When the distributions are locally low-dimensional, the proposed…

机器学习 · 统计学 2018-09-03 Xiuyuan Cheng , Alexander Cloninger , Ronald R. Coifman

In multi-speaker applications is common to have pre-computed models from enrolled speakers. Using these models to identify the instances in which these speakers intervene in a recording is the task of speaker tracking. In this paper, we…

In many contemporary statistical and machine learning methods, one needs to optimize an objective function that depends on the discrepancy between two probability distributions. The discrepancy can be referred to as a metric for…

机器学习 · 计算机科学 2025-02-11 Yijin Ni , Xiaoming Huo

Objective: To enable reliable smartphone-based hearing assessments by developing methods to estimate device calibration offsets using categorical loudness scaling (CLS). Design: Calibration offsets were simulated from a Gaussian…

医学物理 · 物理学 2026-05-11 Chen Xu , Birger Kollmeier

Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have…

音频与语音处理 · 电气工程与系统科学 2022-11-02 Zili Huang , Desh Raj , Paola García , Sanjeev Khudanpur

Keeping in consideration the high demand for clustering, this paper focuses on understanding and implementing K-means clustering using two different similarity measures. We have tried to cluster the documents using two different measures…

信息检索 · 计算机科学 2015-05-04 Manan Mohan Goyal , Neha Agrawal , Manoj Kumar Sarma , Nayan Jyoti Kalita

Noise robustness in keyword spotting remains a challenge as many models fail to overcome the heavy influence of noises, causing the deterioration of the quality of feature embeddings. We proposed a contrastive regularization method called…

声音 · 计算机科学 2022-09-15 Dianwen Ng , Jia Qi Yip , Tanmay Surana , Zhao Yang , Chong Zhang , Yukun Ma , Chongjia Ni , Eng Siong Chng , Bin Ma

Contrastive predictive coding (CPC) aims to learn representations of speech by distinguishing future observations from a set of negative examples. Previous work has shown that linear classifiers trained on CPC features can accurately…

音频与语音处理 · 电气工程与系统科学 2021-08-03 Benjamin van Niekerk , Leanne Nortje , Matthew Baas , Herman Kamper

Clustering short text embeddings is a foundational task in natural language processing, yet remains challenging due to the need to specify the number of clusters in advance. We introduce a scalable spectral method that estimates the number…

机器学习 · 计算机科学 2025-11-26 Nikita Neveditsin , Pawan Lingras , Vijay Mago

Target speaker extraction (TSE) aims to isolate individual speaker voices from complex speech environments. The effectiveness of TSE systems is often compromised when the speaker characteristics are similar to each other. Recent research…

声音 · 计算机科学 2024-10-08 Yun Liu , Xuechen Liu , Junichi Yamagishi

For training the sequence-to-sequence voice conversion model, we need to handle an issue of insufficient data about the number of speech pairs which consist of the same utterance. This study experimentally investigated the effects of…

机器学习 · 计算机科学 2020-06-16 Yeongtae Hwang , Hyemin Cho , Hongsun Yang , Dong-Ok Won , Insoo Oh , Seong-Whan Lee

The potential use of non-linear speech features has not been investigated for music analysis although other commonly used speech features like Mel Frequency Ceptral Coefficients (MFCC) and pitch have been used extensively. In this paper, we…

声音 · 计算机科学 2014-06-11 Sunil Kumar Kopparapu , Meghna Pandharipande , G Sita

Typical deep clustering methods, while achieving notable progress, can only provide one clustering result per dataset. This limitation arises from their assumption of a fixed underlying data distribution, which may fail to meet user needs…

机器学习 · 计算机科学 2025-12-02 Xinyue Wang , Yuheng Jia , Hui Liu , Junhui Hou

Selecting application scenarios matching data is important for the automatic speech recognition (ASR) training, but it is difficult to measure the matching degree of the training corpus. This study proposes a unsupervised target-aware data…

计算与语言 · 计算机科学 2023-02-28 Changfeng Gao , Gaofeng Cheng , Pengyuan Zhang , Yonghong Yan

Word matches are often used in sequence comparison methods, either as a measure of sequence similarity or in the first search steps of algorithms such as BLAST or BLAT. The D2 statistic is the number of matches of words of k letters between…

定量方法 · 定量生物学 2009-09-09 Sylvain Foret , Susan R. Wilson , Conrad J. Burden

Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a…

音频与语音处理 · 电气工程与系统科学 2021-06-11 Yi-Chiao Wu , Cheng-Hung Hu , Hung-Shin Lee , Yu-Huai Peng , Wen-Chin Huang , Yu Tsao , Hsin-Min Wang , Tomoki Toda