中文
相关论文

相关论文: Shared latent subspace modelling within Gaussian-B…

200 篇论文

An extreme learning machine (ELM) is a three-layered feed-forward neural network having untrained parameters, which are randomly determined before training. Inspired by the idea of ELM, a probabilistic untrained layer called a…

机器学习 · 计算机科学 2022-10-28 Yuri Kanno , Muneki Yasuda

In audio applications, one of the most important representations of audio signals is the amplitude spectrogram. It is utilized in many machine-learning-based information processing methods including the ones using the restricted Boltzmann…

音频与语音处理 · 电气工程与系统科学 2020-06-26 Toru Nakashika , Kohei Yatabe

Phonotactic constraints can be employed to distinguish languages by representing a speech utterance as a multinomial distribution or phone events. In the present study, we propose a new learning mechanism based on subspace-based…

声音 · 计算机科学 2022-03-30 Hung-Shin Lee , Yu Tsao , Shyh-Kang Jeng , Hsin-Min Wang

This paper presents a linear regression based back-end for speaker verification. Linear regression is a simple linear model that minimizes the mean squared estimation error between the target and its estimate with a closed form solution,…

声音 · 计算机科学 2018-02-13 Xiao-Lei Zhang

In this paper we describe the recent advancements made in the IBM i-vector speaker recognition system for conversational speech. In particular, we identify key techniques that contribute to significant improvements in performance of our…

声音 · 计算机科学 2016-02-24 Seyed Omid Sadjadi , Sriram Ganapathy , Jason W. Pelecanos

Recently by the development of the Internet and the Web, different types of social media such as web blogs become an immense source of text data. Through the processing of these data, it is possible to discover practical information about…

计算与语言 · 计算机科学 2019-03-12 Masoud Fatemi , Mehran Safayani

The subspace Restricted Boltzmann Machine (subspaceRBM) is a third-order Boltzmann machine where multiplicative interactions are between one visible and two hidden units. There are two kinds of hidden units, namely, gate units and subspace…

机器学习 · 计算机科学 2014-07-17 Jakub M. Tomczak , Adam Gonczarek

This article presents a novel approach for learning domain-invariant speaker embeddings using Generative Adversarial Networks. The main idea is to confuse a domain discriminator so that is can't tell if embeddings are from the source or…

音频与语音处理 · 电气工程与系统科学 2018-11-08 Gautam Bhattacharya , Joao Monteiro , Jahangir Alam , Patrick Kenny

This paper addresses the problem of multiple-speaker localization in noisy and reverberant environments, using binaural recordings of an acoustic scene. A Gaussian mixture model (GMM) is adopted, whose components correspond to all the…

声音 · 计算机科学 2017-10-06 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

In this paper, we propose pass-phrase dependent background models (PBMs) for text-dependent (TD) speaker verification (SV) to integrate the pass-phrase identification process into the conventional TD-SV system, where a PBM is derived from a…

计算与语言 · 计算机科学 2021-03-29 A. K. Sarkar , Zheng-Hua Tan

In this paper, we propose a speaker-verification system based on maximum likelihood linear regression (MLLR) super-vectors, for which speakers are characterized by m-vectors. These vectors are obtained by a uniform segmentation of the…

声音 · 计算机科学 2016-05-13 A. K. Sarkar , C. Barras , V. B. Le , D. Matrouf

The state-of-art approach to speaker verification involves the extraction of discriminative embeddings like x-vectors followed by a generative model back-end using a probabilistic linear discriminant analysis (PLDA). In this paper, we…

音频与语音处理 · 电气工程与系统科学 2020-02-10 Shreyas Ramoji , Prashant Krishnan , Prachi Singh , Sriram Ganapathy

This work proposes a scalable probabilistic latent variable model based on Gaussian processes (Lawrence, 2004) in the context of multiple observation spaces. We focus on an application in astrophysics where data sets typically contain both…

星系天体物理 · 物理学 2025-02-28 Vidhi Lalchand , Anna-Christina Eilers

Understanding the results of deep neural networks is an essential step towards wider acceptance of deep learning algorithms. Many approaches address the issue of interpreting artificial neural networks, but often provide divergent…

机器学习 · 计算机科学 2021-11-16 Vadim Borisov , Johannes Meier , Johan van den Heuvel , Hamed Jalali , Gjergji Kasneci

We present an approach to tackle the speaker recognition problem using Triplet Neural Networks. Currently, the $i$-vector representation with probabilistic linear discriminant analysis (PLDA) is the most commonly used technique to solve…

声音 · 计算机科学 2019-10-07 Kin Wai Cheuk , Balamurali B. T. , Gemma Roig , Dorien Herremans

A novel Gaussian mixture model (GMM) aided sparse Bayesian learning (SBL) framework is proposed for channel state information (CSI) estimation in orthogonal time-frequency space (OTFS) modulated systems. The key attribute of the proposed…

信号处理 · 电气工程与系统科学 2026-03-31 Surbhi Gehlot , Suraj Srivastava , Sandeep Kumar Yadav , Lajos Hanzo

The state-of-art approach for speaker verification consists of a neural network based embedding extractor along with a backend generative model such as the Probabilistic Linear Discriminant Analysis (PLDA). In this work, we propose a neural…

音频与语音处理 · 电气工程与系统科学 2020-05-26 Shreyas Ramoji , Prashant Krishnan , Sriram Ganapathy

The trimming scheme with a prefixed cutoff portion is known as a method of improving the robustness of statistical models such as multivariate Gaussian mixture models (MG- MMs) in small scale tests by alleviating the impacts of outliers.…

计算与语言 · 计算机科学 2014-05-20 Dalei Wu , Haiqing Wu

We revisit the challenging problem of training Gaussian-Bernoulli restricted Boltzmann machines (GRBMs), introducing two innovations. We propose a novel Gibbs-Langevin sampling algorithm that outperforms existing methods like Gibbs…

机器学习 · 计算机科学 2022-10-20 Renjie Liao , Simon Kornblith , Mengye Ren , David J. Fleet , Geoffrey Hinton

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data…