English
Related papers

Related papers: Accommodating Sample Size Effect on Similarity Mea…

200 papers

The distribution of sentence length in ordinary language is not well captured by the existing models. Here we survey previous models of sentence length and present our random walk model that offers both a better fit with the data and a…

Computation and Language · Computer Science 2019-05-23 Gábor Borbély , András Kornai

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data…

We consider quantizing an $Ld$-dimensional sample, which is obtained by concatenating $L$ vectors from datasets of $d$-dimensional vectors, to a $d$-dimensional cluster center. The distortion measure is the weighted sum of $r$th powers of…

Machine Learning · Computer Science 2020-10-26 Erdem Koyuncu

We consider the problem of metric learning subject to a set of constraints on relative-distance comparisons between the data items. Such constraints are meant to reflect side-information that is not expressed directly in the feature vectors…

Machine Learning · Computer Science 2016-12-06 Ehsan Amid , Aristides Gionis , Antti Ukkonen

Large language models (LLMs) have spurred development in multiple industries. However, the growing number of their parameters brings substantial storage and computing burdens, making it essential to explore model compression techniques for…

Machine Learning · Computer Science 2025-01-16 Binrui Zeng , Yongtao Tang , Xiaodong Liu , Xiaopeng Li

While there has been substantial amount of work in speaker diarization recently, there are few efforts in jointly employing lexical and acoustic information for speaker segmentation. Towards that, we investigate a speaker diarization system…

Audio and Speech Processing · Electrical Eng. & Systems 2018-05-29 Tae Jin Park , Panayiotis Georgiou

The Kullback-Leibler (KL) divergence is a foundational measure for comparing probability distributions. Yet in multivariate settings, its single value often obscures the underlying reasons for divergence, conflating mismatches in individual…

Other Computer Science · Computer Science 2025-05-06 William Cook

In speech recognition problems, data scarcity often poses an issue due to the willingness of humans to provide large amounts of data for learning and classification. In this work, we take a set of 5 spoken Harvard sentences from 7 subjects…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-06 Jordan J. Bird , Diego R. Faria , Anikó Ekárt , Cristiano Premebida , Pedro P. S. Ayrosa

The popular i-vector model represents speakers as low-dimensional continuous vectors (i-vectors), and hence it is a way of continuous speaker embedding. In this paper, we investigate binary speaker embedding, which transforms i-vectors to…

Sound · Computer Science 2016-04-01 Lantian Li , Dong Wang , Chao Xing , Kaimin Yu , Thomas Fang Zheng

In the pursuit of supporting more languages around the world, tools that characterize properties of languages play a key role in expanding the existing multilingual NLP research. In this study, we focus on a widely used typological…

Computation and Language · Computer Science 2024-05-21 Hasti Toossi , Guo Qing Huai , Jinyu Liu , Eric Khiu , A. Seza Doğruöz , En-Shiun Annie Lee

This paper introduces kdiff, a novel kernel-based measure for estimating distances between instances of time series, random fields and other forms of structured data. This measure is based on the idea of matching distributions that only…

Machine Learning · Statistics 2021-10-01 Srinjoy Das , Hrushikesh Mhaskar , Alexander Cloninger

Speaker verification is the process by which a speakers claim of identity is tested against a claimed speaker by his or her voice. Speaker verification is done by the use of some parameters (features) from the speakers voice which can be…

Sound · Computer Science 2019-08-16 Bhavana V. S , Pradip K. Das

Pretrained speech representations like wav2vec2 and HuBERT exhibit strong anisotropy, leading to high similarity between random embeddings. While widely observed, the impact of this property on downstream tasks remains unclear. This work…

Sound · Computer Science 2025-06-16 Guillaume Wisniewski , Séverine Guillaume , Clara Rosina Fernández

We present an investigation into a hitherto unexplored systematic that affects the accuracy of galaxy cluster mass estimates with weak gravitational lensing. Specifically, we study the covariance between the weak lensing signal,…

Cosmology and Nongalactic Astrophysics · Physics 2024-06-10 Zhuowen Zhang , Arya Farahi , Daisuke Nagai , Erwin T. Lau , Joshua Frieman , Marina Ricci , Anja von der Linden , Hao-yi Wu

In recent years identity-vector (i-vector) based speaker verification (SV) systems have become very successful. Nevertheless, environmental noise and speech duration variability still have a significant effect on degrading the performance…

Sound · Computer Science 2016-08-09 Ali Khodabakhsh , Seyyed Saeed Sarfjoo , Umut Uludag , Osman Soyyigit , Cenk Demiroglu

Meta-analyses frequently include trials that report multiple effect sizes based on a common set of study participants. These effect sizes will generally be correlated. Cluster-robust variance-covariance estimators are a fruitful approach…

Methodology · Statistics 2022-03-07 Thilo Welz , Wolfgang Viechtbauer , Markus Pauly

We examine the estimation of the Kullback-Leibler (KL) divergence and the use of the goodness-of-fit test for multivariate continuous distributions. Our starting point is the maximum entropy principle for Shannon entropy: among all…

Statistics Theory · Mathematics 2026-03-10 Mehmet Siddik Cadirci , Martin Singull

Speaker diarization is the process of labeling different speakers in a speech signal. Deep speaker embeddings are generally extracted from short speech segments and clustered to determine the segments belong to same speaker identity. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Myungjong Kim , Vijendra Raj Apsingekar , Divya Neelagiri

A recent proposal of data dependent similarity called Isolation Kernel/Similarity has enabled SVM to produce better classification accuracy. We identify shortcomings of using a tree method to implement Isolation Similarity; and propose a…

Machine Learning · Computer Science 2024-01-30 Xiaoyu Qin , Kai Ming Ting , Ye Zhu , Vincent CS Lee

Compact sources can cause scatter in the scaling relationships between the amplitude of the thermal Sunyaev-Zel'dovich Effect (tSZE) in galaxy clusters and cluster mass. Estimates of the importance of this scatter vary - largely due to…

‹ Prev 1 8 9 10 Next ›