中文
相关论文

相关论文: Scoring Formulation for Multi-Condition Joint PLDA

200 篇论文

The literature on clustering for continuous data is rich and wide; differently, that one developed for categorical data is still limited. In some cases, the problem is made more difficult by the presence of noise variables/dimensions that…

统计方法学 · 统计学 2015-04-14 Monia Ranalli , Roberto Rocci

Separating signal from noise is central to experiments. Applying well-established statistical methods effectively to LLM evals requires consideration of their unique noise characteristics. We clearly define and measure three types of noise:…

机器学习 · 计算机科学 2026-03-31 Sida Wang

Parameter estimation connects mathematical models to real-world data and decision making across many scientific and industrial applications. Standard approaches such as maximum likelihood estimation and Markov chain Monte Carlo estimate…

统计方法学 · 统计学 2026-02-06 Matthew J Simpson , James S Bennett , Alexander Johnston , Ruth E Baker

Probabilistic Component Latent Analysis (PLCA) is a statistical modeling method for feature extraction from non-negative data. It has been fruitfully applied to various research fields of information retrieval. However, the EM-solved…

统计方法学 · 统计学 2017-03-16 D. Cazau , G. Nuel

The current research has two aims. First, to demonstrate the utility conditional nonparametric multilevel latent class analysis (NP-MLCA) for multi-site program evaluation using an empirical dataset. Second, to investigate how…

A concrete, stylized example illustrates that inferences may be degraded, rather than improved, by incorporating supplementary data via a joint likelihood. In the example, the likelihood is assumed to be correctly specified, as is the prior…

其他统计学 · 统计学 2019-05-24 P. Richard Hahn

In many data-driven applications, collecting data from different sources is increasingly desirable for enhancing performance. In this paper, we are interested in the problem of probabilistic forecasting with multi-source time series. We…

机器学习 · 计算机科学 2023-02-23 Tian Guo

In many studies of human diseases, multiple omic datasets are measured. Typically, these omic datasets are studied one by one with the disease, thus the relationship between omics are overlooked. Modeling the joint part of multiple omics…

统计方法学 · 统计学 2022-09-02 Zhujie Gu , Said el Bouhaddani , Jeanine Houwing-Duistermaat , Hae-Won Uh

This paper develops a test for homogeneity in finite mixture models where the mixing proportions are known a priori (taken to be 0.5) and a common nuisance parameter is present. Statistical tests based on the notion of Projected Likelihood…

统计理论 · 数学 2008-12-18 Debapriya Sengupta , Rahul Mazumder

In the context of likelihood ratio testing with parameters on the boundary, we revisit two situations for which there are some discrepancies in the literature: the case of two parameters of interest on the boundary, with all other…

统计理论 · 数学 2025-09-03 Clara Bertinelli Salucci , Anders Kvellestad , Riccardo De Bin

This paper compares two different ways of estimating statistical language models. Many statistical NLP tagging and parsing models are estimated by maximizing the (joint) likelihood of the fully-observed training data. However, since these…

计算与语言 · 计算机科学 2007-05-23 Mark Johnson

In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring backends are commonly used, namely cosine scoring or PLDA. Both have advantages and disadvantages, depending on the context. Cosine…

Evaluations of large language models (LLMs) suffer from instability, where small changes of random factors such as few-shot examples can lead to drastic fluctuations of scores and even model rankings. Moreover, different LLMs can have…

机器学习 · 计算机科学 2025-09-17 Yiyang Li , Yonghuang Wu , Ying Luo , Liangtai Sun , Zishu Qin , Lin Qiu , Xuezhi Cao , Xunliang Cai

State-of-art speaker verification (SV) systems use a back-end model to score the similarity of speaker embeddings extracted from a neural network model. The commonly used back-end models are the cosine scoring and the probabilistic linear…

声音 · 计算机科学 2022-04-25 Zhiyuan Peng , Xuanji He , Ke Ding , Tan Lee , Guanglu Wan

In a statistical analysis in Particle Physics, nuisance parameters can be introduced to take into account various types of systematic uncertainties. The best estimate of such a parameter is often modeled as a Gaussian distributed variable…

数据分析、统计与概率 · 物理学 2019-02-25 Glen Cowan

In this paper, we propose a new variant of Linear Discriminant Analysis (LDA) to solve multi-label classification tasks. The proposed method is based on a probabilistic model for defining the weights of individual samples in a weighted…

机器学习 · 计算机科学 2020-04-10 Lei Xu , Jenni Raitoharju , Alexandros Iosifidis , Moncef Gabbouj

In mixture experiments with noise variables or process variables that can not be controlled, investigate and try to control the variability of the response variable is very important for quality improvement in industrial processes. Thus,…

We describe a method for parameter estimation in bipartite probabilistic graphical models for joint prediction of clinical conditions from the electronic medical record. The method does not rely on the availability of gold-standard labels,…

机器学习 · 统计学 2016-09-23 Yoni Halpern , Steven Horng , David Sontag

Multi-criteria decision analysis (MCDA) is a quantitative approach to the drug benefit-risk assessment (BRA) which allows for consistent comparisons by summarising all benefits and risks in a single score. The MCDA consists of several…

统计方法学 · 统计学 2021-07-27 Tom Menzies , Gaelle Saint-Hilary , Pavel Mozgunov

Estimation problems in the presence of deterministic linear nuisance parameters arise in a variety of fields. To cope with those, three common methods are widely considered: (1) jointly estimating the parameters of interest and the nuisance…

最优化与控制 · 数学 2018-07-05 Yongchang Hu , Geert Leus