English
Related papers

Related papers: Analysis of distributional variation through multi…

200 papers

A noniterative sample size procedure is proposed for a general hypothesis test based on the t distribution by modifying and extending Guenther's (1981) approach for the one sample and two sample t tests. The generalized procedure is…

Methodology · Statistics 2018-05-04 Yongqiang Tang

Unsupervised domain adaptation (UDA) is a statistical learning problem when the distribution of training (source) data is different from that of test (target) data. In this setting, one has access to labeled data only from the source domain…

Machine Learning · Computer Science 2026-02-24 Seonghwi Kim , Sung Ho Jo , Wooseok Ha , Minwoo Chae

While machine learning is rapidly being developed and deployed in health settings such as influenza prediction, there are critical challenges in using data from one environment in another due to variability in features; even within disease…

Machine Learning · Statistics 2020-03-10 Vishwali Mhasawade , Nabeel Abdur Rehman , Rumi Chunara

Uplift modeling estimates the causal effect of an intervention as the difference between potential outcomes under treatment and control, whereas counterfactual identification aims to recover the joint distribution of these potential…

Machine Learning · Computer Science 2025-12-10 Théo Verhelst , Gianluca Bontempi

We propose a method that performs anomaly detection and localisation within heterogeneous data using a pairwise undirected mixed graphical model. The data are a mixture of categorical and quantitative variables, and the model is learned…

Machine Learning · Statistics 2016-07-21 Romain Laby , François Roueff , Alexandre Gramfort

Detecting anomalies in large sets of observations is crucial in various applications, such as epidemiological studies, gene expression studies, and systems monitoring. We consider settings where the units of interest result in multiple…

Methodology · Statistics 2025-12-22 Ivo V. Stoepker , Rui M. Castro , Ery Arias-Castro

Anomaly detection is facing with emerging challenges in many important industry domains, such as cyber security and online recommendation and advertising. The recent trend in these areas calls for anomaly detection on time-evolving data…

Machine Learning · Computer Science 2019-07-16 Zheng Gao , Lin Guo , Chi Ma , Xiao Ma , Kai Sun , Hang Xiang , Xiaoqiang Zhu , Hongsong Li , Xiaozhong Liu

Fair and unbiased machine learning is an important and active field of research, as decision processes are increasingly driven by models that learn from data. Unfortunately, any biases present in the data may be learned by the model,…

Machine Learning · Computer Science 2020-02-27 Matthew J. Vowels , Necati Cihan Camgoz , Richard Bowden

Emotion annotation is inherently subjective and cognitively demanding, producing signals that reflect diverse perceptions across annotators rather than a single ground truth. In continuous affect prediction, this variability is typically…

Machine Learning · Computer Science 2026-04-09 Kosmas Pinitas , Ilias Maglogiannis

Tables are an abundant form of data with use cases across all scientific fields. Real-world datasets often contain anomalous samples that can negatively affect downstream analysis. In this work, we only assume access to contaminated data…

Machine Learning · Computer Science 2023-07-25 Guy Zamberg , Moshe Salhov , Ofir Lindenbaum , Amir Averbuch

In many practices, scientists are particularly interested in detecting which of the predictors are truly associated with a multivariate response. It is more accurate to model multiple responses as one vector rather than separating each…

Methodology · Statistics 2021-11-16 Xiaotian Dai , Guifang Fu , Randall Reese , Shaofei Zhao , Zuofeng Shang

This paper addresses detecting anomalous patterns in images, time-series, and tensor data when the location and scale of the pattern is unknown a priori. The multiscale scan statistic convolves the proposed pattern with the image at various…

Statistics Theory · Mathematics 2018-06-22 James Sharpnack

Despite their successes, deep neural networks may make unreliable predictions when faced with test data drawn from a distribution different to that of the training data, constituting a major problem for AI safety. While this has recently…

Machine Learning · Computer Science 2020-07-16 Erik Daxberger , José Miguel Hernández-Lobato

We perform differential expression analysis of high-throughput sequencing count data under a Bayesian nonparametric framework, removing sophisticated ad-hoc pre-processing steps commonly required in existing algorithms. We propose to use…

Applications · Statistics 2017-05-04 Siamak Zamani Dadaneh , Xiaoning Qian , Mingyuan Zhou

The inaccessibility of controlled randomized trials due to inherent constraints in many fields of science has been a fundamental issue in causal inference. In this paper, we focus on distinguishing the cause from effect in the bivariate…

Machine Learning · Statistics 2021-02-23 Jean-Francois Ton , Dino Sejdinovic , Kenji Fukumizu

Analysis of covariance is a crucial method for improving precision of statistical tests for factor effects in randomized experiments. However, existing solutions suffer from one or more of the following limitations: (i) they are not…

Methodology · Statistics 2024-12-24 Konstantin Emil Thiel , Paavo Sattler , Arne C Bathke , Georg Zimmermann

Extending rank-based inference to a multivariate setting such as multiple-output regression or MANOVA with unspecified d-dimensional error density has remained an open problem for more than half a century. None of the many solutions…

Statistics Theory · Mathematics 2025-10-20 Marc Hallin , Daniel Hlubinka , Šárka Hudecová

Diffusion models have emerged as a key pillar of foundation models in visual domains. One of their critical applications is to universally solve different downstream inverse tasks via a single diffusion prior without re-training for each…

Machine Learning · Computer Science 2023-10-03 Morteza Mardani , Jiaming Song , Jan Kautz , Arash Vahdat

Factor analysis provides a canonical framework for imposing lower-dimensional structure such as sparse covariance in high-dimensional data. High-dimensional data on the same set of variables are often collected under different conditions,…

Methodology · Statistics 2024-08-27 Noirrit Kiran Chandra , David B. Dunson , Jason Xu

Conducting genome-wide association studies (GWAS) in copy number variation (CNV) level is a field where few people involves and little statistical progresses have been achieved, traditional methods suffer from many problems such as batch…

Methodology · Statistics 2020-11-17 Han Wang , Changhu Wang , Linjie Wu , Ruibin Xi