English
Related papers

Related papers: Two-sample KS test with approxQuantile in Apache S…

200 papers

The digital era has seen a marked increase in financial fraud. edge ML emerged as a promising solution for smartphone payment services fraud detection, enabling the deployment of ML models directly on edge devices. This approach enables a…

Computational Engineering, Finance, and Science · Computer Science 2024-01-11 Nader Karayanni , Robert J. Shahla , Chieh-Lien Hsiao

We study the problem of closeness testing for continuous distributions and its implications for causal discovery. Specifically, we analyze the sample complexity of distinguishing whether two multidimensional continuous distributions are…

Machine Learning · Computer Science 2025-03-11 Fateme Jamshidi , Sina Akbari , Negar Kiyavash

The Apache Spark stack has enabled fast large-scale data processing. Despite a rich library of statistical models and inference algorithms, it does not give domain users the ability to develop their own models. The emergence of…

Databases · Computer Science 2017-10-10 Zhuoyue Zhao , Jialing Pei , Eric Lo , Kenny Q. Zhu , Chris Liu

Simulation-based techniques such as variants of stochastic Runge-Kutta are the de facto approach for inference with stochastic differential equations (SDEs) in machine learning. These methods are general-purpose and used with parametric and…

Machine Learning · Computer Science 2021-11-01 Arno Solin , Ella Tamir , Prakhar Verma

The univariate quantile-quantile (Q-Q) plot is a well-known graphical tool for examining whether two data sets are generated from the same distribution or not. It is also used to determine how well a specified probability distribution fits…

Statistics Theory · Mathematics 2014-07-07 Subhra Sankar Dhar , Biman Chakraborty , Probal Chaudhuri

A Kernel Adaptive Metropolis-Hastings algorithm is introduced, for the purpose of sampling from a target distribution with strongly nonlinear support. The algorithm embeds the trajectory of the Markov chain into a reproducing kernel Hilbert…

Machine Learning · Statistics 2014-06-16 Dino Sejdinovic , Heiko Strathmann , Maria Lomeli Garcia , Christophe Andrieu , Arthur Gretton

Suppose one has access to oracles generating samples from two unknown probability distributions P and Q on some N-element set. How many samples does one need to test whether the two distributions are close or far from each other in the…

Quantum Physics · Physics 2011-12-01 Sergey Bravyi , Aram W. Harrow , Avinatan Hassidim

This paper presents a convergence analysis of a Krylov subspace spectral (KSS) method applied to a 1-D wave equation in an inhomogeneous medium. It will be shown that for sufficiently regular initial data, this KSS method yields…

Numerical Analysis · Mathematics 2023-07-20 Bailey Rester , Anzhelika Vasilyeva , James V. Lambers

This paper proposes a novel approach to generate samples from target distributions that are difficult to sample from using Markov Chain Monte Carlo (MCMC) methods. Traditional MCMC algorithms often face slow convergence due to the…

Cosmology and Nongalactic Astrophysics · Physics 2023-08-11 Sandro Dias Pinto Vitenti , Eduardo J. Barroso

In this paper, we propose a unified approach to harness quantum conformal methods for multi-output distributions, with a particular emphasis on two experimental paradigms: (i) a standard 2-qubit circuit scenario producing a four-dimensional…

Quantum Physics · Physics 2025-01-22 Emre Tasar

Sampling is a fundamental algorithmic task in wide-ranging applications across multiple disciplines such as scientific computing, statistics and machine learning. In this paper, an efficient stochastic Runge-Kutta scheme is proposed to…

Statistics Theory · Mathematics 2026-05-27 Haotian Lin , Xiaojie Wang , Xiaoyan Zhang

Learning from imbalanced data is among the most challenging areas in contemporary machine learning. This becomes even more difficult when considered the context of big data that calls for dedicated architectures capable of high-performance…

Machine Learning · Computer Science 2022-11-16 William C. Sleeman , Bartosz Krawczyk

Perfect sampling is a technique that uses coupling arguments to provide a sample from the stationary distribution of a Markov chain in a finite time without ever computing the distribution. This technique is very efficient if all the events…

Discrete Mathematics · Computer Science 2015-03-17 Ana Bušić , Bruno Gaujal , Furcy Pin

We propose and demonstrate a novel, effective approach to slice sampling. Using the probability integral transform, we first generalize Neal's shrinkage algorithm, standardizing the procedure to an automatic and universal starting point:…

Computation · Statistics 2025-06-16 Matthew J. Heiner , Samuel B. Johnson , Joshua R. Christensen , David B. Dahl

We introduce the Kernel Calibration Conditional Stein Discrepancy test (KCCSD test), a non-parametric, kernel-based test for assessing the calibration of probabilistic models with well-defined scores. In contrast to previous methods, our…

Machine Learning · Statistics 2025-10-17 Pierre Glaser , David Widmann , Fredrik Lindsten , Arthur Gretton

This paper investigates a statistical procedure for testing the equality of two independently estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

Methodology · Statistics 2020-07-13 Rémy Mariétan , Stephan Morgenthaler

Modern kernel-based two-sample tests have shown great success in distinguishing complex, high-dimensional distributions with appropriate learned kernels. Previous work has demonstrated that this kernel learning procedure succeeds, assuming…

Machine Learning · Statistics 2022-01-06 Feng Liu , Wenkai Xu , Jie Lu , Danica J. Sutherland

Quantum computing, with its potential to enhance various machine learning tasks, allows significant advancements in kernel calculation and model precision. Utilizing the one-class Support Vector Machine alongside a quantum kernel, known for…

Machine learning (ML) models, such as SVM, for tasks like classification and clustering of sequences, require a definition of distance/similarity between pairs of sequences. Several methods have been proposed to compute the similarity…

Machine Learning · Computer Science 2022-09-13 Sarwan Ali , Bikram Sahoo , Muhammad Asad Khan , Alexander Zelikovsky , Imdad Ullah Khan , Murray Patterson

Classical supply chain risk models treat node failures as statistically independent events, systematically underestimating cascade probabilities when supplier dependencies are strongly correlated. At n=40 nodes, the full correlated failure…

Quantum Physics · Physics 2026-04-02 Sumit Tapas Chongder