English
Related papers

Related papers: Perturbation and scaled Cook's distance

200 papers

Subsampling methods have been recently proposed to speed up least squares estimation in large scale settings. However, these algorithms are typically not robust to outliers or corruptions in the observed covariates. The concept of influence…

Machine Learning · Statistics 2014-06-20 Brian McWilliams , Gabriel Krummenacher , Mario Lucic , Joachim M. Buhmann

Testing the independence between random vectors is a fundamental problem in statistics. Distance correlation, a recently popular dependence measure, is universally consistent for testing independence against all distributions with finite…

Methodology · Statistics 2024-08-22 Yuwei Ke , Hok Kan Ling , Yanglei Song

Statistical analysis of large and sparse graphs is a challenging problem in data science due to the high dimensionality and nonlinearity of the problem. This paper presents a fast and scalable algorithm for partitioning such graphs into…

Data Structures and Algorithms · Computer Science 2018-12-24 Hannu Reittu , Lasse Leskelä , Tomi Räty , Marco Fiorucci

Measures of spike train synchrony have become important tools in both experimental and theoretical neuroscience. Three time-resolved measures called the ISI-distance, the SPIKE-distance, and SPIKE-synchronization have already been…

Data Analysis, Statistics and Probability · Physics 2020-01-14 Eero Satuvuori , Irene Malvestio , Thomas Kreuz

Model-Based Diagnosis deals with the identification of the real cause of a system's malfunction based on a formal system model and observations of the system behavior. When a malfunction is detected, there is usually not enough information…

Artificial Intelligence · Computer Science 2017-11-16 Patrick Rodler , Wolfgang Schmid , Konstantin Schekotihin

We consider the inference problem for parameters in stochastic differential equation models from discrete time observations (e.g. experimental or simulation data). Specifically, we study the case where one does not have access to…

Numerical Analysis · Mathematics 2018-04-10 Sebastian Krumscheid

With current high precision collider data, the reliable estimation of theoretical uncertainties due to missing higher orders (MHOs) in perturbation theory has become a pressing issue for collider phenomenology. Traditionally, the size of…

High Energy Physics - Phenomenology · Physics 2021-09-30 Claude Duhr , Alexander Huss , Aleksas Mazeliauskas , Robert Szafron

Measurement is a fundamental building block of numerous scientific models and their creation. This is in particular true for data driven science. Due to the high complexity and size of modern data sets, the necessity for the development of…

Artificial Intelligence · Computer Science 2022-04-26 Tom Hanika , Johannes Hirth

Accurate estimation for extent of cross{sectional dependence in large panel data analysis is paramount to further statistical analysis on the data under study. Grouping more data with weak relations (cross{sectional dependence) together…

Econometrics · Economics 2019-04-16 Jiti Gao , Guangming Pan , Yanrong Yang , Bo Zhang

The double slit experiment provides a classic example of both interference and the effect of observation in quantum physics. When particles are sent individually through a pair of slits, a wave-like interference pattern develops, but no…

Quantum Physics · Physics 2016-07-01 Joshua Kincaid , Kyle McLelland , Michael Zwolak

The paper algorithmizes the problem of regime change point identification for data measured in a system exhibiting impulsive behaviors. This is a fundamental challenge for annotation of measurement data relevant, e.g., for designing…

Quantifying the distance between datasets is a fundamental question in mathematics and machine learning. We propose \textit{magnitude distance}, a novel distance metric defined on finite datasets using the notion of the \emph{magnitude} of…

Machine Learning · Computer Science 2026-02-10 Sahel Torkamani , Henry Gouk , Rik Sarkar

The concept of information has emerged as a language in its own right, bridging several disciplines that analyze natural phenomena and man-made systems. Integrated information has been introduced as a metric to quantify the amount of…

Neurons and Cognition · Quantitative Biology 2019-06-10 Alberto Hernández-Espinosa , Héctor Zenil , Narsis A. Kiani , Jesper Tegnér

Data-collapse is a way of establishing scaling and extracting associated exponents in problems showing self-similar or self-affine characteristics as e.g. in equilibrium or non-equilibrium phase transitions, in critical phases, in dynamics…

Soft Condensed Matter · Physics 2009-11-07 Somendra M. Bhattacharjee , Flavio Seno

The gradual patterns that model the complex co-variations of attributes of the form "The more/less X, The more/less Y" play a crucial role in many real world applications where the amount of numerical data to manage is important, this is…

Machine Learning · Computer Science 2020-05-25 Michaël Chirmeni Boujike , Jerry Lonlac , Norbert Tsopze , Engelbert Mephu Nguifo

Study samples often differ from the target populations of inference and policy decisions in non-random ways. Researchers typically believe that such departures from random sampling -- due to changes in the population over time and space, or…

Methodology · Statistics 2023-07-20 Tamara Broderick , Ryan Giordano , Rachael Meager

By means of the present geometrical and dynamical observational data, it is very hard to establish, from a statistical perspective, a clear preference among the vast majority of the proposed models for the dynamical dark energy and/or…

Cosmology and Nongalactic Astrophysics · Physics 2017-11-07 Tomasz Denkiewicz , Vincenzo Salzano

In community detection on graphs, the semi-supervised learning problem entails inferring the ground-truth membership of each node in a graph, given the connectivity structure and a limited number of revealed node labels. Different subsets…

Disordered Systems and Neural Networks · Physics 2022-03-22 Hugo Cui , Luca Saglietti , Lenka Zdeborová

One fundamental statistical question for research areas such as precision medicine and health disparity is about discovering effect modification of treatment or exposure by observed covariates. We propose a semiparametric framework for…

Methodology · Statistics 2020-08-04 Muxuan Liang , Menggang Yu

In order for clinicians to manage disease progression and make effective decisions about drug dosage, treatment regimens or scheduling follow up appointments, it is necessary to be able to identify both short and long-term trends in…

Quantitative Methods · Quantitative Biology 2016-12-06 Norman Poh , Simon Bull , Santosh Tirunagari , Nicholas Cole , Simon de Lusignan