English
Related papers

Related papers: On Statistical Properties of A Veracity Scoring Me…

200 papers

We consider the problem of fitting a relationship (e.g. a potential scientific law) to data involving multiple variables. Ordinary (least squares) regression is not suitable for this because the estimated relationship will differ according…

Methodology · Statistics 2024-09-05 Chris Tofallis

This paper introduces and analyzes a framework that accommodates general heterogeneity in regression modeling. It demonstrates that regression models with fixed or time-varying parameters can be estimated using the OLS and time-varying OLS…

Econometrics · Economics 2025-11-11 Liudas Giraitis , George Kapetanios , Yufei Li , Alexia Ventouri

We develop a data-driven approach for signal denoising that utilizes variational mode decomposition (VMD) algorithm and Cramer Von Misses (CVM) statistic. In comparison with the classical empirical mode decomposition (EMD), VMD enjoys…

Signal Processing · Electrical Eng. & Systems 2020-06-02 Khuram Naveed , Muhammad Tahir Akhtar , Muhammad Faisal Siddiqui , Naveed ur Rehman

In this work we study the problem of measuring the fairness of a machine learning model under noisy information. Focusing on group fairness metrics, we investigate the particular but common situation when the evaluation requires controlling…

Machine Learning · Computer Science 2021-05-24 Flavien Prost , Pranjal Awasthi , Nick Blumm , Aditee Kumthekar , Trevor Potter , Li Wei , Xuezhi Wang , Ed H. Chi , Jilin Chen , Alex Beutel

Evaluating the performance of causal discovery algorithms that aim to find causal relationships between time-dependent processes remains a challenging topic. In this paper, we show that certain characteristics of datasets, such as…

Artificial Intelligence · Computer Science 2025-08-12 Christopher Lohse , Jonas Wahl

Many modern datasets, such as those in ecology and geology, are composed of samples with spatial structure and dependence. With such data violating the usual independent and identically distributed (IID) assumption in machine learning and…

Methodology · Statistics 2023-10-18 Kevin Fry , Jonathan E. Taylor

The use of simulated data in the field of causal discovery is ubiquitous due to the scarcity of annotated real data. Recently, Reisach et al., 2021 highlighted the emergence of patterns in simulated linear data, which displays increasing…

Methodology · Statistics 2023-10-24 Francesco Montagna , Nicoletta Noceti , Lorenzo Rosasco , Francesco Locatello

Dynamic models of occupancy patterns have shown to be effective in optimizing building-systems operations. Previous research has relied on CO$_2$ sensors and vision-based techniques to determine occupancy patterns. Vision-based techniques…

Machine Learning · Computer Science 2022-03-10 Mahsa Pahlavikhah Varnosfaderani , Arsalan Heydarian , Farrokh Jazizadeh

High dimensional Vector Autoregressions (VAR) have received a lot of interest recently due to novel applications in health, engineering, finance and the social sciences. Three issues arise when analyzing VAR's: (a) The high dimensional…

Statistics Theory · Mathematics 2022-11-15 Sagnik Halder , George Michailidis

Online knowledge repositories typically rely on their users or dedicated editors to evaluate the reliability of their content. These evaluations can be viewed as noisy measurements of both information reliability and information source…

Social and Information Networks · Computer Science 2017-04-04 Behzad Tabibian , Isabel Valera , Mehrdad Farajtabar , Le Song , Bernhard Schölkopf , Manuel Gomez-Rodriguez

Remotely sensed data are sparse, which means that data have missing values, for instance due to cloud cover. This is problematic for applications and signal processing algorithms that require complete data sets. To address the sparse data…

The Princeton Variability Survey (PVS) is a robotic survey which makes use of readily available, ``off-the-shelf'' type hardware products, in conjunction with a powerful set of commercial software products, in order to monitor and discover…

Astrophysics · Physics 2009-11-07 Cullen Blake

The "crowd" has become a very important geospatial data provider. Subsumed under the term Volunteered Geographic Information (VGI), non-expert users have been providing a wealth of quantitative geospatial data online. With spatial reasoning…

Databases · Computer Science 2014-08-27 Georgios Skoumas , Dieter Pfoser , Anastasios Kyrillidis

We propose a new, computationally efficient, sparsity adaptive changepoint estimator for detecting changes in unknown subsets of a high-dimensional data sequence. Assuming the data sequence is Gaussian, we prove that the new method…

Methodology · Statistics 2023-11-27 Per August Jarval Moen , Ingrid Kristine Glad , Martin Tveten

Given full or partial information about a collection of points that lie close to a union of several subspaces, subspace clustering refers to the process of clustering the points according to their subspace and identifying the subspaces. One…

Machine Learning · Statistics 2018-01-16 Zachary Charles , Amin Jalali , Rebecca Willett

Spatial statistics is concerned with the analysis of data that have spatial locations associated with them, and those locations are used to model statistical dependence between the data. The spatial data are treated as a single realisation…

Methodology · Statistics 2022-02-09 Noel Cressie , Matthew Sainsbury-Dale , Andrew Zammit-Mangion

Shapley data valuation provides a principled, axiomatic framework for assigning importance to individual datapoints, and has gained traction in dataset curation, pruning, and pricing. However, it is a combinatorial measure that requires…

Machine Learning · Computer Science 2025-11-05 Rodrigo Mendoza-Smith

We analyze the performance of a data-assimilation algorithm based on a linear feedback control when used with observational data that contains measurement errors. Our model problem consists of dynamics governed by the two-dimension…

Analysis of PDEs · Mathematics 2015-06-19 Hakima Bessaih , Eric Olson , E. S. Titi

The Meteorology is a field where huge amounts of data are generated, mainly collected by sensors at weather stations, where different variables can be measured. Those data have some particularities such as high volume and dimensionality,…

Machine Learning · Computer Science 2025-10-27 Shadi Aljawarneh , Juan A. Lara , Muneer Bani Yassein

Stochastic models are widely used to verify whether systems satisfy their reliability, performance and other nonfunctional requirements. However, the validity of the verification depends on how accurately the parameters of these models can…

Software Engineering · Computer Science 2022-02-22 Naif Alasmari , Radu Calinescu , Colin Paterson , Raffaela Mirandola