English
Related papers

Related papers: Temporal Wasserstein Imputation: A Versatile Metho…

200 papers

Missing data is a common problem in real-world settings and particularly relevant in healthcare applications where researchers use Electronic Health Records (EHR) and results of observational studies to apply analytics methods. This issue…

Machine Learning · Statistics 2018-12-04 Dimitris Bertsimas , Agni Orfanoudaki , Colin Pawlowski

In this paper, we expand upon the theory of trend filtering by introducing the use of the Wasserstein metric as a means to control the amount of spatiotemporal variation in filtered time series data. While trend filtering utilizes…

Signal Processing · Electrical Eng. & Systems 2019-10-25 Erdem Varol , Amin Nejatbakhsh

This paper introduces a novel iterative method for missing data imputation that sequentially reduces the mutual information between data and the corresponding missingness mask. Inspired by GAN-based approaches that train generators to…

Machine Learning · Statistics 2025-11-26 Jiahao Yu , Qizhen Ying , Leyang Wang , Ziyue Jiang , Song Liu

Multivariate time series data suffer from the problem of missing values, which hinders the application of many analytical methods. To achieve the accurate imputation of these missing values, exploiting inter-correlation by employing the…

Machine Learning · Computer Science 2024-09-17 Kohei Obata , Koki Kawabata , Yasuko Matsubara , Yasushi Sakurai

Data values in a dataset can be missing or anomalous due to mishandling or human error. Analysing data with missing values can create bias and affect the inferences. Several analysis methods, such as principle components analysis or…

Artificial Intelligence · Computer Science 2022-05-11 Sandeep Hans , Diptikalyan Saha , Aniya Aggarwal

Detecting relevant changes in dynamic time series data in a timely manner is crucially important for many data analysis tasks in real-world settings. Change point detection methods have the ability to discover changes in an unsupervised…

Artificial Intelligence · Computer Science 2022-01-19 Kamil Faber , Roberto Corizzo , Bartlomiej Sniezynski , Michael Baron , Nathalie Japkowicz

Advancements in data collection techniques and the heterogeneity of data resources can yield high percentages of missing observations on variables, such as block-wise missing data. Under missing-data scenarios, traditional methods such as…

Methodology · Statistics 2022-05-17 Wei Lan , Xuerong Chen , Tao Zou , Chih-Ling Tsai

This paper considers the problem of regression over distributions, which is becoming increasingly important in machine learning. Existing approaches often ignore the geometry of the probability space or are computationally expensive. To…

Machine Learning · Computer Science 2025-10-31 Maksim Maslov , Alexander Kugaevskikh , Matthew Ivanov

We present a simple yet novel time series imputation technique with the goal of constructing an irregular time series that is uniform across every sample in a data set. Specifically, we fix a grid defined by the midpoints of non-overlapping…

Machine Learning · Computer Science 2022-01-19 Andrew Baumgartner , Sevda Molani , Qi Wei , Jennifer Hadlock

Adversarial examples are crafted by adding indistinguishable perturbations to normal examples in order to fool a well-trained deep learning model to misclassify. In the context of computer vision, this notion of indistinguishability is…

Machine Learning · Computer Science 2023-03-23 Wenjie Wang , Li Xiong , Jian Lou

Randomised signature has been proposed as a flexible and easily implementable alternative to the well-established path signature. In this article, we employ randomised signature to introduce a generative model for financial time series data…

Machine Learning · Computer Science 2024-09-09 Francesca Biagini , Lukas Gonon , Niklas Walter

Diffusion models (DMs) have gained attention in Missing Data Imputation (MDI), but there remain two long-neglected issues to be addressed: (1). Inaccurate Imputation, which arises from inherently sample-diversification-pursuing generative…

Machine Learning · Computer Science 2024-06-25 Zhichao Chen , Haoxuan Li , Fangyikang Wang , Odin Zhang , Hu Xu , Xiaoyu Jiang , Zhihuan Song , Eric H. Wang

Sensor data has been playing an important role in machine learning tasks, complementary to the human-annotated data that is usually rather costly. However, due to systematic or accidental mis-operations, sensor data comes very often with a…

Machine Learning · Computer Science 2017-11-22 Jingguang Zhou , Zili Huang

Errors are prevalent in time series data, such as GPS trajectories or sensor readings. Existing methods focus more on anomaly detection but not on repairing the detected anomalies. By simply filtering out the dirty data via anomaly…

Databases · Computer Science 2020-03-30 Aoqian Zhang , Shaoxu Song , Jianmin Wang , Philip S. Yu

Wasserstein distances provide a powerful framework for comparing data distributions. They can be used to analyze processes over time or to detect inhomogeneities within data. However, simply calculating the Wasserstein distance or analyzing…

Machine Learning · Computer Science 2026-03-03 Philip Naumann , Jacob Kauffmann , Grégoire Montavon

Imputation methods play a critical role in enhancing the quality of practical time-series data, which often suffer from pervasive missing values. Recently, diffusion-based generative imputation methods have demonstrated remarkable success…

Machine Learning · Computer Science 2025-10-03 Zeqi Ye , Minshuo Chen

In this paper, we study statistical inference for the Wasserstein distance, which has attracted much attention and has been applied to various machine learning tasks. Several studies have been proposed in the literature, but almost all of…

Machine Learning · Statistics 2022-01-21 Vo Nguyen Le Duy , Ichiro Takeuchi

In real-world scenarios like traffic and energy, massive time-series data with missing values and noises are widely observed, even sampled irregularly. While many imputation methods have been proposed, most of them work with a local…

Machine Learning · Computer Science 2024-06-03 Shikai Fang , Qingsong Wen , Yingtao Luo , Shandian Zhe , Liang Sun

The prevalence of missing values in data science poses a substantial risk to any further analyses. Despite a wealth of research, principled nonparametric methods to deal with general non-monotone missingness are still scarce. Instead,…

Machine Learning · Statistics 2026-05-04 Gitte Kremling , Jeffrey Näf , Johannes Lederer

We consider a data-driven robust hypothesis test where the optimal test will minimize the worst-case performance regarding distributions that are close to the empirical distributions with respect to the Wasserstein distance. This leads to a…

Statistics Theory · Mathematics 2021-06-01 Liyan Xie , Rui Gao , Yao Xie