中文
相关论文

相关论文: Temporal Wasserstein Imputation: A Versatile Metho…

200 篇论文

We study data-driven decision problems where historical observations are generated by a time-evolving distribution whose consecutive shifts are bounded in Wasserstein distance. We address this nonstationarity using a distributionally robust…

最优化与控制 · 数学 2025-12-25 Dominic S. T. Keehan , Edward J. Anderson , Wolfram Wiesemann

Missing data is a fundamental challenge in data science, significantly hindering analysis and decision-making across a wide range of disciplines, including healthcare, bioinformatics, social science, e-commerce, and industrial monitoring.…

机器学习 · 统计学 2026-05-12 Jicong Fan

This article proposes a Bayesian nonparametric method for forecasting, imputation, and clustering in sparsely observed, multivariate time series data. The method is appropriate for jointly modeling hundreds of time series with widely…

统计方法学 · 统计学 2019-02-27 Feras A. Saad , Vikash K. Mansinghka

Inspired by recent advancements in large language models (LLMs) for Natural Language Processing (NLP), there has been a surge in research focused on developing foundational models for time series forecasting. One approach involves training…

机器学习 · 计算机科学 2024-11-19 Andrei Chernov

In this paper, we present an ensemble data assimilation paradigm over a Riemannian manifold equipped with the Wasserstein metric. Unlike the Eulerian penalization of error in the Euclidean space, the Wasserstein metric can capture…

统计方法学 · 统计学 2021-10-11 Sagar K. Tamang , Ardeshir Ebtehaj , Peter J. Van Leeuwen , Dongmian Zou , Gilad Lerman

Anomaly detection in multivariate time series data is of paramount importance for ensuring the efficient operation of large-scale systems across diverse domains. However, accurately detecting anomalies in such data poses significant…

The Wasserstein distance is a distance between two probability distributions and has recently gained increasing popularity in statistics and machine learning, owing to its attractive properties. One important approach to extending this…

统计方法学 · 统计学 2022-02-14 Ryo Okano , Masaaki Imaizumi

Distributional ambiguity sets provide quantifiable ways to characterize the uncertainty about the true probability distribution of random variables of interest. This makes them a key element in data-driven robust optimization by exploiting…

最优化与控制 · 数学 2019-09-26 Dimitris Boskos , Jorge Cortés , Sonia Martínez

Imputation of missing values is a strategy for handling non-responses in surveys or data loss in measurement processes, which may be more effective than ignoring them. When the variable represents a count, the literature dealing with this…

应用统计 · 统计学 2020-07-31 Gilma Hernández-Herrera , Albert Navarro , David Moriña

In this work we study systems consisting of a group of moving particles. In such systems, often some important parameters are unknown and have to be estimated from observed data. Such parameter estimation problems can often be solved via a…

应用统计 · 统计学 2023-07-11 Chen Cheng , Linjie Wen , Jinglai Li

Multivariate time series with missing values are common in areas such as healthcare and finance, and have grown in number and complexity over the years. This raises the question whether deep learning methodologies can outperform classical…

机器学习 · 统计学 2020-02-21 Vincent Fortuin , Dmitry Baranchuk , Gunnar Rätsch , Stephan Mandt

Time series forecasting using historical data has been an interesting and challenging topic, especially when the data is corrupted by missing values. In many industrial problem, it is important to learn the inference function between the…

机器学习 · 计算机科学 2023-06-02 Trang H. Tran , Lam M. Nguyen , Kyongmin Yeo , Nam Nguyen , Dzung Phan , Roman Vaculin , Jayant Kalagnanam

Data imputation is a critical step in data pre-processing, particularly for datasets with missing or unreliable values. This study introduces a novel quantum-inspired imputation framework evaluated on the UCI Diabetes dataset, which…

量子物理 · 物理学 2025-05-13 Nishikanta Mohanty , Bikash K. Behera , Badshah Mukherjee , Christopher Ferrie

Dynamical systems governed by ordinary differential equations (ODEs) serve as models for a vast number of natural and social phenomena. In this work, we offer a fresh perspective on the classical problem of imputing missing time series…

机器学习 · 计算机科学 2025-03-17 Patrick Seifner , Kostadin Cvejoski , Antonia Körner , Ramsés J. Sánchez

Spatiotemporal data mining plays an important role in air quality monitoring, crowd flow modeling, and climate forecasting. However, the originally collected spatiotemporal data in real-world scenarios is usually incomplete due to sensor…

机器学习 · 计算机科学 2023-02-21 Mingzhe Liu , Han Huang , Hao Feng , Leilei Sun , Bowen Du , Yanjie Fu

Sensor data streams occur widely in various real-time applications in the context of the Internet of Things (IoT). However, sensor data streams feature missing values due to factors such as sensor failures, communication errors, or depleted…

数据库 · 计算机科学 2023-11-15 Xiao Li , Huan Li , Hua Lu , Christian S. Jensen , Varun Pandey , Volker Markl

Probabilistic forecasting of multivariate time series is essential for various downstream tasks. Most existing approaches rely on the sequences being uniformly spaced and aligned across all variables. However, real-world multivariate time…

机器学习 · 计算机科学 2025-02-18 Yijun Li , Cheuk Hang Leung , Qi Wu

The challenge of handling missing data is widespread in modern data analysis, particularly during the preprocessing phase and in various inferential modeling tasks. Although numerous algorithms exist for imputing missing data, the…

统计方法学 · 统计学 2024-03-28 Marcos Matabuena , Carla Díaz-Louzao , Rahul Ghosal , Francisco Gude

Events in the world may be caused by other, unobserved events. We consider sequences of events in continuous time. Given a probability model of complete sequences, we propose particle smoothing---a form of sequential importance…

机器学习 · 计算机科学 2019-05-15 Hongyuan Mei , Guanghui Qin , Jason Eisner

The problem of missing data, usually absent incurated and competition-standard datasets, is an unfortunate reality for most machine learning models used in industry applications. Recent work has focused on understanding the nature and the…