English
Related papers

Related papers: Internal Data Imputation in Data Warehouse Dimensi…

200 papers

Methods to handle missing data have been extensively explored in the context of estimation and descriptive studies, with multiple imputation being the most widely used method in clinical research. However, in the context of clinical risk…

Methodology · Statistics 2024-11-25 Junhui Mi , Rahul D. Tendulkar , Sarah M. C. Sittenfeld , Sujata Patil , Emily C. Zabor

In this paper, we suggest a multi-dimensional approach towards intrusion detection. Network and system usage parameters like source and destination IP addresses; source and destination ports; incoming and outgoing network traffic data rate…

Cryptography and Security · Computer Science 2012-05-11 Manoj Rameshchandra Thakur , Sugata Sanyal

Multimodal data, where different types of data are collected from the same subjects, are fast emerging in a large variety of scientific applications. Factor analysis is commonly used in integrative analysis of multimodal data, and is…

Statistics Theory · Mathematics 2021-03-31 Quefeng Li , Lexin Li

We propose a multifidelity dimension reduction method to identify a low-dimensional structure present in many engineering models. The structure of interest arises when functions vary primarily on a low-dimensional subspace of the…

Numerical Analysis · Mathematics 2020-01-08 Rémi Lam , Olivier Zahm , Youssef Marzouk , Karen Willcox

We endeavour to estimate numerous multi-dimensional means of various probability distributions on a common space based on independent samples. Our approach involves forming estimators through convex combinations of empirical means derived…

Machine Learning · Statistics 2025-03-11 Gilles Blanchard , Jean-Baptiste Fermanian , Hannah Marienwald

The size of datasets has been increasing rapidly both in terms of number of variables and number of events. As a result, the empty space phenomenon and the curse of dimensionality complicate the extraction of useful information. But, in…

Data Analysis, Statistics and Probability · Physics 2015-05-07 Jean Golay , Mikhail Kanevski

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events,…

Methodology · Statistics 2023-01-12 Tuo Wang , Rachel Zilinskas , Ying Li , Yongming Qu

This article presents the implementation process of a Data Warehouse and a multidimensional analysis of business data for a holding company in the financial sector. The goal is to create a business intelligence system that, in a simple,…

Databases · Computer Science 2017-09-19 Jose Ferreira , Fernando Almeida , Jose Monteiro

Data imputation is crucial for addressing challenges posed by missing values in multivariate time series data across various fields, such as healthcare, traffic, and economics, and has garnered significant attention. Among various methods,…

Machine Learning · Computer Science 2025-01-14 Chunjing Xiao , Xue Jiang , Xianghe Du , Wei Yang , Wei Lu , Xiaomin Wang , Kevin Chetty

An efficient monotone data augmentation (MDA) algorithm is proposed for missing data imputation for incomplete multivariate nonnormal data that may contain variables of different types, and are modeled by a sequence of regression models…

Methodology · Statistics 2018-11-21 Yongqiang Tang

Data from discovery proteomic and phosphoproteomic experiments typically include missing values that correspond to proteins that have not been identified in the analyzed sample. Replacing the missing values with random numbers, a process…

Quantitative Methods · Quantitative Biology 2019-10-01 Matus Medo , Daniel M. Aebersold , Michaela Medova

It is time to renew old ways of thinking about dimensional analysis. Specifically, more than $n-r$ invariants and more than one functional relation between invariants need to be considered simultaneously. Thus generalized, dimensional…

History and Overview · Mathematics 2014-11-12 Dan Jonsson

Multivariate categorical data nested within households often include reported values that fail edit constraints---for example, a participating household reports a child's age as older than his biological parent's age---as well as missing…

Methodology · Statistics 2018-09-21 Olanrewaju Akande , Andrés Barrientos , Jerome P. Reiter

Various applications leverage location data to increase transparency, efficiency, and safety in intralogistics. There are several properties of location data, such as the data's degrees of freedom, system latency, update rate, or accuracy.…

Systems and Control · Electrical Eng. & Systems 2023-04-25 Jakob Schyga , Markus Knitt , Johannes Hinckeldeyn , Jochen Kreutzfeldt

We present an approach for imputation of missing items in multivariate categorical data nested within households. The approach relies on a latent class model that (i) allows for household level and individual level variables, (ii) ensures…

Methodology · Statistics 2018-07-05 Olanrewaju Akande , Jerome Reiter , Andrés F. Barrientos

We consider the problem of reconstructing missing data on a smooth manifold from incomplete and nonuniform samples. While classical methods for manifold approximation typically assume quasi-uniform data, their performance deteriorates…

Numerical Analysis · Mathematics 2026-04-15 David Levin

One of the purposes of Big Data systems is to support analysis of data gathered from heterogeneous data sources. Since data warehouses have been used for several decades to achieve the same goal, they could be leveraged also to provide…

Databases · Computer Science 2018-09-13 Darja Solodovnikova , Laila Niedrite

The quality of training data for knowledge discovery in databases (KDD) and data mining depends upon many factors, but handling missing values is considered to be a crucial factor in overall data quality. Today real world datasets contains…

Databases · Computer Science 2009-04-22 Shariq Bashir , Saad Razzaq , Umer Maqbool , Sonya Tahir , Abdul Rauf Baig

The intrinsic dimensionality refers to the ``true'' dimensionality of the data, as opposed to the dimensionality of the data representation. For example, when attributes are highly correlated, the intrinsic dimensionality can be much lower…

Machine Learning · Statistics 2020-11-30 Erik Thordsen , Erich Schubert

The concept of dimension is essential to grasp the complexity of data. A naive approach to determine the dimension of a dataset is based on the number of attributes. More sophisticated methods derive a notion of intrinsic dimension (ID)…

Machine Learning · Computer Science 2023-04-18 Maximilian Stubbemann , Tom Hanika , Friedrich Martin Schneider