English
Related papers

Related papers: A critical comparison of handling zeros in high-di…

200 papers

Dealing with missing data in data analysis is inevitable. Although powerful imputation methods that address this problem exist, there is still much room for improvement. In this study, we examined single imputation based on deep…

Machine Learning · Computer Science 2020-04-07 Najmeh Abiri , Björn Linse , Patrik Edén , Mattias Ohlsson

The frequency scarcity imposed by fast growing demand for mobile data service requires promising spectrum aggregation systems. The so-called higher-order statistics (HOS) of the channel capacity is a suitable metric on the system…

Information Theory · Computer Science 2016-12-07 Jiayi Zhang , Xiaoyu Chen , Kostas P. Peppas , Xu Li , Ying Liu

This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…

Methodology · Statistics 2024-09-18 D. S. Martinez-Lobo , O. O. Melo , N. A. Cruz

Mediation analysis seeks to understand the mechanism by which a treatment affects an outcome. Count or zero-inflated count outcome are common in many studies in which mediation analysis is of interest. For example, in dental studies,…

Methodology · Statistics 2016-07-12 Zijian Guo , Dylan S. Small , Stuart A. Gansky , Jing Cheng

This work presents a sum-of-squares (SOS) based framework to perform data-driven stabilization and robust control tasks on discrete-time linear systems where the full-state observations are corrupted by L-infinity bounded input,…

Optimization and Control · Mathematics 2023-03-31 Jared Miller , Tianyu Dai , Mario Sznaier

A key problem in computational sustainability is to understand the distribution of species across landscapes over time. This question gives rise to challenging large-scale prediction problems since (i) hundreds of species have to be…

Machine Learning · Computer Science 2020-11-02 Shufeng Kong , Junwen Bai , Jae Hee Lee , Di Chen , Andrew Allyn , Michelle Stuart , Malin Pinsky , Katherine Mills , Carla P. Gomes

In this paper, we investigate the optimal statistical performance and the impact of computational constraints for independent component analysis (ICA). Our goal is twofold. On the one hand, we characterize the precise role of dimensionality…

Statistics Theory · Mathematics 2023-04-03 Arnab Auddy , Ming Yuan

Standard approaches for variable selection in linear models are not tailored to deal properly with high-dimensional and incomplete data. Currently, methods dedicated to high-dimensional data handle missing values by ad-hoc strategies, like…

Methodology · Statistics 2021-06-09 Avner Bar-Hen , Vincent Audigier

The development of data-driven heart sound classification models has been an active area of research in recent years. To develop such data-driven models in the first place, heart sound signals need to be captured using a signal acquisition…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-15 Davoud Shariat Panah , Andrew Hines , Susan McKeever

Topological data analysis is a relatively new branch of machine learning that excels in studying high dimensional data, and is theoretically known to be robust against noise. Meanwhile, data objects with mixed numeric and categorical…

Algebraic Topology · Mathematics 2020-06-15 Chengyuan Wu , Carol Anne Hargreaves

Topological data analysis (TDA) aims to extract noise-robust features from a data set by examining the number and persistence of holes in its topology. We show that a computational problem closely related to a core task in TDA --…

Quantum Physics · Physics 2024-10-29 Casper Gyurik , Alexander Schmidhuber , Robbie King , Vedran Dunjko , Ryu Hayakawa

Persistent homology is an area within topological data analysis (TDA) that can uncover different dimensional holes (connected components, loops, voids, etc.) in data. The holes are characterized, in part, by how long they persist across…

Methodology · Statistics 2025-04-08 Sixtus Dakurah , Jessi Cisewski-Kehe

The advent of plant phenomics, coupled with the wealth of genotypic data generated by next-generation sequencing technologies, provides exciting new resources for investigations into and improvement of complex traits. However, these new…

Genomics · Quantitative Biology 2019-04-30 Gota Morota , Diego Jarquin , Malachy T. Campbell , Hiroyoshi Iwata

Harmonic drive systems (HDS) are high-precision robotic transmissions featuring compact size and high gear ratios. However, issues like kinematic transmission errors hamper their precision performance. This article focuses on data-driven…

Robotics · Computer Science 2023-12-14 Ju Wu

Tactical selection of experiments to estimate an underlying model is an innate task across various fields. Since each experiment has costs associated with it, selecting statistically significant experiments becomes necessary. Classic linear…

Optimization and Control · Mathematics 2021-03-30 Raj K. Velicheti , Amber Srivastava , Srinivasa M. Salapaka

This study analyzes the impact of heterogeneity ("Variety") in Big Data by comparing classification strategies across structured (Epsilon) and unstructured (Rest-Mex, IMDB) domains. A dual methodology was implemented: evolutionary and…

A common problem faced by statistical institutes is that data may be missing from collected data sets. The typical way to overcome this problem is to impute the missing data. The problem of imputing missing data is complicated by the fact…

Applications · Statistics 2014-01-09 Jeroen Pannekoek , Natalie Shlomo , Ton De Waal

Sequencing-based technologies provide an abundance of high-dimensional biological datasets with skewed and zero-inflated measurements. Classification of such data with linear discriminant analysis leads to poor performance due to the…

Methodology · Statistics 2022-08-09 Hee Cheol Chung , Yang Ni , Irina Gaynanova

Missing values are largely inevitable in gene expression microarray studies. Data sets often have significant omissions due to individuals dropping out of experiments, errors in data collection, image corruptions, and so on. Missing data…

Quantitative Methods · Quantitative Biology 2018-09-18 Marie Li

Data mining and machine learning techniques such as classification and regression trees (CART) represent a promising alternative to conventional logistic regression for propensity score estimation. Whereas incomplete data preclude the…

Machine Learning · Statistics 2018-07-26 Bas B. L. Penning de Vries , Maarten van Smeden , Rolf H. H. Groenwold