中文
相关论文

相关论文: A Mixed-effects Model for Incomplete Data With Bat…

200 篇论文

A data mixture refers to how different data sources are combined to train large language models, and selecting an effective mixture is crucial for optimal downstream performance. Existing methods either conduct costly searches directly on…

机器学习 · 计算机科学 2026-05-07 Jingwei Li , Xinran Gu , Jingzhao Zhang

Biclustering is a powerful data mining technique that allows simultaneously clustering rows (observations) and columns (features) in a matrix-format data set, which can provide results in a checkerboard-like pattern for visualization and…

统计方法学 · 统计学 2021-06-09 Binhuan Wang , Lanqiu Yao , Jiyuan Hu , Huilin Li

Missing data are a common problem for both the construction and implementation of a prediction algorithm. Pattern mixture kernel submodels (PMKS) - a series of submodels for every missing data pattern that are fit using only data from that…

统计方法学 · 统计学 2017-04-27 Sarah Fletcher Mercaldo , Jeffrey D. Blume

Complex biological processes are usually experimented along time among a collection of individuals. Longitudinal data are then available and the statistical challenge is to better understand the underlying biological mechanisms. The…

统计理论 · 数学 2015-06-11 Pierre Barbillon , Célia Barthélémy , Adeline Samson

Target trial emulation (TTE) enables causal questions to be studied with observational data when randomized controlled trials (RCTs) are infeasible. Yet treatment-effect methods often address causal estimation, missingness, and temporal…

For multi-source data, blocks of variable information from certain sources are likely missing. Existing methods for handling missing data do not take structures of block-wise missing data into consideration. In this paper, we propose a…

统计方法学 · 统计学 2020-04-07 Fei Xue , Annie Qu

Stochastic gradient descent-based algorithms are widely used for training deep neural networks but often suffer from slow convergence. To address the challenge, we leverage the framework of the alternating direction method of multipliers…

机器学习 · 计算机科学 2025-02-03 Ouya Wang , Shenglong Zhou , Geoffrey Ye Li

We present a nonparametric Bayesian joint model for multivariate continuous and categorical variables, with the intention of developing a flexible engine for multiple imputation of missing values. The model fuses Dirichlet process mixtures…

应用统计 · 统计学 2015-10-14 Jared S. Murray , Jerome P. Reiter

Dropout represents a typical issue to be addressed when dealing with longitudinal studies. If the mechanism leading to missing information is non-ignorable, inference based on the observed data only may be severely biased. A frequent…

统计方法学 · 统计学 2018-03-23 Maria Francesca Marino , Marco Alfo'

Regression analysis with missing data is a long-standing and challenging problem, particularly when there are many missing variables with arbitrary missing patterns. Likelihood-based methods, although theoretically appealing, are often…

统计方法学 · 统计学 2024-10-16 Ngok Sang Kwok , Kin Yau Wong

We evaluate the performance of targeted maximum likelihood estimation (TMLE) for estimating the average treatment effect in missing data scenarios under varying levels of positivity violations. We employ model- and design-based simulations,…

统计方法学 · 统计学 2026-05-12 Christoph Wiederkehr , Christian Heumann , Michael Schomaker

Data augmentation plays a crucial role in enhancing the robustness and performance of machine learning models across various domains. In this study, we introduce a novel mixed-sample data augmentation method called RandoMix. RandoMix is…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Xiaoliang Liu , Furao Shen , Jian Zhao , Changhai Nie

Estimating conditional dependence graphs and precision matrices are some of the most common problems in modern statistics and machine learning. When data are fully observed, penalized maximum likelihood-type estimators have become standard…

机器学习 · 统计学 2019-04-09 Roger Fan , Byoungwook Jang , Yuekai Sun , Shuheng Zhou

Bridging the gap between internal and external validity is crucial for heterogeneous treatment effect estimation. Randomised controlled trials (RCTs), favoured for their internal validity due to randomisation, often encounter challenges in…

统计方法学 · 统计学 2026-04-03 Evangelos Dimitriou , Edwin Fong , Jens Magelund Tarp , Karla Diaz-Ordaz , Brieuc Lehmann

This paper develops an inferential framework for matrix completion when missing is not at random and without the requirement of strong signals. Our development is based on the observation that if the number of missing entries is small…

统计方法学 · 统计学 2023-08-07 Jungjun Choi , Ming Yuan

Analysis of multivariate healthcare time series data is inherently challenging: irregular sampling, noisy and missing values, and heterogeneous patient groups with different dynamics violating exchangeability. In addition, interpretability…

机器学习 · 计算机科学 2023-11-15 Onur Poyraz , Pekka Marttinen

We consider the task of identifying and estimating a parameter of interest in settings where data is missing not at random (MNAR). In general, such parameters are not identified without strong assumptions on the missing data model. In this…

统计方法学 · 统计学 2024-02-29 Zixiao Wang , AmirEmad Ghassami , Ilya Shpitser

Processing high-volume, streaming data is increasingly common in modern statistics and machine learning, where batch-mode algorithms are often impractical because they require repeated passes over the full dataset. This has motivated…

Integrating deep learning with latent state space models has the potential to yield temporal models that are powerful, yet tractable and interpretable. Unfortunately, current models are not designed to handle missing data or multiple data…

机器学习 · 计算机科学 2019-11-25 Tan Zhi-Xuan , Harold Soh , Desmond C. Ong

Matrix completion tackles the task of predicting missing values in a low-rank matrix based on a sparse set of observed entries. It is often assumed that the observation pattern is generated uniformly at random or has a very specific…

机器学习 · 统计学 2025-03-18 Yudong Chen , Xumei Xi , Christina Lee Yu
‹ 上一页 1 8 9 10 下一页 ›