English
Related papers

Related papers: Highly Irregular Functional Generalized Linear Reg…

200 papers

For the past decades, face recognition (FR) has been actively studied in computer vision and pattern recognition society. Recently, due to the advances in deep learning, the FR technology shows high performance for most of the benchmark…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Hyung-Il Kim , Kimin Yun , Yong Man Ro

Evaluating treatment effect heterogeneity widely informs treatment decision making. At the moment, much emphasis is placed on the estimation of the conditional average treatment effect via flexible machine learning algorithms. While these…

Methodology · Statistics 2021-05-07 Lihua Lei , Emmanuel J. Candès

In the present paper we investigate the predictive risk of possibly misspecified quantile regression functions. The in-sample risk is well-known to be an overly optimistic estimate of the predictive risk and we provide two relatively simple…

Statistics Theory · Mathematics 2018-11-05 Alexander Giessing , Xuming He

Benign overfitting is a phenomenon in machine learning where a model perfectly fits (interpolates) the training data, including noisy examples, yet still generalizes well to unseen data. Understanding this phenomenon has attracted…

Machine Learning · Computer Science 2025-05-20 Junhyung Park , Patrick Bloebaum , Shiva Prasad Kasiviswanathan

The analysis of multivariate functional curves has the potential to yield important scientific discoveries in domains such as healthcare, medicine, economics and social sciences. However, it is common for real-world settings to present…

Methodology · Statistics 2024-07-23 Tui Nolan , Sylvia Richardson , Hélène Ruffieux

Semiparametric accelerated failure time (AFT) models directly relate the predicted failure times to covariates and are a useful alternative to models that work on the hazard function or the survival function. For case-cohort data, much less…

Computation · Statistics 2022-12-15 Steven Chiou , Sangwook Kang , Jun Yan

This article describes a multivariate polynomial regression method where the uncertainty of the input parameters are approximated with Gaussian distributions, derived from the central limit theorem for large weighted sums, directly from the…

Machine Learning · Statistics 2013-10-04 Peter Kovesarki , Ian C. Brock

Newton-step approximations to pseudo maximum likelihood estimates of spatial autoregressive models with a large number of parameters are examined, in the sense that the parameter space grows slowly as a function of sample size. These have…

Econometrics · Economics 2021-05-25 Abhimanyu Gupta

Overfitting is a phenomenon that occurs when a machine learning model is trained for too long and focused too much on the exact fitness of the training samples to the provided training labels and cannot keep track of the predictive rules…

Machine Learning · Computer Science 2025-09-22 Nuri Korhan , Samet Bayram

In this paper, we present a novel and effective inference approach to conduct both finite- and large-sample inference for high-dimensional linear regression models. This approach is developed under the so-called repro samples framework, in…

Methodology · Statistics 2025-12-01 Peng Wang , Min-Ge Xie , Linjun Zhang

The infrequent occurrence of overfit in deep neural networks is perplexing. On the one hand, theory predicts that as models get larger they should eventually become too specialized for a specific training set, with ensuing decrease in…

Machine Learning · Computer Science 2023-12-29 Uri Stern , Daphna Weinshall

A time series represents a set of observations collected over time. Typically, these observations are captured with a uniform sampling frequency (e.g. daily). When data points are observed in uneven time intervals the time series is…

Machine Learning · Computer Science 2022-01-03 Pedro Costa , Vitor Cerqueira , João Vinagre

Synthetic data has been proposed as a solution to address the issue of high-quality data scarcity in the training of large language models (LLMs). Studies have shown that synthetic data can effectively improve the performance of LLMs on…

Computation and Language · Computer Science 2024-06-19 Jie Chen , Yupeng Zhang , Bingning Wang , Wayne Xin Zhao , Ji-Rong Wen , Weipeng Chen

Traditionally in regression one minimizes the number of fitting parameters or uses smoothing/regularization to trade training (TE) and generalization error (GE). Driving TE to zero by increasing fitting degrees of freedom (dof) is expected…

Machine Learning · Computer Science 2019-06-11 Partha P Mitra

Estimating the conditional average treatment effect (CATE) from observational data plays a crucial role in areas such as e-commerce, healthcare, and economics. Existing studies mainly rely on the strong ignorability assumption that there…

Machine Learning · Computer Science 2025-01-28 Chuan Zhou , Yaxuan Li , Chunyuan Zheng , Haiteng Zhang , Haoxuan Li , Mingming Gong

Irregular time series, where data points are recorded at uneven intervals, are prevalent in healthcare settings, such as emergency wards where vital signs and laboratory results are captured at varying times. This variability, which…

Machine Learning · Computer Science 2024-10-16 Hrishikesh Patel , Ruihong Qiu , Adam Irwin , Shazia Sadiq , Sen Wang

We consider a class of semiparametric regression models which are one-parameter extensions of the Cox [J. Roy. Statist. Soc. Ser. B 34 (1972) 187-220] model for right-censored univariate failure times. These models assume that the hazard…

Statistics Theory · Mathematics 2007-06-13 Michael R. Kosorok , Bee Leng Lee , Jason P. Fine

With advances in digital technology, the classification of medical images has become a crucial step for image-based clinical decision support systems. Automatic medical image classification represents a pivotal domain where the use of AI…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Abu Adnan Sadi , Labib Chowdhury , Nusrat Jahan , Mohammad Newaz Sharif Rafi , Radeya Chowdhury , Faisal Ahamed Khan , Nabeel Mohammed

A reliable fault diagnosis system should not only accurately classify known health states but also effectively identify unknown faults. In multimode processes, samples belonging to the same health state often show multiple cluster…

Machine Learning · Computer Science 2025-11-13 Guangqiang Li , M. Amine Atoui , Xiangshun Li

We consider the problem of precision matrix estimation where, due to extraneous confounding of the underlying precision matrix, the data are independent but not identically distributed. While such confounding occurs in many scientific…

Machine Learning · Statistics 2019-07-01 Sinong Geng , Mladen Kolar , Oluwasanmi Koyejo