中文
相关论文

相关论文: First-Mover Bias in Gradient Boosting Explanations…

200 篇论文

A fundamental problem in supervised learning is to find a good set of features or distance measures. If the new set of features is of lower dimensionality and can be obtained by a simple transformation of the original data, they can make…

机器学习 · 计算机科学 2024-05-15 Anri Patron , Ayush Prasad , Hoang Phuc Hau Luu , Kai Puolamäki

Variance reduction is a family of powerful mechanisms for stochastic optimization that appears to be helpful in many machine learning tasks. It is based on estimating the exact gradient with some recursive sequences. Previously, many papers…

最优化与控制 · 数学 2025-11-07 Aleksandr Shestakov , Valery Parfenov , Aleksandr Beznosikov

Variable selection in the linear regression model takes many apparent faces from both frequentist and Bayesian standpoints. In this paper we introduce a variable selection method referred to as a rescaled spike and slab model. We study the…

统计理论 · 数学 2007-06-13 Hemant Ishwaran , J. Sunil Rao

This work proposes a framework, embedded within the Performance Estimation framework (PEP), for obtaining worst-case performance guarantees on stochastic first-order methods. Given a first-order method, a function class, and a noise model…

最优化与控制 · 数学 2026-01-05 Anne Rubbens , Sébastien Colla , Julien M. Hendrickx

Stacking is a general approach for combining multiple models toward greater predictive accuracy. It has found various application across different domains, ensuing from its meta-learning nature. Our understanding, nevertheless, on how and…

机器学习 · 计算机科学 2019-01-29 Nino Arsov , Martin Pavlovski , Ljupco Kocarev

We study the trade-offs between convergence rate and robustness to gradient errors in designing a first-order algorithm. We focus on gradient descent (GD) and accelerated gradient (AG) methods for minimizing strongly convex functions when…

最优化与控制 · 数学 2019-11-07 Necdet Serhat Aybat , Alireza Fallah , Mert Gurbuzbalaban , Asuman Ozdaglar

A curious phenomenon observed in some dynamical generative models is the following: despite learning errors in the score function or the drift vector field, the generated samples appear to shift \emph{along} the support of the data…

机器学习 · 计算机科学 2025-08-12 Nisha Chandramoorthy , Adriaan de Clercq

Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting their value as synthetic augmentation for rare conditions. We trace this to low head-versus-tail…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Felix Nützel , Mischa Dombrowski , Bernhard Kainz

We study first-hitting times in Differential Evolution (DE) through a conditional hazard frame work. Instead of analyzing convergence via Markov-chain transition kernels or drift arguments, we ex press the survival probability of a…

神经与进化计算 · 计算机科学 2026-01-19 Dimitar Nedanovski , Svetoslav Nenov , Dimitar Pilev

The activation memory required for exact backpropagation scales linearly with network depth, context length, and feature dimensionality, forming an O(L * BN ) spatial bottleneck (where B is the sequence-batch cardinality and N is the…

机器学习 · 计算机科学 2026-04-21 Vladimer Khasia

We consider the problem of minimizing a strongly convex smooth function where the gradients are subject to additive worst-case deterministic errors that are square-summable. We study the trade-offs between the convergence rate and…

最优化与控制 · 数学 2023-10-23 Mert Gurbuzbalaban

Pretrained diffusion models have revolutionized real-world image super-resolution (Real-ISR) but suffer from computational bottlenecks due to iterative sampling. Recent single-step distillation accelerates inference but faces a stark…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Shyang-En Weng , Yi-Cheng Liao , Yu-Syuan Xu , Wei-Chen Chiu , Ching-Chun Huang

Gradient boosted trees are competition-winning, general-purpose, non-parametric regressors, which exploit sequential model fitting and gradient descent to minimize a specific loss function. The most popular implementations are tailored to…

机器学习 · 计算机科学 2022-08-23 Lorenzo Nespoli , Vasco Medici

We consider random perturbations of discrete-time dynamical systems. We give sufficient conditions for the stochastic stability of certain classes of maps, in a strong sense. This improves the main result in J. F. Alves, V. Araujo, Random…

动力系统 · 数学 2010-03-01 Jose F. Alves , Helder Vilarinho

Data augmentation has been an indispensable tool to improve the performance of deep neural networks, however the augmentation can hardly transfer among different tasks and datasets. Consequently, a recent trend is to adopt AutoML technique…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Aoming Liu , Zehao Huang , Zhiwu Huang , Naiyan Wang

This paper continues the systematic investigation of diffusive shear instabilities initiated in Part I of this series. In this work, we primarily focus on quantifying the impact of non-local mixing, which is not taken into account in Zahn's…

太阳与恒星天体物理 · 物理学 2018-07-25 D. Gagnier , P. Garaud

Spectral gradient methods, such as the Muon optimizer, modify gradient updates by preserving directional information while discarding scale, and have shown strong empirical performance in deep learning. We investigate the mechanisms…

机器学习 · 统计学 2026-02-02 Guillaume Braun , Han Bao , Wei Huang , Masaaki Imaizumi

Currently, widely used first-order deep learning optimizers include non-adaptive learning rate optimizers and adaptive learning rate optimizers. The former is represented by SGDM (Stochastic Gradient Descent with Momentum), while the latter…

机器学习 · 计算机科学 2024-09-25 Honglin Qin , Hongye Zheng , Bingxing Wang , Zhizhong Wu , Bingyao Liu , Yuanfang Yang

This dissertation explores the impact of bias in deep neural networks and presents methods for reducing its influence on model performance. The first part begins by categorizing and describing potential sources of bias and errors in data…

机器学习 · 计算机科学 2023-08-21 Agnieszka Mikołajczyk-Bareła

Learning a Bayesian network (BN) from data can be useful for decision-making or discovering causal relationships. However, traditional methods often fail in modern applications, which exhibit a larger number of observed variables than data…

统计计算 · 统计学 2018-06-26 Raj Agrawal , Tamara Broderick , Caroline Uhler