English
Related papers

Related papers: Hierarchical Importance Sampling for Estimating Oc…

200 papers

The substantial memory demands of pre-training and fine-tuning large language models (LLMs) require memory-efficient optimization algorithms. One promising approach is layer-wise optimization, which treats each transformer block as a single…

Machine Learning · Computer Science 2026-01-15 Yuxi Liu , Renjia Deng , Yutong He , Xue Wang , Tao Yao , Kun Yuan

We consider estimation of a functional parameter of a realistically modeled data distribution based on observing independent and identically distributed observations. We define an $m$-th order Spline Highly Adaptive Lasso Minimum Loss…

Statistics Theory · Mathematics 2021-07-05 Mark J. van der Laan , David Benkeser , Weixin Cai

To schedule LLM inference, the \textit{shortest job first} (SJF) principle is favorable by prioritizing requests with short output lengths to avoid head-of-line (HOL) blocking. Existing methods usually predict a single output length for…

Machine Learning · Computer Science 2026-05-26 Haoyu Zheng , Yongqiang Zhang , Fangcheng Fu , Xiaokai Zhou , Hao Luo , Hongchao Zhu , Yuanyuan Zhu , Hao Wang , Xiao Yan , Jiawei Jiang

This paper addresses a continuous-time continuous-space chance-constrained stochastic optimal control (SOC) problem via a Hamilton-Jacobi-Bellman (HJB) partial differential equation (PDE). Through Lagrangian relaxation, we convert the…

Optimization and Control · Mathematics 2022-05-03 Apurva Patil , Alfredo Duarte , Aislinn Smith , Takashi Tanaka , Fabrizio Bisetti

Non-uniform sampling arises when an experimenter does not have full control over the sampling characteristics of the process under investigation. Moreover, it is introduced intentionally in algorithms such as Bayesian optimization and…

Machine Learning · Statistics 2020-07-03 Stijn de Waele

We present a multilevel stochastic gradient descent method for the optimal control of systems governed by partial differential equations under uncertain input data. The gradient descent method used to find the optimal control leverages a…

Optimization and Control · Mathematics 2025-06-04 Niklas Baumgarten , David Schneiderhan

Instance selection (IS) addresses the critical challenge of reducing dataset size while keeping informative characteristics, becoming increasingly important as datasets grow to millions of instances. Current IS methods often struggle with…

Machine Learning · Computer Science 2025-09-25 Zahiriddin Rustamov , Ayham Zaitouny , Nazar Zaki

Imitation learning (IL) aims to mimic the behavior of an expert policy in a sequential decision-making problem given only demonstrations. In this paper, we focus on understanding the minimax statistical limits of IL in episodic Markov…

Machine Learning · Computer Science 2020-09-15 Nived Rajaraman , Lin F. Yang , Jiantao Jiao , Kannan Ramachandran

Maximum likelihood estimation (MLE) is a well-known estimation method used in many robotic and computer vision applications. Under Gaussian assumption, the MLE converts to a nonlinear least squares (NLS) problem. Efficient solutions to NLS…

Robotics · Computer Science 2016-08-11 Viorela Ila , Lukas Polok , Marek Solony , Pavel Svoboda

Iterative Hessian sketch (IHS) is an effective sketching method for modeling large-scale data. It was originally proposed by Pilanci and Wainwright (2016; JMLR) based on randomized sketching matrices. However, it is computationally…

Machine Learning · Statistics 2020-03-10 Aijun Zhang , Hengtao Zhang , Guosheng Yin

Rather than traditional position control, impedance control is preferred to ensure the safe operation of industrial robots programmed from demonstrations. However, variable stiffness learning studies have focused on task performance rather…

Robotics · Computer Science 2023-07-31 Masashi Okada , Mayumi Komatsu , Ryo Okumura , Tadahiro Taniguchi

The paper proposes a systematic framework for building data-driven stochastic differential equation (SDE) models from sparse, noisy observations. Unlike traditional parametric approaches, which assume a known functional form for the drift,…

Machine Learning · Statistics 2025-08-18 Arnab Ganguly , Riten Mitra , Jinpu Zhou

We propose SLIM (Stochastic Learning and Inference in overidentified Models), a scalable stochastic approximation framework for nonlinear GMM. SLIM forms iterative updates from independent mini-batches of moments and their derivatives,…

Econometrics · Economics 2025-11-03 Xiaohong Chen , Min Seong Kim , Sokbae Lee , Myung Hwan Seo , Myunghyun Song

The increasing complexity of distribution network calls for advancement in distribution system state estimation (DSSE) to monitor the operating conditions more accurately. Sufficient number of measurements is imperative for a reliable and…

Applications · Statistics 2018-11-09 Mehdi Shafiei , Ghavameddin Nourbakhsh , Ali Arefi , Gerard Ledwich , Houman Pezeshki

This letter proposes a novel and highly efficient distribution system state estimation (DSSE) algorithm with nonlinear measurements from supervisory control and data acquisition (SCADA) systems. Conventional DSSE, i.e., a weighted least…

Systems and Control · Electrical Eng. & Systems 2020-01-14 Ying Zhang , Jianhui Wang

The $H_2$ norm is a commonly used performance metric in the design of estimators. However, $H_2$-optimal estimation of most PDEs is complicated by the lack of transfer function and state-space representations. To address this problem, we…

Optimization and Control · Mathematics 2026-05-19 Danio Braghini , Sachin Shivakumar , Matthew M. Peet

The challenging problem of conducting fully Bayesian inference for the reaction rate constants governing stochastic kinetic models (SKMs) is considered. Given the challenges underlying this problem, the Markov jump process representation is…

Computation · Statistics 2019-01-10 Andrew Golightly , Emma Bradley , Tom Lowe , Colin S. Gillespie

This paper presents a novel Importance Sampling (IS) scheme for estimating distribution tails of performance measures modeled with a rich set of tools such as linear programs, integer linear programs, piecewise linear/quadratic objectives,…

Machine Learning · Statistics 2023-07-11 Anand Deo , Karthyek Murthy

The likelihood-informed subspace (LIS) method offers a viable route to reducing the dimensionality of high-dimensional probability distributions arising in Bayesian inference. LIS identifies an intrinsic low-dimensional linear subspace…

Computation · Statistics 2021-10-22 Tiangang Cui , Xin T. Tong

We generalize the derivation of model predictive path integral control (MPPI) to allow for a single joint distribution across controls in the control sequence. This reformation allows for the implementation of adaptive importance sampling…

Systems and Control · Electrical Eng. & Systems 2023-03-02 Dylan M. Asmar , Ransalu Senanayake , Shawn Manuel , Mykel J. Kochenderfer