English
Related papers

Related papers: High-Dimensional Importance-Weighted Information C…

200 papers

In multi-objective optimization problems, there might exist hidden objectives that are important to the decision-maker but are not being optimized. On the other hand, there might also exist irrelevant objectives that are being optimized but…

Optimization and Control · Mathematics 2022-06-06 Seyed Mahdi Shavarani , Manuel López-Ibáñez , Richard Allmendinger

We study the fundamental problem of calibrating a linear binary classifier of the form $\sigma(\hat{w}^\top x)$, where the feature vector $x$ is Gaussian, $\sigma$ is a link function, and $\hat{w}$ is an estimator of the true linear weight…

Statistics Theory · Mathematics 2025-10-03 Yufan Li , Pragya Sur

Optimal experimental design (OED) concerns itself with identifying ideal methods of data collection, e.g.~via sensor placement. The \emph{greedy algorithm}, that is, placing one sensor at a time, in an iteratively optimal manner, stands as…

Optimization and Control · Mathematics 2025-10-15 Christian Aarset

Learning to optimize is an approach that leverages training data to accelerate the solution of optimization problems. Many approaches use unrolling to parametrize the update step and learn optimal parameters. Although L2O has shown…

Optimization and Control · Mathematics 2025-07-15 Patrick Fahy , Mohammad Golbabaee , Matthias J. Ehrhardt

Importance weighting is widely applicable in machine learning in general and in techniques dealing with data covariate shift problems in particular. A novel, direct approach to determine such importance weighting is presented. It relies on…

Machine Learning · Computer Science 2021-02-05 Marco Loog

We study the numerical integration problem for functions with infinitely many variables. The function spaces of integrands we consider are weighted reproducing kernel Hilbert spaces with norms related to the ANOVA decomposition of the…

Numerical Analysis · Mathematics 2021-09-21 Josef Dick , Michael Gnewuch

Full Waveform Inversion (FWI) is a standard algorithm in seismic imaging. Its implementation requires the a priori choice of a number of "design parameters", such as the positions of sensors for the actual measurements and one (or more)…

Numerical Analysis · Mathematics 2024-06-25 Shaunagh Downing , Silvia Gazzola , Ivan G. Graham , Euan A. Spence

Effective features can improve the performance of a model, which can thus help us understand the characteristics and underlying structure of complex data. Previous feature selection methods usually cannot keep more local structure…

Machine Learning · Computer Science 2019-10-10 Xia Wu , Xueyuan Xu , Jianhong Liu , Hailing Wang , Bin Hu , Feiping Nie

A common problem in data analysis is the separation of signal and background. We revisit and generalise the so-called $sWeights$ method, which allows one to calculate an empirical estimate of the signal density of a control variable using a…

Methodology · Statistics 2022-08-24 Hans Dembinski , Matthew Kenzie , Christoph Langenbruch , Michael Schmelling

We consider the problem of heteroscedastic linear regression, where, given $n$ samples $(\mathbf{x}_i, y_i)$ from $y_i = \langle \mathbf{w}^{*}, \mathbf{x}_i \rangle + \epsilon_i \cdot \langle \mathbf{f}^{*}, \mathbf{x}_i \rangle$ with…

Machine Learning · Statistics 2023-07-04 Dheeraj Baby , Aniket Das , Dheeraj Nagaraj , Praneeth Netrapalli

We provide lower error bounds for randomized algorithms that approximate integrals of functions depending on an unrestricted or even infinite number of variables. More precisely, we consider the infinite-dimensional integration problem on…

Numerical Analysis · Mathematics 2021-02-09 Michael Gnewuch

While the inverse probability of treatment weighting (IPTW) is a commonly used approach for treatment comparisons in observational data, the resulting estimates may be subject to bias and excessively large variance when there is lack of…

Methodology · Statistics 2024-02-13 Zhiqiang Cao , Lama Ghazi , Claudia Mastrogiacomo , Laura Forastiere , F. Perry Wilson , Fan Li

A platform trial with a master protocol provides an infrastructure to ethically and efficiently evaluate multiple treatment options in multiple diseases. Given that certain study drugs can enter or exit a platform trial, the randomization…

Methodology · Statistics 2025-07-15 Tianyu Zhan , Jane Zhang , Lei Shu , Yihua Gu

The ability to design effective experiments is crucial for obtaining data that can substantially reduce the uncertainty in the predictions made using computational models. An optimal experimental design (OED) refers to the choice of a…

Methodology · Statistics 2025-06-17 Troy Butler , John Jakeman , Michael Pilosov , Scott Walsh , Timothy Wildey

A recurring challenge in high energy physics is inference of the signal component from a distribution for which observations are assumed to be a mixture of signal and background events. A standard assumption is that there exists information…

Applications · Statistics 2025-10-21 Chad Schafer , Larry Wasserman , Mikael Kuusela

The Bayesian and Akaike information criteria aim at finding a good balance between under- and over-fitting. They are extensively used every day by practitioners. Yet we contend they suffer from at least two afflictions: their penalty…

Statistics Theory · Mathematics 2026-03-20 Sylvain Sardy , Maxime van Cutsem , Sara van de Geer

Recently mean field theory has been successfully used to analyze properties of wide, random neural networks. It gave rise to a prescriptive theory for initializing feed-forward neural networks with orthogonal weights, which ensures that…

Machine Learning · Statistics 2019-06-05 Piotr A. Sokol , Il Memming Park

Top-$N$ recommender systems typically utilize side information to address the problem of data sparsity. As nowadays side information is growing towards high dimensionality, the performances of existing methods deteriorate in terms of both…

Information Retrieval · Computer Science 2017-05-17 Yifan Chen , Xiang Zhao

The large-scale multiple testing inherent to high throughput biological data necessitates very high statistical stringency and thus true effects in data are difficult to detect unless they have high effect sizes. One solution to this…

Methodology · Statistics 2017-12-21 Mohamad S. Hasan

While large-scale robot datasets have propelled recent progress in imitation learning, learning from smaller task specific datasets remains critical for deployment in new environments and unseen tasks. One such approach to few-shot…

Robotics · Computer Science 2025-09-03 Amber Xie , Rahul Chand , Dorsa Sadigh , Joey Hejna