English
Related papers

Related papers: Splitting Localization and Prediction Numbers

200 papers

Predict-then-Optimize is a framework for using machine learning to perform decision-making under uncertainty. The central research question it asks is, "How can the structure of a decision-making task be used to tailor ML models for that…

Machine Learning · Computer Science 2024-02-20 Sanket Shah , Andrew Perrault , Bryan Wilder , Milind Tambe

Identifiability, and the closely related idea of partial smoothness, unify classical active set methods and more general notions of solution structure. Diverse optimization algorithms generate iterates in discrete time that are eventually…

Optimization and Control · Mathematics 2024-06-24 Adrian Lewis , Tonghua Tian

The connection between dependency trees and spanning trees is exploited by the NLP community to train and to decode graph-based dependency parsers. However, the NLP literature has missed an important difference between the two structures:…

Computation and Language · Computer Science 2020-10-08 Ran Zmigrod , Tim Vieira , Ryan Cotterell

In this paper we analyze a hash function for $k$-partitioning a set into bins, obtaining strong concentration bounds for standard algorithms combining statistics from each bin. This generic method was originally introduced by Flajolet and…

Data Structures and Algorithms · Computer Science 2016-02-16 Søren Dahlgaard , Mathias Bæk Tejs Knudsen , Eva Rotenberg , Mikkel Thorup

We propose a modification that corrects for split-improvement variable importance measures in Random Forests and other tree-based methods. These methods have been shown to be biased towards increasing the importance of features with more…

Machine Learning · Statistics 2020-03-25 Zhengze Zhou , Giles Hooker

Optimization is widely used in statistics, and often efficiently delivers point estimates on useful spaces involving structural constraints or combinatorial structure. To quantify uncertainty, Gibbs posterior exponentiates the negative loss…

Methodology · Statistics 2025-07-23 Cheng Zeng , Eleni Dilma , Jason Xu , Leo L Duan

Problem 1.5.7 from Pitman's Saint-Flour lecture notes: Does there exist for each n a fragmentation process (\Pi_{n,k}, 1 \leq k \leq n) taking values in the space of partitions of {1,2,...,n} such that \Pi_{n,k} is distributed like the…

Probability · Mathematics 2007-12-05 Christina Goldschmidt , James B. Martin , Dario Spanò

We consider a probability distribution on the set of Boolean functions in n variables which is induced by random Boolean expressions. Such an expression is a random rooted plane tree where the internal vertices are labelled with connectives…

Combinatorics · Mathematics 2015-09-28 Antoine Genitrini , Bernhard Gittenberger , Veronika Kraus , Cécile Mailler

Set prediction is about learning to predict a collection of unordered variables with unknown interrelations. Training such models with set losses imposes the structure of a metric space over sets. We focus on stochastic and underdefined…

Machine Learning · Computer Science 2021-02-23 David W. Zhang , Gertjan J. Burghouts , Cees G. M. Snoek

Let F be a uniformly distributed random k-SAT formula with n variables and m clauses. Non-constructive arguments show that F is satisfiable for clause/variable ratios m/n< r(k)~2^k ln 2 with high probability. Yet no efficient algorithm is…

Combinatorics · Mathematics 2017-11-29 Amin Coja-Oghlan

The estimation of probability densities based on available data is a central task in many statistical applications. Especially in the case of large ensembles with many samples or high-dimensional sample spaces, computationally efficient…

Methodology · Statistics 2017-05-04 Daniel W. Meyer

Standard clustering techniques assume a common configuration for all features in a dataset. However, when dealing with multi-view or longitudinal data, the clusters' number, frequencies, and shapes may need to vary across features to…

Methodology · Statistics 2025-03-26 Beatrice Franzolini , Maria De Iorio , Johan Eriksson

The problem of categorical data analysis in high dimensions is considered. A discussion of the fundamental difficulties of probability modeling is provided, and a solution to the derivation of high dimensional probability distributions…

Machine Learning · Computer Science 2017-08-24 Cetin Savkli , J. Ryan Carr , Philip Graff , Lauren Kennell

Starting from a discussion of the concrete representations of the coordinates of the k-Minkowski spacetime (in 1+1 dimensions, for simplicity), we explicitly compute the associated Weyl operators as functions of a pair of Schroedinger…

High Energy Physics - Theory · Physics 2010-02-22 Ludwik Dabrowski , Gherardo Piacitelli

We initiate a novel approach to explain the predictions and out of sample performance of random forest (RF) regression and classification models by exploiting the fact that any RF can be mathematically formulated as an adaptive weighted K…

We address the task of estimating multiple trajectories from unlabeled data. This problem arises in many settings, one could think of the construction of maps of transport networks from passive observation of travellers, or the…

Statistics Theory · Mathematics 2016-11-07 Matthew Thorpe , Adam M. Johansen

We consider a Galton-Watson tree where each node is marked independently of each others with a probability depending on itsout-degree. Using a penalization method, we exhibit new martingales where the number of marks up to level n -- 1…

Probability · Mathematics 2024-03-04 Romain Abraham , Sonia Boulal , Pierre Debs

Plucking polynomial of a plane rooted tree with a delay function $\alpha$ was introduced in 2014 by J.H.~Przytycki. As shown in this paper, plucking polynomial factors when $\alpha$ satisfies additional conditions. We use this result and…

Geometric Topology · Mathematics 2022-09-14 Mieczyslaw K. Dabkowski , Cheyu Wu

With recent developments in remote sensing technologies, plot-level forest resources can be predicted utilizing airborne laser scanning (ALS). The prediction is often assisted by mostly vertical summaries of the ALS point clouds. We present…

Applications · Statistics 2021-05-03 Henrike Häbel , András Balázs , Mari Myllymäki

Gorman and Bedrick (2019) argued for using random splits rather than standard splits in NLP experiments. We argue that random splits, like standard splits, lead to overly optimistic performance estimates. We can also split data in biased or…

Computation and Language · Computer Science 2021-04-27 Anders Søgaard , Sebastian Ebert , Jasmijn Bastings , Katja Filippova