Related papers: Wavelet-Packet Content for Positive Operators
Solving ill-posed inverse problems necessitates effective regularization strategies to stabilize the inversion process against measurement noise. While classical methods like Tikhonov regularization require heuristic parameter tuning, and…
In distributed optimization problems, a technique called gradient coding, which involves replicating data points, has been used to mitigate the effect of straggling machines. Recent work has studied approximate gradient coding, which…
Many websites with an underlying database containing structured data provide the richest and most dense source of information relevant for topical data integration. The real data integration requires sustainable and reliable pattern…
We study contractive projections, isometries, and real positive maps on algebras of operators on a Hilbert space. For example we find generalizations and variants of certain classical results on contractive projections on C*-algebras and…
We study mixed-state localization operators from the perspective of Werner's operator convolutions which allows us to extend known results from the rank-one case to trace class operators. The idea of localizing a signal to a domain in phase…
We investigate compactness properties of weighted summation operators $V_{\alpha,\sigma}$ as mapping from $\ell_1(T)$ into $\ell_q(T)$ for some $q\in (1,\infty)$. Those operators are defined by $$ (V_{\alpha,\sigma} x)(t)…
A tree-packing is a collection of spanning trees of a graph. It has been a useful tool for computing the minimum cut in static, dynamic, and distributed settings. In particular, [Thorup, Comb. 2007] used them to obtain his dynamic min-cut…
Decision trees are a popular technique in statistical data classification. They recursively partition the feature space into disjoint sub-regions until each sub-region becomes homogeneous with respect to a particular class. The basic…
Parameter-efficient fine-tuning (PEFT) of pre-trained foundation models is increasingly attracting interest in medical imaging due to its effectiveness and computational efficiency. Among these methods, Low-Rank Adaptation (LoRA) is a…
Scalable packet classification is a key requirement to support scalable network applications like firewalls, intrusion detection, and differentiated services. With ever increasing in the line-rate in core networks, it becomes a great…
Qualitative and quantitative aspects for variational inequalities governed by strongly pseudomonotone operators on Hilbert space are investigated in this paper. First, we establish a global error bound for the solution set of the given…
Symbolic regression aims to recover closed-form expressions from numerical data, but in differentiable symbolic regression the recovered expression depends not only on the grammar but also on the fixed architecture through which variables…
We introduce an adaptive tree search algorithm, that can find high-scoring outputs under translation models that make no assumptions about the form or structure of the search objective. This algorithm -- a deterministic variant of Monte…
General treebank analyses are graph structured, but parsers are typically restricted to tree structures for efficiency and modeling reasons. We propose a new representation and algorithm for a class of graph structures that is flexible…
Sequential fine-tuning of pretrained language encoders often overwrites previously acquired capabilities, but the forgetting behavior of parameter-efficient updates remains under-characterized. We present a controlled empirical study of…
In this article we study quantitative rigidity properties for the compatible and incompatible two-state problems for suitable classes of $\mathcal{A}$-free operators and for a singularly perturbed $T_3$-structure for the divergence…
Let $\{(X_i,Y_i)\}_{i\in \{1,..., n\}}$ be an i.i.d. sample from the random design regression model $Y=f(X)+\epsilon$ with $(X,Y)\in [0,1]\times [-M,M]$. In dealing with such a model, adaptation is naturally to be intended in terms of…
When trained on language data, do transformers learn some arbitrary computation that utilizes the full capacity of the architecture or do they learn a simpler, tree-like computation, hypothesized to underlie compositional meaning systems…
We provide a simple recipe for obtaining all self-adjoint extensions, together with their resolvent, of the symmetric operator $S$ obtained by restricting the self-adjoint operator $A:\D(A)\subseteq\H\to\H$ to the dense, closed with respect…
We study the fundamental optimization principles of self-attention, the defining mechanism of transformers, by analyzing the implicit bias of gradient-based optimizers in training a self-attention layer with a linear decoder in binary…