Related papers: When does third order efficiency imply fourth orde…
Recently, there has been a surge in research in multimodal machine translation (MMT), where additional modalities such as images are used to improve translation quality of textual systems. A particular use for such multimodal systems is the…
Large Language Models (LLMs) have driven significant progress, yet their growing parameter counts and context windows incur prohibitive compute, energy, and monetary costs. We introduce EfficientLLM, a novel benchmark and the first…
The task of word-level quality estimation (QE) consists of taking a source sentence and machine-generated translation, and predicting which words in the output are correct and which are wrong. In this paper, propose a method to effectively…
We apply the theory of Bruhat-Tits trees to the study of optimal embeddings of two and three dimensional commutative orders into quaternion algebras. Specifically, we determine how many conjugacy classes of global Eichler orders in a…
A key challenge of modern machine learning systems is to achieve Out-of-Distribution (OOD) generalization -- generalizing to target data whose distribution differs from that of source data. Despite its significant importance, the…
In large-scale modern data analysis, first-order optimization methods are usually favored to obtain sparse estimators in high dimensions. This paper performs theoretical analysis of a class of iterative thresholding based estimators defined…
We use a characterization of symmetry in terms of extremal order statistics which enables to build several new nonparametric tests of symmetry. We discuss their limiting distributions and calculate their local exact Bahadur efficiency under…
A preference order or ranking aggregated from pairwise comparison data is commonly understood as a strict total order. However, in real-world scenarios, some items are intrinsically ambiguous in comparisons, which may very well be an…
In this paper, we study a class of non-parametric density estimators under Bayesian settings. The estimators are piecewise constant functions on binary partitions. We analyze the concentration rate of the posterior distribution under a…
In this work at first the relation the Mittag-Lefler function to the exponential is given. The results are applied to the construction of the solution of Cauchy problem for ordinary linear operator differential equations with constant…
In this note, we consider the complexity of optimizing a highly smooth (Lipschitz $k$-th order derivative) and strongly convex function, via calls to a $k$-th order oracle which returns the value and first $k$ derivatives of the function at…
We study the problem of learning multivariate log-concave densities with respect to a global loss function. We obtain the first upper bound on the sample complexity of the maximum likelihood estimator (MLE) for a log-concave density on…
In Transformer-based neural machine translation (NMT), the positional encoding mechanism helps the self-attention networks to learn the source representation with order dependency, which makes the Transformer-based NMT achieve…
Simultaneous machine translation (SiMT) outputs translation while receiving the streaming source inputs, and hence needs a policy to determine where to start translating. The alignment between target and source words often implies the most…
In this paper we introduce new distributions which are solutions of higher-order Laplace equations. It is proved that their densities can be obtained by folding and symmetrizing Cauchy distributions. Another class of probability laws…
We consider the problem of the estimation of the mean function of an inhomogeneous Poisson process when its intensity function is periodic. For the mean integrated squared error (MISE) there is a classical lower bound for all estimators and…
The main objective of this article is to present $\nu$-fractional derivative $\mu$-differentiable functions by considering 4-parameters extended Mittag-Leffler function (MLF). We investigate that the new $\nu$-fractional derivative…
In-context learning (ICL) enables large language models to perform new tasks by conditioning on a sequence of examples. Most prior work reasonably and intuitively assumes that which examples are chosen has a far greater effect on…
Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises concerns among users about the LLM's stability and reliability. In this work, we consider…
With the advent of the Transformer architecture, Neural Machine Translation (NMT) results have shown great improvement lately. However, results in low-resource conditions still lag behind in both bilingual and multilingual setups, due to…