Related papers: When does third order efficiency imply fourth orde…
Zeroth-order optimization is a fundamental research topic that has been a focus of various learning tasks, such as black-box adversarial attacks, bandits, and reinforcement learning. However, in theory, most complexity results assert a…
This paper describes the third place submission to the shared task on simultaneous translation and paraphrasing for language education at the 4th workshop on Neural Generation and Translation (WNGT) for ACL 2020. The final system leverages…
We consider ordered logit models for directed network data that allow for flexible sender and receiver fixed effects that can vary arbitrarily across outcome categories. This structure poses a significant incidental parameter problem,…
We consider the estimation of two-sample integral functionals, of the type that occur naturally, for example, when the object of interest is a divergence between unknown probability densities. Our first main result is that, in wide…
We consider the torsion function for the Dirichlet Laplacian $-\Delta$, and for the Schr\"odinger operator $- \Delta + V$ on an open set $\Omega\subset \R^m$ of finite Lebesgue measure $0<|\Omega|<\infty$ with a real-valued, non-negative,…
In this article, we study the performance of the estimator that minimizes $L_{2k}- $ order loss function (for $ k \ge \; 2 )$ against the estimators which minimizes the $L_2-$ order loss function (or the least squares estimator). Commonly…
Large language model (LLM) shows promising performances in a variety of downstream tasks, such as machine translation (MT). However, using LLMs for translation suffers from high computational costs and significant latency. Based on our…
Zeroth-order optimization is the process of minimizing an objective $f(x)$, given oracle access to evaluations at adaptively chosen inputs $x$. In this paper, we present two simple yet powerful GradientLess Descent (GLD) algorithms that do…
Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the output distributions of multiple models and selecting the most probable label. This idea…
We provide a theoretical treatment of over-specified Gaussian mixtures of experts with covariate-free gating networks. We establish the convergence rates of the maximum likelihood estimation (MLE) for these models. Our proof technique is…
This study examines the generalization performance and interpretability of machine learning (ML) models used for predicting crop yield and yield anomalies in Germany's NUTS-3 regions. Using a high-quality, long-term dataset, the study…
The master equation for a charged harmonic oscillator coupled to an electromagnetic reservoir is investigated up to fourth-order in the interaction strength by using Krylov averaging method. The interaction is in the velocity-coupling form…
The independent component model is a latent variable model where the components of the observed random vector are linear combinations of latent independent variables. The aim is to find an estimate for a transformation matrix back to…
In this paper we explore evaluation of LLM capabilities. We present measurements of GPT-4 performance on several deterministic tasks; each task involves a basic calculation and takes as input parameter some element drawn from a large…
The global attention mechanism is one of the keys to the success of transformer architecture, but it incurs quadratic computational costs in relation to the number of tokens. On the other hand, equivariant models, which leverage the…
This work reports a quantitative analysis to predicting the efficiency of distributed computing running in three models of complex networks: Barab\'asi-Albert, Erd\H{o}s-R\'enyi and Watts-Strogatz. A master/slave computing model is…
We describe here the properties expected of a higher (with emphasis on the order fourth) order phase transition. The order is identified in the sense first noted by Ehrenfest, namely in terms of the temperature dependence of the ordered…
The log-concave maximum likelihood estimator (MLE) problem answers: for a set of points $X_1,...X_n \in \mathbb R^d$, which log-concave density maximizes their likelihood? We present a characterization of the log-concave MLE that leads to…
We study the Hamiltonian truncation for the two-dimensional $\lambda\phi^4$ theory within the framework of Hamiltonian truncation effective theory, where truncation artifacts are mitigated through a systematic inclusion of corrective terms…
In this paper, we study the log-likelihood function and Maximum Likelihood Estimate (MLE) for the matrix normal model for both real and complex models. We describe the exact number of samples needed to achieve (almost surely) three…