English
Related papers

Related papers: When does third order efficiency imply fourth orde…

200 papers

We consider a one-dimensional singularly perturbed 4th order problem with the additional feature of a shift term. An expansion into a smooth term, boundary layers and an inner layer yields a formal solution decomposition, and together with…

Numerical Analysis · Mathematics 2023-09-22 Sebastian Franz , Kleio Liotati

An important challenge in machine translation (MT) is to generate high-quality and diverse translations. Prior work has shown that the estimated likelihood from the MT model correlates poorly with translation quality. In contrast, quality…

Computation and Language · Computer Science 2024-10-17 Gonçalo R. A. Faria , Sweta Agrawal , António Farinhas , Ricardo Rei , José G. C. de Souza , André F. T. Martins

Large-Momentum Effective Theory (LaMET) is a physics-guided systematic expansion to calculate light-cone parton distributions, including collinear (PDFs) and transverse-momentum-dependent ones, at any fixed momentum fraction $x$ within a…

Distributed learning of probabilistic models from multiple data repositories with minimum communication is increasingly important. We study a simple communication-efficient learning framework that first calculates the local maximum…

Machine Learning · Statistics 2014-10-13 Qiang Liu , Alexander Ihler

We prove concentration inequalities for $f\left( X\right) $ about its median, where $X$ is a random vector in $\mathbb{R}^n$ with independent heavy tailed coordinates of Weibull or power type, and $f:\mathbb{R}^n\rightarrow\mathbb{R}$ is a…

Probability · Mathematics 2022-08-12 Daniel J. Fresen

In this paper, we re-examine the Markov property in the context of neural machine translation. We design a Markov Autoregressive Transformer~(MAT) and undertake a comprehensive assessment of its performance across four WMT benchmarks. Our…

Computation and Language · Computer Science 2024-02-06 Cunxiao Du , Hao Zhou , Zhaopeng Tu , Jing Jiang

The scaling trend in Large Language Models (LLMs) has prioritized increasing the maximum context window to facilitate complex, long-form reasoning and document analysis. However, managing this expanded context introduces severe…

Computation and Language · Computer Science 2026-01-21 Ahilan Ayyachamy Nadar Ponnusamy , Karthic Chandran , M Maruf Hossain

Parameter-efficient transfer learning (PETL) aims to adapt pre-trained models to new downstream tasks while minimizing the number of fine-tuned parameters. Adapters, a popular approach in PETL, inject additional capacity into existing…

Machine Learning · Computer Science 2024-10-22 Aleksandra I. Nowak , Otniel-Bogdan Mercea , Anurag Arnab , Jonas Pfeiffer , Yann Dauphin , Utku Evci

In fair division of indivisible goods, using sequences of sincere choices (or picking sequences) is a natural way to allocate the objects. The idea is the following: at each stage, a designated agent picks one object among those that…

Computer Science and Game Theory · Computer Science 2016-04-07 Sylvain Bouveret , Michel Lemaître

The distribution of the spacing, or the difference between consecutive order statistics, is known only for uniform and exponential random variates. We add here logistic and Gumbel variates, and present an estimator for distributions with a…

Methodology · Statistics 2026-01-30 Greg Kreider

This paper proves an impossibility result for stochastic network utility maximization for multi-user wireless systems, including multiple access and broadcast systems. Every time slot an access point observes the current channel states for…

Optimization and Control · Mathematics 2020-03-18 Michael J. Neely

In this paper the Gaussian quasi maximum likelihood estimator (GQMLE) is generalized by applying a transform to the probability distribution of the data. The proposed estimator, called measure-transformed GQMLE (MT-GQMLE), minimizes the…

Methodology · Statistics 2016-10-19 Koby Todros , Alfred O. Hero

We prove maximum and comparison principles for fractional discrete derivatives in the integers. Regularity results when the space is a mesh of length $h$, and approximation theorems to the continuous fractional derivatives are shown. When…

Analysis of PDEs · Mathematics 2016-05-24 Luciano Abadías , Marta de León-Contreras , José L. Torrea

Suppose we are given observations, where each observation is drawn independently from one of $k$ known distributions. The goal is to match each observation to the distribution from which it was drawn. We observe that the maximum likelihood…

Data Structures and Algorithms · Computer Science 2019-10-01 Sinho Chewi , Forest Yang , Avishek Ghosh , Abhay Parekh , Kannan Ramchandran

Throughout this article we develop and change the definitions and the ideas in "arXiv:1006.4939", in order to consider the efficiency of functions and complexity time problems. The central idea here is effective enumeration and listing, and…

Computational Complexity · Computer Science 2010-11-30 Saeed Asaeedi , Farzad Didehvar

Correlated outcomes are common in many practical problems. In some settings, one outcome is of particular interest, and others are auxiliary. To leverage information shared by all the outcomes, traditional multi-task learning (MTL)…

Methodology · Statistics 2023-03-23 Muxuan Liang , Jaeyoung Park , Qing Lu , Xiang Zhong

In supervised machine learning, the choice of loss function implicitly assumes a particular noise distribution over the data. For example, the frequently used mean squared error (MSE) loss assumes a Gaussian noise distribution. The choice…

Machine Learning · Computer Science 2023-02-15 Thamsanqa Mlotshwa , Heinrich van Deventer , Anna Sergeevna Bosman

Multimodal Large Language Models (MLLMs) utilize multimodal contexts consisting of text, images, or videos to solve various multimodal tasks. However, we find that changing the order of multimodal input can cause the model's performance to…

Artificial Intelligence · Computer Science 2024-10-23 Zhijie Tan , Xu Chu , Weiping Li , Tong Mo

In this paper, by treating in-context learning (ICL) as a meta-optimization process, we explain why LLMs are sensitive to the order of ICL examples. This understanding leads us to the development of Batch-ICL, an effective, efficient, and…

Machine Learning · Computer Science 2024-06-06 Kaiyi Zhang , Ang Lv , Yuhan Chen , Hansen Ha , Tao Xu , Rui Yan

An optimized Rayleigh-Schr\"{o}dinger expansion scheme of solving the functional Schr\"odinger equation with an external source is proposed to calculate the effective potential beyond the Gaussian approximation. For a scalar field theory…

High Energy Physics - Theory · Physics 2009-11-07 Wen-Fa Lu , Chul Koo Kim , Kyun Nahm
‹ Prev 1 3 4 5 6 7 10 Next ›