English
Related papers

Related papers: VORT: Adaptive Power-Law Memory for NLP Transforme…

200 papers

Inference in both brains and machines can be formalized by optimizing a shared objective: maximizing the evidence lower bound (ELBO) in machine learning, or minimizing variational free energy (F) in neuroscience (ELBO = -F). While this…

Artificial Intelligence · Computer Science 2025-10-27 Hadi Vafaii , Dekel Galor , Jacob L. Yates

We calculate next-to-leading-order (NLO) corrections to exclusive processes in $k_T$ factorization theorem, taking $\pi\gamma^*\to\gamma$ as an example. Partons off-shell by $k_T^2$ are considered in both the quark diagrams from full QCD…

High Energy Physics - Phenomenology · Physics 2008-11-26 Soumitra Nandi , Hsiang-nan Li

Standard transformer attention computes pairwise token similarity but treats all tokens as equally salient and all positions as equally local, regardless of the informational structure of the input. We identify two complementary inductive…

Machine Learning · Computer Science 2026-05-27 Athanasios Zeris

We study reinforcement learning for partially observed Markov decision processes (POMDPs) with infinite observation and state spaces, which remains less investigated theoretically. To this end, we make the first attempt at bridging partial…

Machine Learning · Computer Science 2024-04-02 Qi Cai , Zhuoran Yang , Zhaoran Wang

Vision Transformers have been tremendously successful in computer vision tasks. However, their large computational, memory, and energy demands are a challenge for edge inference on FPGAs -- a field that has seen a recent surge in demand. We…

Transformer classifiers such as BERT deliver impressive closed-set accuracy, yet they remain brittle when confronted with inputs from unseen categories--a common scenario for deployed NLP systems. We investigate Open-Set Recognition (OSR)…

Machine Learning · Computer Science 2026-01-06 Tianshuo Yang , Ryan Rabinowitz , Terrance E. Boult , Jugal Kalita

Owing to the impressive dot-product attention, the Transformers have been the dominant architectures in various natural language processing (NLP) tasks. Recently, the Receptance Weighted Key Value (RWKV) architecture follows a…

Computation and Language · Computer Science 2024-09-16 Leilei Wang

Vision transformer has emerged as a new paradigm in computer vision, showing excellent performance while accompanied by expensive computational cost. Image token pruning is one of the main approaches for ViT compression, due to the facts…

Computer Vision and Pattern Recognition · Computer Science 2023-07-07 Xiangcheng Liu , Tianyi Wu , Guodong Guo

Long and short memory in economic processes is usually described by the so-called discrete fractional differencing and fractional integration. We prove that the discrete fractional differencing and integration are the Grunwald-Letnikov…

Economics · Quantitative Finance 2017-08-08 Vasily E. Tarasov , Valentina V. Tarasova

Earth introduces strong attenuation and dispersion to propagating waves. The time-fractional wave equation with very small fractional exponent, based on Kjartansson's constant-Q theory, is widely recognized in the field of geophysics as a…

Numerical Analysis · Mathematics 2023-09-12 Xu Guo , Shidong Jiang , Yunfeng Xiong , Jiwei Zhang

Attention based language models have become a critical component in state-of-the-art natural language processing systems. However, these models have significant computational requirements, due to long training times, dense operations and…

Computation and Language · Computer Science 2021-06-11 Ivan Chelombiev , Daniel Justus , Douglas Orr , Anastasia Dietrich , Frithjof Gressmann , Alexandros Koliousis , Carlo Luschi

Transformer is a new kind of neural architecture which encodes the input data as powerful features via the attention mechanism. Basically, the visual transformers first divide the input images into several local patches and then calculate…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Kai Han , An Xiao , Enhua Wu , Jianyuan Guo , Chunjing Xu , Yunhe Wang

We conduct a systematic study of the approximation properties of Transformer for sequence modeling with long, sparse and complicated memory. We investigate the mechanisms through which different components of Transformer, such as the…

Machine Learning · Computer Science 2024-10-31 Mingze Wang , Weinan E

This article is the second work in our series of papers dedicated to image processing models based on the fractional order total variation $TV^r$. In our first work of this series, we studied key analytic properties of these semi-norms.…

Optimization and Control · Mathematics 2019-03-21 Pan Liu , Xin Yang Lu

The transformer architecture has revolutionized Natural Language Processing (NLP) and other machine-learning tasks, due to its unprecedented accuracy. However, their extensive memory and parameter requirements often hinder their practical…

Computation and Language · Computer Science 2023-11-01 Subhadra Vadlamannati , Ryan Solgi

Transformers trained on modular arithmetic exhibit sharp transitions between memorization, generalization, and collapse. We show that weight decay acts as a scalar empirical control parameter for these regimes, and introduce two cheap…

Machine Learning · Computer Science 2026-05-21 Lucky Verma

Existing techniques to encode spatial invariance within deep convolutional neural networks (CNNs) apply the same warping field to all the feature channels. This does not account for the fact that the individual feature channels can…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Seungryong Kim , Sabine Süsstrunk , Mathieu Salzmann

Text recognition in natural images remains a challenging yet essential task, with broad applications spanning computer vision and natural language processing. This paper introduces a novel end-to-end framework that combines ResNet and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Naphat Nithisopa , Teerapong Panboonyuen

Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a…

Artificial Intelligence · Computer Science 2026-05-14 Duc Anh Le , Tien-Phat Nguyen , Thien Huu Nguyen , Linh Ngo Van , Trung Le

We discuss the problem of adaptive discrete-time signal denoising in the situation where the signal to be recovered admits a "linear oracle" -- an unknown linear estimate that takes the form of convolution of observations with a…

Statistics Theory · Mathematics 2021-02-15 Zaid Harchaoui , Anatoli Juditsky , Arkadi Nemirovski , Dmitrii Ostrovskii
‹ Prev 1 4 5 6 7 8 10 Next ›