English
Related papers

Related papers: A completely uniform transformer for parity

200 papers

Effectively processing long contexts is a critical challenge for language models. While standard Transformers are limited by quadratic complexity and poor length extrapolation, alternative architectures like sliding window attention and…

Computation and Language · Computer Science 2026-05-01 Jiaqi Leng , Xiang Hu , Junxiong Wang , Jianguo Li , Wei Wu , Yucheng Lu

Many NLP models operate over sequences of subword tokens produced by hand-crafted tokenization rules and heuristic subword induction algorithms. A simple universal alternative is to represent every computerized text as a sequence of bytes…

Computation and Language · Computer Science 2021-04-13 Uri Shaham , Omer Levy

We propose a method for the construction of sets of variable dimension strong non-overlapping matrices basing on any strong non-overlapping set of strings.

Combinatorics · Mathematics 2023-09-07 Elena Barcucci , Antonio Bernini , Stefano Bilotta , Renzo Pinzani

The advent of Transformer-based models has surpassed the barriers of text. When working with speech, we must face a problem: the sequence length of an audio input is not suitable for the Transformer. To bypass this problem, a usual approach…

Computation and Language · Computer Science 2021-07-08 Belen Alastruey , Gerard I. Gállego , Marta R. Costa-jussà

In odd dimensions the lattice overlap formalism is simpler than in even dimensions. Masslessness of fermions can still be preserved without fine tuning and gauge invariance without gauge averaging can be maintained, although, sometimes,…

High Energy Physics - Lattice · Physics 2009-10-30 Y. Kikukawa , H. Neuberger

The study on the expressive power of transformers shows that transformers are permutation equivariant, and they can approximate all permutation-equivariant continuous functions on a compact domain. However, these results are derived under…

Machine Learning · Computer Science 2026-01-26 Sejun Park , Yeachan Park , Geonho Hwang

Chain of thought is a natural inference-time method for increasing the computational power of transformer-based large language models (LLMs), but comes at the cost of sequential decoding. Are there more efficient alternatives to expand a…

Machine Learning · Computer Science 2025-11-07 William Merrill , Ashish Sabharwal

We consider spatially coupled low-density parity-check codes with finite smoothing parameters. A finite smoothing parameter is important for designing practical codes that are decoded using low-complexity windowed decoders. By optimizing…

Information Theory · Computer Science 2019-07-09 Laurent Schmalen , Vahid Aref

The tongue's intricate 3D structure, comprising localized functional units, plays a crucial role in the production of speech. When measured using tagged MRI, these functional units exhibit cohesive displacements and derived quantities that…

Transformer becomes the state-of-the-art translation model, while it is not well studied how each intermediate component contributes to the model performance, which poses significant challenges for designing optimal architectures. In this…

Computation and Language · Computer Science 2020-11-10 Wenxuan Wang , Zhaopeng Tu

A novel method guaranteeing nondecreasing girth is presented for constructing longer low-density parity-check (LDPC) codes from shorter ones. The parity-check matrix of a shorter base code is decomposed into N (N>=2) non-overlapping…

Information Theory · Computer Science 2018-01-29 Guohua Zhang , Yulin Hu , Qinwei He

We show that the Parikh image of the language of an NFA with n states over an alphabet of size k can be described as a finite union of linear sets with at most k generators and total size 2^{O(k^2 log n)}, i.e., polynomial for all fixed k…

Logic in Computer Science · Computer Science 2010-02-12 Anthony Widjaja To

Let $K$ denote a field and let $V$ denote a vector space over $K$ with finite positive dimension. We consider a pair of linear transformations $A : V \to V$ and $A^* : V \to V$ that satisfy (i) and (ii) below: (i) There exists a basis for…

Rings and Algebras · Mathematics 2007-05-23 Kazumasa Nomura , Paul Terwilliger

We determine the generic consistency, dimension and nondegeneracy of the zero locus over $\mathbb{C}^*$, $\mathbb{R}^*$ and $\mathbb{R}_{>0}$ of vertically parametrized systems: parametric polynomial systems consisting of linear…

Algebraic Geometry · Mathematics 2025-10-06 Elisenda Feliu , Oskar Henriksson , Beatriz Pascual-Escudero

We re-evaluate the standard practice of sharing weights between input and output embeddings in state-of-the-art pre-trained language models. We show that decoupled embeddings provide increased modeling flexibility, allowing us to…

Computation and Language · Computer Science 2020-10-27 Hyung Won Chung , Thibault Févry , Henry Tsai , Melvin Johnson , Sebastian Ruder

Deep learning models have been widely applied in various aspects of daily life. Many variant models based on deep learning structures have achieved even better performances. Attention-based architectures have become almost ubiquitous in…

Machine Learning · Computer Science 2022-02-25 Zhiying Fang , Yidong Ouyang , Ding-Xuan Zhou , Guang Cheng

We present Layout Anything, a transformer-based framework for indoor layout estimation that adapts the OneFormer's universal segmentation architecture to geometric structure prediction. Our approach integrates OneFormer's task-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Md Sohag Mia , Muhammad Abdullah Adnan

We present numerical evidences using overlap fermions for a scale-invariant behavior of parity-invariant three-dimensional QED with two flavors of massless two-component fermions. Using finite-size scaling of the low-lying eigenvalues of…

High Energy Physics - Theory · Physics 2016-09-22 Nikhil Karthik , Rajamani Narayanan

Some Transformer-based models can perform cross-lingual transfer learning: those models can be trained on a specific task in one language and give relatively good results on the same task in another language, despite having been pre-trained…

Computation and Language · Computer Science 2022-07-20 Félix Gaschi , François Plesse , Parisa Rastin , Yannick Toussaint

Recent developments have shown the existence of quantum low-density parity check (qLDPC) codes with constant rate and linear distance. A natural question concerns the efficient decodability of these codes. In this paper, we present a linear…

Quantum Physics · Physics 2022-06-15 Shouzhen Gu , Christopher A. Pattison , Eugene Tang
‹ Prev 1 4 5 6 7 8 10 Next ›