English
Related papers

Related papers: Manifold Trajectories in Next-Token Prediction: Fr…

200 papers

Speculative decoding has emerged as a promising approach to accelerate autoregressive inference in large language models (LLMs). Self-draft methods, which leverage the base LLM itself for speculation, avoid the overhead of auxiliary draft…

Computation and Language · Computer Science 2026-04-15 Zhuofan Wen , Yang Feng

The finite symmetric group S_n provides a natural domain for permutations, yet learning probability distributions on S_n is challenging due to its factorially growing size and discrete, non-Euclidean structure. Recent permutation diffusion…

Machine Learning · Computer Science 2026-03-19 Sizhuang He , Yangtian Zhang , Shiyang Zhang , David van Dijk

This work aims to develop a measure that can accurately rank the performance of various classifiers when they are tested on unlabeled data from out-of-distribution (OOD) distributions. We commence by demonstrating that conventional…

Machine Learning · Computer Science 2024-06-17 Weijie Tu , Weijian Deng , Liang Zheng , Tom Gedeon

In high dimensions, reflective Hamiltonian Monte Carlo with inexact reflections exhibits slow mixing when the particle ensemble is initialised from a Dirac delta distribution and the uniform distribution is targeted. By quantifying the…

Machine Learning · Statistics 2026-03-20 Namu Kroupa , Gábor Csányi , Will Handley

Composite minimization involves a collection of smooth functions which are aggregated in a nonsmooth manner. In the convex setting, we design an algorithm by linearizing each smooth component in accordance with its main curvature. The…

Optimization and Control · Mathematics 2019-03-26 Jérôme Bolte , Zheng Chen , Edouard Pauwels

Human movement prediction is difficult as humans naturally exhibit complex behaviors that can change drastically from one environment to the next. In order to alleviate this issue, we propose a prediction framework that decouples short-term…

Robotics · Computer Science 2020-03-19 Philipp Kratzer , Marc Toussaint , Jim Mainprice

Certified verification of transformer attention requires bounding the softmax function over interval constraints on the pre-softmax scores. Existing verifiers relax softmax ndependently of the downstream objective, leaving avoidable slack.…

Machine Learning · Computer Science 2026-05-13 Navid Rezazadeh , Arash Gholami Davoodi

Recursive learning -- where models are trained on data generated by previous versions of themselves -- is increasingly common in large language models, autonomous agents, and self-supervised systems. However, standard performance metrics…

Machine Learning · Computer Science 2026-05-20 Zhipeng Zhang

We propose a novel computational strategy to study the glass transition of molecular fluids. Our approach combines the construction of simple yet realistic models with the development of Monte Carlo algorithms to accelerate equilibration…

Statistical Mechanics · Physics 2026-03-31 Romain Simon , Jean-Louis Barrat , Ludovic Berthier

Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at…

Analysis of PDEs · Mathematics 2026-05-12 Albert Alcalde , Leon Bungert , Konstantin Riedl , Tim Roith

Token-based transformer world models have shown strong performance in visual reinforcement learning, but often suffer from temporal inconsistency in long-horizon rollouts, including object duplication, disappearance, and transmutation. A…

Machine Learning · Computer Science 2026-05-27 Youngin Kim , Ray Sun , Inho Kim , Bumsoo Park , Hyun Oh Song

Diffusion models often generate novel samples even when the learned score is only \emph{coarse} -- a phenomenon not accounted for by the standard view of diffusion training as density estimation. In this paper, we show that, under the…

Machine Learning · Computer Science 2026-03-26 Zebang Shen , Ya-Ping Hsieh , Niao He

This work presents self-rewarding sequential Monte Carlo (SMC), an inference-time scaling algorithm enabling effective sampling of masked diffusion language models (MDLMs). Our algorithm stems from the observation that most existing MDLMs…

Machine Learning · Computer Science 2026-02-03 Ziwei Luo , Ziqi Jin , Lei Wang , Lidong Bing , Thomas B. Schön

Nearly every recent image synthesis approach, including diffusion, masked-token prediction, and next-token prediction, uses a Transformer network architecture. Despite this common backbone, there has been no direct, compute controlled…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Maciej Kilian , Varun Jampani , Luke Zettlemoyer

Effective movement primitives should be capable of encoding and generating a rich repertoire of trajectories -- typically collected from human demonstrations -- conditioned on task-defining parameters such as vision or language inputs.…

Robotics · Computer Science 2025-01-09 Yonghyeon Lee , Byeongho Lee , Seungyeon Kim , Frank C. Park

The paper studies sub and super-replication price bounds for contingent claims defined on general trajectory based market models. No prior probabilistic or topological assumptions are placed on the trajectory space, trading is assumed to…

Mathematical Finance · Quantitative Finance 2018-02-22 Ivan Degano , Sebastian Ferrando , Alfredo Gonzalez

Riemannian diffusion models draw inspiration from standard Euclidean space diffusion models to learn distributions on general manifolds. Unfortunately, the additional geometric complexity renders the diffusion transition term inexpressible…

Machine Learning · Computer Science 2023-11-01 Aaron Lou , Minkai Xu , Stefano Ermon

This paper concerns the mathematical analyses of the diffusion model in machine learning. The drift term of the backward sampling process is represented as a conditional expectation involving the data distribution and the forward diffusion.…

Machine Learning · Computer Science 2024-12-11 Yubin Lu , Zhongjian Wang , Guillaume Bal

In this paper, we present a maximum likelihood estimation approach to determine the value vector in transformer models. We model the sequence of value vectors, key vectors, and the query vector as a sequence of Gaussian distributions. The…

Machine Learning · Computer Science 2025-09-17 Jiyong Ma

While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks. Multi-token prediction (MTP) has recently emerged as a promising alternative,…

Machine Learning · Computer Science 2026-04-15 Jianhao Huang , Zhanpeng Zhou , Renqiu Xia , Baharan Mirzasoleiman , Weijie Su , Wei Huang
‹ Prev 1 3 4 5 6 7 10 Next ›