English
Related papers

Related papers: Two-Scale Latent Dynamics for Recurrent-Depth Tran…

200 papers

The objective of this work is to improve the accuracy of building demand forecasting. This is a more challenging task than grid level forecasting. For the said purpose, we develop a new technique called recurrent transform learning (RTL).…

Machine Learning · Computer Science 2019-12-12 Megha Gupta , Angshul Majumdar

It is known that input-output approaches based on scaled small-gain theorems with constant $D$-scalings and integral linear constraints are non-conservative for the analysis of some classes of linear positive systems interconnected with…

Optimization and Control · Mathematics 2017-03-02 Corentin Briat

Recurrent neural networks (RNNs) have been used extensively and with increasing success to model various types of sequential data. Much of this progress has been achieved through devising recurrent units and architectures with the…

Machine Learning · Statistics 2017-03-06 Yacine Jernite , Edouard Grave , Armand Joulin , Tomas Mikolov

We introduce a two-state non-conserving driven-diffusive system in one-dimension under a discrete-time updating scheme. We show that the steady-state of the system can be obtained using a matrix product approach. On the other hand, the…

Statistical Mechanics · Physics 2014-03-07 S. R. Masharian , F. H. Jafarpour , A. Aghamohammadi

Scaling Transformer-based click-through rate (CTR) models by stacking more parameters brings growing computational and storage overhead, creating a widening gap between scaling ambitions and the stringent industrial deployment constraints.…

Information Retrieval · Computer Science 2026-04-22 Jiakai Tang , Runfeng Zhang , Weiqiu Wang , Yifei Liu , Chuan Wang , Xu Chen , Yeqiu Yang , Jian Wu , Yuning Jiang , Bo Zheng

One-dimensional run-and-tumble processes may converge towards some localized non-equilibrium steady state when the two velocities and/or the two switching rates are space-dependent. A long dynamical trajectory can be then analyzed via the…

Statistical Mechanics · Physics 2021-08-23 Cecile Monthus

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction. However, practical dLLM decoding still suffers from high inference latency, which limits…

Computation and Language · Computer Science 2026-04-22 Zhenbang Du , Kejing Xia , Xinrui Zhong , Yonggan Fu , Nicolai Oswald , Binfei Ji , Brucek Khailany , Pavlo Molchanov , Yingyan Lin

Transformer-based models have dramatically increased their size and parameter count to tackle increasingly complex tasks. At the same time, there is a growing demand for high performance, low-latency inference on devices with limited…

Machine Learning · Computer Science 2026-04-01 Ginés Carreto Picón , Peng Yuan Zhou , Qi Zhang , Alexandros Iosifidis

A wide range of implicit time integration methods, including multi-step, implicit Runge-Kutta, and Galerkin finite-time element schemes, is evaluated in the context of chaotic dynamical systems. The schemes are applied to solve the Lorenz…

Computational Physics · Physics 2024-01-02 Viktoriya Morozova , James G. Coder , Kevin Holst

Transformer-based models generally allocate the same amount of computation for each token in a given sequence. We develop a simple but effective "token dropping" method to accelerate the pretraining of transformer models, such as BERT,…

Computation and Language · Computer Science 2022-03-25 Le Hou , Richard Yuanzhe Pang , Tianyi Zhou , Yuexin Wu , Xinying Song , Xiaodan Song , Denny Zhou

Quantum coherence reflects the origin of quantumness and might be capable of extracting the subtle nature of a system. We investigate the ground-state coherence and steered coherence in the Lipkin-Meshkov-Glick model and show that they…

Quantum Physics · Physics 2021-12-14 Ming-Liang Hu , Fan Fang , Heng Fan

Transformer models have emerged as fundamental tools across various scientific and engineering disciplines, owing to their outstanding performance in diverse applications. Despite this empirical success, the theoretical foundations of…

Machine Learning · Computer Science 2026-04-14 Zhen Qin , Jinxin Zhou , Jiachen Jiang , Zhihui Zhu

Large reasoning models achieve strong performance on complex tasks by generating extended chains of thought, but they often "overthink": continuing to reason long after they have enough information to answer correctly. This wastes…

Computation and Language · Computer Science 2025-12-08 Ömer Faruk Akgül , Yusuf Hakan Kalaycı , Rajgopal Kannan , Willie Neiswanger , Viktor Prasanna

Sequence modelling requires determining which past tokens are causally relevant from the context and their importance: a process inherent to the attention layers in transformers, yet whose underlying learned mechanisms remain poorly…

Machine Learning · Computer Science 2026-04-14 Francesco D'Angelo , Nicolas Flammarion

This paper presents a simple, effective, and cost-efficient strategy to improve LLM performance by scaling test-time compute. Our strategy builds upon the repeated-sampling-then-voting framework, with a novel twist: incorporating multiple…

Artificial Intelligence · Computer Science 2025-11-11 Jianhao Chen , Zishuo Xun , Bocheng Zhou , Han Qi , Hangfan Zhang , Qiaosheng Zhang , Yang Chen , Wei Hu , Yuzhong Qu , Wanli Ouyang , Shuyue Hu

We present a finite-size scaling analysis of the droplet condensation-evaporation transition of a lattice gas (in two and three dimensions) and a Lennard-Jones gas (in three dimensions) at fixed density. Parallel multicanonical simulations…

Soft Condensed Matter · Physics 2015-08-04 Johannes Zierenberg , Wolfhard Janke

In this paper, we propose to model the video dynamics by learning the trajectory of independently inverted latent codes from GANs. The entire sequence is seen as discrete-time observations of a continuous trajectory of the initial latent…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Weihao Xia , Yujiu Yang , Jing-Hao Xue

Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each…

Machine Learning · Computer Science 2026-05-15 Jesseba Fernando , Grigori Guitchounts

Long Short-Term Memory (LSTM) is a special class of recurrent neural network, which has shown remarkable successes in processing sequential data. The typical architecture of an LSTM involves a set of states and gates: the states retain…

Machine Learning · Computer Science 2018-12-03 Arash Ardakani , Zhengyun Ji , Warren J. Gross

In two-dimensional decaying homogeneous isotropic turbulence, kinetic energy and enstrophy are respectively transferred to larger and smaller scales. In such spatiotemporally complex dynamics, it is challenging to identify the important…

Fluid Dynamics · Physics 2023-12-08 Aditya G. Nair , James Hanna , Matteo Aureli