中文
相关论文

相关论文: Two-Scale Latent Dynamics for Recurrent-Depth Tran…

200 篇论文

Recent advances in natural language processing highlight two key factors for improving reasoning in large language models (LLMs): (i) allocating more test-time compute tends to help on harder problems but often introduces redundancy in the…

计算与语言 · 计算机科学 2025-11-04 Riccardo Alberghi , Elizaveta Demyanenko , Luca Biggio , Luca Saglietti

In this paper, we propose a method of designing low-dimensional retrofit controllers for interconnected linear systems. In the proposed method, by retrofitting an additional low-dimensional controller to a preexisting control system, we aim…

系统与控制 · 计算机科学 2017-06-09 Takayuki Ishizaki , Masakazu Koike , Jun-ichi Imura

Large Transformer models routinely achieve state-of-the-art results on a number of tasks but training these models can be prohibitively costly, especially on long sequences. We introduce two techniques to improve the efficiency of…

机器学习 · 计算机科学 2020-02-19 Nikita Kitaev , Łukasz Kaiser , Anselm Levskaya

Test-time scaling via explicit reasoning trajectories significantly boosts large language model (LLM) performance but often triggers overthinking. To explore this, we analyze reasoning through two lenses: Reasoning Length Dynamics, which…

计算与语言 · 计算机科学 2026-01-14 Zihao Wei , Liang Pang , Jiahao Liu , Wenjie Shi , Jingcheng Deng , Shicheng Xu , Zenghao Duan , Fei Sun , Huawei Shen , Xueqi Cheng

We introduce a new class of time-continuous recurrent neural network models. Instead of declaring a learning system's dynamics by implicit nonlinearities, we construct networks of linear first-order dynamical systems modulated via nonlinear…

机器学习 · 计算机科学 2020-12-16 Ramin Hasani , Mathias Lechner , Alexander Amini , Daniela Rus , Radu Grosu

We introduce a data-driven approach to building reduced dynamical models through manifold learning; the reduced latent space is discovered using Diffusion Maps (a manifold learning technique) on time series data. A second round of Diffusion…

Looped Transformers provide advantages in parameter efficiency, computational capabilities, and generalization for reasoning tasks. However, their expressive power regarding function approximation remains underexplored. In this paper, we…

机器学习 · 计算机科学 2025-06-06 Kevin Xu , Issei Sato

Recent advances in large language models (LLMs), such as OpenAI-o1 and DeepSeek-R1, have demonstrated the effectiveness of test-time scaling, where extended reasoning processes substantially enhance model performance. Despite this, current…

计算与语言 · 计算机科学 2025-03-26 Xiaoyu Tian , Sitong Zhao , Haotian Wang , Shuaiting Chen , Yunjie Ji , Yiping Peng , Han Zhao , Xiangang Li

Current approaches for scaling inference-time compute in transformers train them to emit explicit chain-of-thought tokens before producing an answer. While these methods are powerful, they are limited because they cannot be applied during…

机器学习 · 计算机科学 2026-02-02 Houjun Liu , Shikhar Murty , Christopher D. Manning , Róbert Csordás

Encoder-decoder models offer substantial inference-time savings over decoder-only models, but their pretraining objectives suffer from sparse supervision and dynamic sequence lengths, keeping them out of practice at scale. We propose…

机器学习 · 计算机科学 2026-05-20 Asher Labovich , Benjamin Bradley , Vanessa Alexander , Chaitanya Harsha

We tackle the question of how to scale more efficiently across the many, ever-growing stages of current LLM training pipelines. Our guiding intuition stems from the fact that the dynamics of later stages of the pipeline, e.g. post-training,…

We investigate the transient times for the onset of control of steady states by time-delayed feedback. The optimization of control by minimising the transient time before control becomes effective is discussed analytically and numerically,…

适应与自组织系统 · 物理学 2009-12-10 Robert C. Hinz , Philipp Hövel , Eckehard Schöll

A stochastic process, when subject to resetting to its initial condition at a constant rate, generically reaches a non-equilibrium steady state. We study analytically how the steady state is approached in time and find an unusual relaxation…

统计力学 · 物理学 2015-05-29 Satya N. Majumdar , Sanjib Sabhapandit , Gregory Schehr

A dynamical model is proposed for isotropic turbulence driven by steady forcing that yields a viscosity independent dynamics for the small-scale (inertial) regime. This reproduces the Kolmogorov spectrum for the two-point velocity…

统计力学 · 物理学 2011-07-04 Mohammad Mehrafarin

Noise-induced escape from a metastable state of a dynamical system is studied close to a saddle-node bifurcation point, but in the region where the system remains underdamped. The activation energy of escape scales as a power of the…

介观与纳米尺度物理 · 物理学 2009-11-11 M. I. Dykman , I. B. Schwartz , M. Shapiro

Temporal difference (TD) learning is a foundational algorithm in reinforcement learning (RL). For nearly forty years, TD learning has served as a workhorse for applied RL as well as a building block for more complex and specialized…

机器学习 · 计算机科学 2025-06-24 Hwanwoo Kim , Panos Toulis , Eric Laber

We observe a novel 'multiple-descent' phenomenon during the training process of LSTM, in which the test loss goes through long cycles of up and down trend multiple times after the model is overtrained. By carrying out asymptotic stability…

机器学习 · 计算机科学 2025-05-27 Wenbo Wei , Nicholas Chong Jia Le , Choy Heng Lai , Ling Feng

Transformer architectures, particularly Diffusion Transformers (DiTs), have become widely used in diffusion and flow-matching models due to their strong performance compared to convolutional UNets. However, the isotropic design of DiTs…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Quan Dao , Dimitris Metaxas

To model time-varying nonlinear temporal dynamics in sequential data, a recurrent network capable of varying and adjusting the recurrence depth between input intervals is examined. The recurrence depth is extended by several intermediate…

机器学习 · 计算机科学 2017-08-15 Hyunsin Park , Chang D. Yoo

Autoregressive language models trained with next-token prediction generate text by sampling one discrete token at a time. Although very scalable, this objective forces the model to commit at every step, preventing it from exploring or…

计算与语言 · 计算机科学 2026-03-24 Lorenzo Noci , Gregor Bachmann , Seyed-Mohsen Moosavi-Dezfooli , Moin Nabi