中文
相关论文

相关论文: Mpemba Effect in Large-Language Model Training Dyn…

200 篇论文

Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates…

计算与语言 · 计算机科学 2026-05-22 Meimingwei Li , Yuanhao Ding , Esteban Garces Arias , Christian Heumann

Training large models is both resource-intensive and time-consuming, making it crucial to understand the quantitative relationship between model performance and hyperparameters. In this paper, we present an empirical law that describes how…

机器学习 · 计算机科学 2025-03-18 Kairong Luo , Haodong Wen , Shengding Hu , Zhenbo Sun , Zhiyuan Liu , Maosong Sun , Kaifeng Lyu , Wenguang Chen

Recent analysis on the training dynamics of Transformers has unveiled an interesting characteristic: the training loss plateaus for a significant number of training steps, and then suddenly (and sharply) drops to near--optimal values. To…

机器学习 · 计算机科学 2024-10-30 Pulkit Gopalani , Ekdeep Singh Lubana , Wei Hu

Slow relaxation processes spanning widely separated timescales pose fundamental challenges for probing steady-state properties and engineering functional quantum systems, such as quantum heat engines and quantum computing devices. We…

量子物理 · 物理学 2025-10-20 Ruicheng Bao , Zhonghuai Hou

The so-called Mpemba effect, i.e. the observation that the warmer of two otherwise identical systems cools faster when both are refrigerated in the same thermal reservoir, is a hotly debated topic in condensed mater physics and statistical…

材料科学 · 物理学 2019-09-11 A. Gijón , A. Lasanta , E. R. Hernández

The ability of Large Language Models (LLMs) to extract context from natural language problem descriptions naturally raises questions about their suitability in autonomous decision-making settings. This paper studies the behaviour of these…

人工智能 · 计算机科学 2025-07-22 Xiao Yang , Juxi Leitner , Michael Burke

Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text's reasoning capability, undermining…

Under certain conditions, two samples of fluid at different initial temperatures present a counterintuitive behavior known as the Mpemba effect: it is the hotter system that cools sooner. Here, we show that the Mpemba effect is present in…

软凝聚态物质 · 物理学 2017-10-06 Antonio Lasanta , Francisco Vega Reyes , Antonio Prados , Andrés Santos

Schedule-Free Learning has shown promise as a practical anytime training method for machine learning, showing success across dozens of standard benchmark problems. However, strong performance for LLM training has only been demonstrated at…

机器学习 · 计算机科学 2026-05-20 Aaron Defazio

Beyond neural scaling laws, little is known about the laws underlying large language models (LLMs). We introduce Neural Thermodynamic Laws (NTL) -- a new framework that offers fresh insights into LLM training dynamics. On the theoretical…

机器学习 · 计算机科学 2025-05-16 Ziming Liu , Yizhou Liu , Jeff Gore , Max Tegmark

This study investigates the relationships which deep learning methods can identify between the input and output data. As a case study, rainfall-runoff modeling in a snow-dominated watershed by means of a long- and short-term memory (LSTM)…

大气与海洋物理 · 物理学 2021-11-11 Kazuki Yokoo , Kei Ishida , Ali Ercan , Tongbi Tu , Takeyoshi Nagasato , Masato Kiyama , Motoki Amagasaki

The Mpemba effect has initially been noticed in macroscopic systems -- namely that hot water can freeze faster than cold water -- but recently its extension to open quantum systems has attracted significant attention. This phenomenon can be…

介观与纳米尺度物理 · 物理学 2025-04-08 Juliane Graf , Janine Splettstoesser , Juliette Monsel

Increasing the number of parameters in language models is a common strategy to enhance their performance. However, smaller language models remain valuable due to their lower operational costs. Despite their advantages, smaller models…

计算与语言 · 计算机科学 2024-10-16 Richard Diehl Martinez , Pietro Lesci , Paula Buttery

Parameter-efficient fine-tuning (PEFT), particularly Low-Rank Adaptation (LoRA), adapts large language models (LLMs) by training only a small fraction of parameters. However, as the rank of the low-rank matrices used for adaptation…

计算与语言 · 计算机科学 2025-09-29 Yupeng Chang , Chenlu Guo , Yi Chang , Yuan Wu

Transferring the optimal learning rate from small to large neural networks can enable efficient training at scales where hyperparameter tuning is otherwise prohibitively expensive. To this end, the Maximal Update Parameterization (muP)…

机器学习 · 计算机科学 2026-02-16 Atli Kosson , Jeremy Welborn , Yang Liu , Martin Jaggi , Xi Chen

The Mpemba effect and its inverse can be understood as a result of nonequilibrium thermodynamics. In polymers, changes of state are generally non-equilibrium processes. However, the Mpemba effect has been rarely reported in the…

软凝聚态物质 · 物理学 2023-04-28 Jinghua Liu , Jingqing Li , Binyuan Liu , Ian W. Hamley , Shichun Jiang

Setting the learning rate (LR) for a deep learning model is a critical part of successful training. Choosing LRs is often done empirically with trial and error. In this work, we explore a solvable model of optimal LR schedules for a…

无序系统与神经网络 · 物理学 2026-05-11 Blake Bordelon , Francesco Mori

When a hot system cools down faster than an equivalent cold one, it exhibits the Mpemba Effect. This counterintuitive phenomenon was observed in several systems including water, magnetic alloys and polymers. In most experiments the system…

统计力学 · 物理学 2023-07-07 Gianluca Teza , Ran Yaacoby , Oren Raz

Training large language models (LLMs) typically involves pre-training on massive corpora, only to restart the process entirely when new data becomes available. A more efficient and resource-conserving approach would be continual…

Scaling laws have transformed our understanding of large language models by linking upstream metrics like cross-entropy loss to design factors such as model size, training data, and compute. However, these conventional laws fail to capture…

计算与语言 · 计算机科学 2025-10-17 Kyle Montgomery , David Park , Jianhong Tu , Michael Bendersky , Beliz Gunel , Dawn Song , Chenguang Wang