中文
相关论文

相关论文: Mpemba Effect in Large-Language Model Training Dyn…

200 篇论文

During long-duration Large Language Model (LLM) training runs the gradient norm increases rapidly near the end of training. In this short note, we show that this increase is due to an unintended interaction between weight decay,…

机器学习 · 计算机科学 2025-06-11 Aaron Defazio

Despite decades of research, the Mpemba Effect challenges scientists, prompting further investigation and refinement of existing hypotheses. This work uses optical tools such as thermography to analyze and study the Mpemba effect on drops.…

流体动力学 · 物理学 2024-09-05 Argelia Balbuena Ortega , Emiliano Hernández-Figueroa , J. Antonio del Río

Large language models (LLMs) are typically optimized for resource-rich languages like English, exacerbating the gap between high-resource and underrepresented languages. This work presents a detailed analysis of strategies for developing a…

计算与语言 · 计算机科学 2024-12-19 Ander Corral , Ixak Sarasua , Xabier Saralegi

Recent deep learning approaches for river discharge forecasting have improved the accuracy and efficiency in flood forecasting, enabling more reliable early warning systems for risk management. Nevertheless, existing deep learning…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Mohamad Hakam Shams Eddin , Yikui Zhang , Stefan Kollet , Juergen Gall

The multi-stage phenomenon in the training loss curves of neural networks has been widely observed, reflecting the non-linearity and complexity inherent in the training process. In this work, we investigate the training dynamics of neural…

机器学习 · 计算机科学 2024-11-07 Zheng-An Chen , Tao Luo , GuiHong Wang

The Mpemba effect occurs when a hot system cools faster than an initially colder one, when both are refrigerated in the same thermal reservoir. Using the custom built supercomputer Janus II, we study the Mpemba effect in spin glasses and…

Recent advances in NLP have significantly improved the performance of language models on a variety of tasks. While these advances are largely driven by the availability of large amounts of data and computational power, they also benefit…

计算与语言 · 计算机科学 2023-06-05 Wissam Antoun , Benoît Sagot , Djamé Seddah

Recent advancements in large language models (LLMs) reveal a perplexing phenomenon in continual learning: despite extensive training, models experience significant performance declines, raising questions about task alignment and underlying…

机器学习 · 计算机科学 2025-01-24 Junhao Zheng , Xidi Cai , Shengjie Qiu , Qianli Ma

This paper investigates the one-epoch overfitting phenomenon in Click-Through Rate (CTR) models, where performance notably declines at the start of the second epoch. Despite extensive research, the efficacy of multi-epoch training over the…

机器学习 · 计算机科学 2024-07-03 Zhongxiang Fan , Zhaocheng Liu , Jian Liang , Dongying Kong , Han Li , Peng Jiang , Shuang Li , Kun Gai

We analyze the convergence behavior of stochastic gradient descent with momentum (SGDM) under dynamic learning-rate and batch-size schedules by introducing a novel and simpler Lyapunov function. We extend the existing theoretical framework…

机器学习 · 计算机科学 2025-10-13 Yuichi Kondo , Hideaki Iiduka

Self-Rewarding Language Models propose an architecture in which the Large Language Models(LLMs) both generates responses and evaluates its own outputs via LLM-as-a-Judge prompting, dynamically improving its generative capabilities through…

We find analytically the complete set of eigenvalues and eigenvectors associated with Metropolis dynamics on a complete graph. As an application, we use this information to study a counter-intuitive relaxation phenomenon, called the Mpemba…

数学物理 · 物理学 2019-02-07 Israel Klich , Marija Vucelja

Detecting pre-training data in Large Language Models (LLMs) is crucial for auditing data privacy and copyright compliance, yet it remains challenging in black-box, zero-shot settings where computational resources and training data are…

计算与语言 · 计算机科学 2026-01-13 Jinhan Liu , Yibo Yang , Ruiying Lu , Piotr Piekos , Yimeng Chen , Peng Wang , Dandan Guo

The prevailing paradigm in large language model (LLM) development is to pretrain a base model, then perform further training to improve performance and model behavior. However, hyperparameter optimization and scaling laws have been studied…

机器学习 · 计算机科学 2026-02-12 Tessa Han , Sebastian Bordt , Hanlin Zhang , Sham Kakade

Diffusion models (DMs) are a powerful generative framework that have attracted significant attention in recent years. However, the high computational cost of training DMs limits their practical applications. In this paper, we start with a…

机器学习 · 计算机科学 2024-04-12 Tianshuo Xu , Peng Mi , Ruilin Wang , Yingcong Chen

Quantum thermometry provides a key capability for nanoscale devices and quantum technologies, but most existing strategies rely on probes initialized near equilibrium. This equilibrium paradigm imposes intrinsic limitations: sensitivity is…

量子物理 · 物理学 2026-01-09 Pritam Chattopadhyay , Jonas F. G. Santos , Avijit Misra

We study the local relaxation of closed quantum systems through the relative entropy between the reduced density matrix and its long time limit. We show, using analytic arguments combined with numerical checks, that this relative entropy…

统计力学 · 物理学 2025-10-29 Filiberto Ares , Colin Rylands , Pasquale Calabrese

Fine-tuning large language models (LLMs) for reasoning tasks using reinforcement learning methods like Group Relative Policy Optimization (GRPO) is computationally expensive. To address this, we propose a predictive framework that models…

机器学习 · 计算机科学 2026-03-23 Datta Nimmaturi , Vaishnavi Bhargava , Rajat Ghosh , Johnu George , Debojyoti Dutta

Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forces them to absorb task-specific information, which can result in catastrophic forgetting and loss of…

LLMs are commonly trained with a learning rate (LR) warmup, followed by cosine decay to 10% of the maximum (10x decay). In a large-scale empirical study, we show that under an optimal peak LR, a simple linear decay-to-zero (D2Z) schedule…

机器学习 · 计算机科学 2025-11-25 Shane Bergsma , Nolan Dey , Gurpreet Gosal , Gavia Gray , Daria Soboleva , Joel Hestness