中文
相关论文

相关论文: Analyzing the Role of Permutation Invariance in Li…

200 篇论文

In this paper, we conjecture that if the permutation invariance of neural networks is taken into account, SGD solutions will likely have no barrier in the linear interpolation between them. Although it is a bold conjecture, we show how…

机器学习 · 计算机科学 2022-07-06 Rahim Entezari , Hanie Sedghi , Olga Saukh , Behnam Neyshabur

Linear Mode Connectivity (LMC) refers to the phenomenon that performance remains consistent for linearly interpolated models in the parameter space. For independently optimized model pairs from different random initializations, achieving…

机器学习 · 计算机科学 2025-02-17 Ryuichi Kanoh , Mahito Sugiyama

Recently, Ainsworth et al. empirically demonstrated that, given two independently trained models, applying a parameter permutation that preserves the input-output behavior allows the two models to be connected by a low-loss linear path.…

机器学习 · 计算机科学 2026-03-09 Akira Ito , Masanori Yamada , Daiki Chijiwa , Atsutoshi Kumagai

Neural networks typically exhibit permutation symmetries which contribute to the non-convexity of the networks' loss landscapes, since linearly interpolating between two permuted versions of a trained network tends to encounter a high loss…

机器学习 · 计算机科学 2024-04-10 Ekansh Sharma , Devin Kwok , Tom Denton , Daniel M. Roy , David Rolnick , Gintare Karolina Dziugaite

Understanding the geometry of neural network loss landscapes is a central question in deep learning, with implications for generalization and optimization. A striking phenomenon is linear mode connectivity (LMC), where independently trained…

Recently, Ainsworth et al. showed that using weight matching (WM) to minimize the $L^2$ distance in a permutation search of model parameters effectively identifies permutations that satisfy linear mode connectivity (LMC), where the loss…

机器学习 · 计算机科学 2025-04-09 Akira Ito , Masanori Yamada , Atsutoshi Kumagai

The presence of linear paths in parameter space between two different network solutions in certain cases, i.e., linear mode connectivity (LMC), has garnered interest from both theoretical and practical fronts. There has been significant…

机器学习 · 计算机科学 2024-06-25 Sidak Pal Singh , Linara Adilova , Michael Kamp , Asja Fischer , Bernhard Schölkopf , Thomas Hofmann

The phenomenon of linear mode connectivity (LMC) links several aspects of deep learning, including training stability under noisy stochastic gradients, the smoothness and generalization of local minima (basins), the similarity and…

机器学习 · 计算机科学 2025-11-07 C. Hepburn , T. Zielke , A. P. Raulf

The question of how and why the phenomenon of mode connectivity occurs in training deep neural networks has gained remarkable attention in the research community. From a theoretical perspective, two possible explanations have been proposed:…

机器学习 · 计算机科学 2021-10-22 Quynh Nguyen , Pierre Brechet , Marco Mondelli

Linear mode-connectivity (LMC) (or lack thereof) is one of the intriguing characteristics of neural network loss landscapes. While empirically well established, it unfortunately still lacks a proper theoretical understanding. Even worse,…

机器学习 · 计算机科学 2023-12-18 Gul Sena Altintas , Gregor Bachmann , Lorenzo Noci , Thomas Hofmann

In this paper we look into the conjecture of Entezari et al. (2021) which states that if the permutation invariance of neural networks is taken into account, then there is likely no loss barrier to the linear interpolation between SGD…

机器学习 · 计算机科学 2023-09-26 Keller Jordan , Hanie Sedghi , Olga Saukh , Rahim Entezari , Behnam Neyshabur

Recent work has revealed many intriguing empirical phenomena in neural network training, despite the poorly understood and highly complex loss landscapes and training dynamics. One of these phenomena, Linear Mode Connectivity (LMC), has…

机器学习 · 计算机科学 2023-11-14 Zhanpeng Zhou , Yongyi Yang , Xiaojiang Yang , Junchi Yan , Wei Hu

Understanding the properties of neural networks trained via stochastic gradient descent (SGD) is at the heart of the theory of deep learning. In this work, we take a mean-field view, and consider a two-layer ReLU network trained via SGD for…

机器学习 · 计算机科学 2022-05-02 Alexander Shevchenko , Vyacheslav Kungurtsev , Marco Mondelli

We explore element-wise convex combinations of two permutation-aligned neural network parameter vectors $\Theta_A$ and $\Theta_B$ of size $d$. We conduct extensive experiments by examining various distributions of such model combinations…

Managing group-delay (GD) spread is vital for reducing the complexity of digital signal processing (DSP) in long-haul systems using multi-mode fibers. GD compensation through mode permutation, which involves periodically exchanging power…

光学 · 物理学 2025-05-13 Oleksiy Krutko , Rebecca Refaee , Anirudh Vijay , Nika Zahedi , J. M. Kahn

The energy landscape of high-dimensional non-convex optimization problems is crucial to understanding the effectiveness of modern deep neural network architectures. Recent works have experimentally shown that two different solutions found…

机器学习 · 计算机科学 2024-03-04 Damien Ferbach , Baptiste Goujaud , Gauthier Gidel , Aymeric Dieuleveut

Implicit deep learning has recently become popular in the machine learning community since these implicit models can achieve competitive performance with state-of-the-art deep networks while using significantly less memory and computational…

机器学习 · 计算机科学 2022-05-17 Tianxiang Gao , Hongyang Gao

The success of deep learning is due in large part to our ability to solve certain massive non-convex optimization problems with relative ease. Though non-convex optimization is NP-hard, simple algorithms -- often variants of stochastic…

机器学习 · 计算机科学 2023-03-03 Samuel K. Ainsworth , Jonathan Hayase , Siddhartha Srinivasa

We extend the concept of loss landscape mode connectivity to the input space of deep neural networks. Mode connectivity was originally studied within parameter space, where it describes the existence of low-loss paths between different…

机器学习 · 计算机科学 2024-09-10 Jakub Vrabel , Ori Shem-Ur , Yaron Oz , David Krueger

The loss landscapes of deep neural networks are not well understood due to their high nonconvexity. Empirically, the local minima of these loss functions can be connected by a learned curve in model space, along which the loss remains…

机器学习 · 计算机科学 2020-12-11 N. Joseph Tatro , Pin-Yu Chen , Payel Das , Igor Melnyk , Prasanna Sattigeri , Rongjie Lai
‹ 上一页 1 2 3 10 下一页 ›