中文
相关论文

相关论文: Disentangling Linear Mode-Connectivity

200 篇论文

Pre-trained language models (LMs), such as BERT (Devlin et al., 2018) and its variants, have led to significant improvements on various NLP tasks in past years. However, a theoretical framework for studying their relationships is still…

计算与语言 · 计算机科学 2022-10-24 Hao Zhang

We consider the problem of predicting link formation in Social Learning Networks (SLN), a type of social network that forms when people learn from one another through structured interactions. While link prediction has been studied for…

社会与信息网络 · 计算机科学 2023-01-05 Rajeev Sahay , Serena Nicoll , Minjun Zhang , Tsung-Yen Yang , Carlee Joe-Wong , Kerrie A. Douglas , Christopher G Brinton

Linear interpolation between initial neural network parameters and converged parameters after training with stochastic gradient descent (SGD) typically leads to a monotonic decrease in the training objective. This Monotonic Linear…

机器学习 · 计算机科学 2021-04-26 James Lucas , Juhan Bae , Michael R. Zhang , Stanislav Fort , Richard Zemel , Roger Grosse

The ability to remove features from the input of machine learning models is very important to understand and interpret model predictions. However, this is non-trivial for vision models since masking out parts of the input image typically…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Sriram Balasubramanian , Soheil Feizi

In multimode fiber transmission systems, mode-dependent loss and gain (collectively referred to as MDL) pose fundamental performance limitations. In the regime of strong mode coupling, the statistics of MDL (expressed in decibels or log…

光学 · 物理学 2013-01-15 Keang-Po Ho , Joseph M. Kahn

Seeking effective neural networks is a critical and practical field in deep learning. Besides designing the depth, type of convolution, normalization, and nonlinearities, the topological connectivity of neural networks is also important.…

计算机视觉与模式识别 · 计算机科学 2020-08-20 Kun Yuan , Quanquan Li , Jing Shao , Junjie Yan

This paper discusses two main themes. First, it investigates the formation of a spatiotemporal cognitive map (mental image) of a road network in travelers memory, which entails the travelers global conceptual understanding of congestion or…

物理与社会 · 物理学 2020-02-25 Navid Khademi , Ramin Saedi

With the proliferation of network devices and rapid development in information technology, networks such as Internet of Things are increasing in size and becoming more complex with heterogeneous wired and wireless links. In such networks,…

网络与互联网体系结构 · 计算机科学 2019-03-28 Srinikethan Madapuzi Srinivasan , Tram Truong-Huu , Mohan Gurusamy

We survey the model merging literature through the lens of loss landscape geometry to connect observations from empirical studies on model merging and loss landscape analysis to phenomena that govern neural network training and the…

Averaging neural network parameters is an intuitive method for fusing the knowledge of two independent models. It is most prominently used in federated learning. If models are averaged at the end of training, this can only lead to a good…

机器学习 · 计算机科学 2024-03-20 Linara Adilova , Maksym Andriushchenko , Michael Kamp , Asja Fischer , Martin Jaggi

Hidden parameters are latent variables in reinforcement learning (RL) environments that are constant over the course of a trajectory. Understanding what, if any, hidden parameters affect a particular environment can aid both the development…

机器学习 · 计算机科学 2022-11-30 Christopher Reale , Rebecca Russell

Disentanglement via mechanism sparsity was introduced recently as a principled approach to extract latent factors without supervision when the causal graph relating them in time is sparse, and/or when actions are observed and affect them…

机器学习 · 统计学 2022-07-19 Sébastien Lachapelle , Simon Lacoste-Julien

In this paper, we propose a method to perform empirical analysis of the loss landscape of machine learning (ML) models. The method is applied to two ML models for scientific sensing, which necessitates quantization to be deployed and are…

机器学习 · 计算机科学 2025-02-17 Tommaso Baldi , Javier Campos , Olivia Weng , Caleb Geniesse , Nhan Tran , Ryan Kastner , Alessandro Biondi

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot explain many…

机器学习 · 计算机科学 2022-06-23 Chao Ma , Daniel Kunin , Lei Wu , Lexing Ying

Visual long-range interaction refers to modeling dependencies between distant feature points or blocks within an image, which can significantly enhance the model's robustness. Both CNN and Transformer can establish long-range interactions…

计算机视觉与模式识别 · 计算机科学 2023-03-13 Xiuxiu Bai , Shuaishuai Zhao , Yao Gao , Zhe Liu

Many natural systems are organized as networks, in which the nodes (be they cells, individuals or populations) interact in a time-dependent fashion. The dynamic behavior of these networks depends on how these nodes are connected, which can…

神经元与认知 · 定量生物学 2015-06-22 Anca Radulescu , Sergio Verduzco-Flores

Link prediction (LP) is an important problem in network science and machine learning research. The state-of-the-art LP methods are usually evaluated in a uniform setup, ignoring several factors associated with the data and application…

社会与信息网络 · 计算机科学 2025-07-21 Bhargavi Kalyani , A Rama Prasad Mathi , Niladri Sett

Understanding the internal mechanisms of large language models (LLMs) remains a challenging and complex endeavor. Even fundamental questions, such as how fine-tuning affects model behavior, often require extensive empirical evaluation. In…

During multimodal model training and testing, certain data modalities may be absent due to sensor limitations, cost constraints, privacy concerns, or data loss, negatively affecting performance. Multimodal learning techniques designed to…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Renjie Wu , Hu Wang , Hsiang-Ting Chen , Gustavo Carneiro

In these six lectures, we examine what can be learnt about the behavior of multi-layer neural networks from the analysis of linear models. We first recall the correspondence between neural networks and linear models via the so-called lazy…

机器学习 · 统计学 2023-08-28 Theodor Misiakiewicz , Andrea Montanari