中文
相关论文

相关论文: Spectral Condition for $\mu$P under Width-Depth Sc…

200 篇论文

Deep learning models have become a cornerstone of modern AI research, yet their initializations and learning rates may at times be set in an opaque or ad-hoc fashion due to the high cost of hyperparameter sweeps. The $\mu$-Parameterization…

机器学习 · 计算机科学 2025-02-17 Lucas Lingle

Recent advances in Remote Sensing Foundation Models (RSFMs) have led to significant breakthroughs in the field. While many RSFMs have been pretrained with massive optical imagery, more multispectral/hyperspectral data remain lack of the…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Yuxiang Zhang , Wei Li , Mengmeng Zhang , Jiawei Han , Ran Tao , Shunlin Liang

The Single Program Multiple Data (SPMD) paradigm provides a unified abstraction to annotate various parallel dimensions in distributed deep learning (DL) training. With SPMD, users can write training programs from the viewpoint of a single…

分布式、并行与集群计算 · 计算机科学 2025-04-30 Haoyang Li , Fangcheng Fu , Hao Ge , Sheng Lin , Xuanyu Wang , Jiawen Niu , Xupeng Miao , Bin Cui

In recent years, there has been a growing interest in deep learning-based pansharpening. Thus far, research has mainly focused on architectures. Nonetheless, model training is an equally important issue. A first problem is the absence of…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Matteo Ciotola , Sergio Vitale , Antonio Mazza , Giovanni Poggi , Giuseppe Scarpa

Spectral image reconstruction is an important task in snapshot compressed imaging. This paper aims to propose a new end-to-end framework with iterative capabilities similar to a deep unfolding network to improve reconstruction accuracy,…

图像与视频处理 · 电气工程与系统科学 2023-05-09 Zeyu Cai , Jian Yu , Ziyu Zhang , Chengqian Jin , Feipeng Da

This paper proposes a parameter collaborative optimization algorithm for large language models, enhanced with graph spectral analysis. The goal is to improve both fine-tuning efficiency and structural awareness during training. In the…

机器学习 · 计算机科学 2025-06-03 Hanlu Zhang , Yumeng Ma , Shuo Wang , Guiran Liu , Binrong Zhu

We present a modular, extensible likelihood framework for spectroscopic inference based on synthetic model spectra. The subtraction of an imperfect model from a continuously sampled spectrum introduces covariance between adjacent datapoints…

太阳与恒星天体物理 · 物理学 2015-10-21 Ian Czekala , Sean M. Andrews , Kaisey S. Mandel , David W. Hogg , Gregory M. Green

We extend our two-scale neural-network method for scalar singularly perturbed problems with one small parameter to dynamical systems with multiple small parameters. To accommodate multiple small parameters, we use a single effective scale…

数值分析 · 数学 2026-05-05 Qiao Zhuang , Taorui Wang , Rita Wanjiku , Majid Bani-Yaghoub , Zhongqiang Zhang

In this work, we present a novel non-rigid shape matching framework based on multi-resolution functional maps with spectral attention. Existing functional map learning methods all rely on the critical choice of the spectral resolution…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Lei Li , Nicolas Donati , Maks Ovsjanikov

A robust $hp$-adaptive finite element framework is presented for the investigation of static cracks in materials characterized by complex, pointwise density variations. Within such heterogeneous media, the equilibrium equation governed by…

数值分析 · 数学 2025-12-29 S. M. Mallikarjunaiah

We provide the first proof of learning rate transfer with width in a linear multi-layer perceptron (MLP) parametrized with $\mu$P, a neural network parameterization designed to ``maximize'' feature learning in the infinite-width limit. We…

机器学习 · 统计学 2026-02-26 Soufiane Hayou

This paper presents a novel holistic deep learning framework that simultaneously addresses the challenges of vulnerability to input perturbations, overparametrization, and performance instability from different train-validation splits. The…

For the accurate representation and reconstruction of band-limited signals on the sphere, an optimal-dimensionality sampling scheme has been recently proposed which requires the optimal number of samples equal to the number of degrees of…

信息论 · 计算机科学 2017-09-11 Wajeeha Nafees , Zubair Khalid , Rodney A. Kennedy , Jason D. McEwen

Recent works have highlighted scale invariance or symmetry present in the weight space of a typical deep network and the adverse effect it has on the Euclidean gradient based stochastic gradient descent optimization. In this work, we show…

机器学习 · 计算机科学 2015-11-04 Vijay Badrinarayanan , Bamdev Mishra , Roberto Cipolla

Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovered Maximal Update Parametrization (muP), many optimal HPs…

Adapting pre-trained foundation models for various downstream tasks has been prevalent in artificial intelligence. Due to the vast number of tasks and high costs, adjusting all parameters becomes unfeasible. To mitigate this, several…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Chongjie Si , Xuehui Wang , Xue Yang , Zhengqin Xu , Qingyun Li , Jifeng Dai , Yu Qiao , Xiaokang Yang , Wei Shen

While spectral-based optimizers like Muon operate directly on the spectrum of updates, standard adaptive methods such as AdamW do not account for the spectral structure of weights and gradients, leaving them vulnerable to two empirical…

机器学习 · 计算机科学 2026-05-29 Xiaowen Jiang , Andrei Semenov , Sebastian U. Stich

Stability evaluation of a weight-update system of higher-order neural units (HONUs) with polynomial aggregation of neural inputs (also known as classes of polynomial neural networks) for adaptation of both feedforward and recurrent HONUs by…

神经与进化计算 · 计算机科学 2016-06-24 Ivo Bukovsky , Noriyasu Homma

Recent developments in Parameter-Efficient Fine-Tuning (PEFT) methods for pretrained deep neural networks have captured widespread interest. In this work, we study the enhancement of current PEFT methods by incorporating the spectral…

机器学习 · 计算机科学 2024-11-05 Fangzhao Zhang , Mert Pilanci

Large language models (LLMs) have significantly advanced natural language processing, but their massive parameter counts create substantial computational and memory challenges during deployment. Post-training quantization (PTQ) has emerged…

机器学习 · 计算机科学 2025-11-25 Cuong Pham , Hoang Anh Dung , Cuong C. Nguyen , Trung Le , Gustavo Carneiro , Thanh-Toan Do