English
Related papers

Related papers: Spectral Condition for $\mu$P under Width-Depth Sc…

200 papers

We propose Complete-muE, a framework which targets hyperparameter transfer across dense FFN and any Mixture-of-Experts (MoE) setups in transformer blocks. Existing tools such as $\mu$P (requires fixed architectue) or SDE (requires fixed…

Machine Learning · Computer Science 2026-05-25 Hongwu Peng , Ohiremen Dibua , Yuanjun Xiong , Yifan Gong , Jianming Zhang , Yan Kang

Geometric data pruning methods, while practical for leveraging pretrained models, are fundamentally unstable. Their reliance on extrinsic geometry renders them highly sensitive to latent space perturbations, causing performance to degrade…

Machine Learning · Computer Science 2026-05-11 Arjun Roy , Prajna G. Malettira , Manish Nagaraj , Kaushik Roy

In this paper, we consider the approximate weighted graph matching problem and introduce stable and informative first and second order compatibility terms suitable for inclusion into the popular integer quadratic program formulation. Our…

Computer Vision and Pattern Recognition · Computer Science 2018-04-12 Nan Hu , Raif M. Rustamov , Leonidas Guibas

Hidden Markov models (HMMs) are widely used statistical models for modeling sequential data. The parameter estimation for HMMs from time series data is an important learning problem. The predominant methods for parameter estimation are…

Machine Learning · Computer Science 2014-04-30 Carl Mattfeld

In modern engineering scenarios, there is often a strict upper bound on the number of algorithm iterations that can be performed within a given time limit. This raises the question of optimal algorithmic configuration for a fixed and finite…

Optimization and Control · Mathematics 2024-12-31 Yushun Zhang , Dmitry Rybin , Zhi-Quan Luo

Synthetic Aperture Radar (SAR) imagery enables all-weather, day-and-night Earth observation; however, it remains difficult to interpret due to speckle noise and other intrinsic imaging artifacts. Sentinel-1 (S1) constitutes one of the most…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Juan Francisco Amieva , Christian Ayala , Roberto Del Prete , Mikel Galar

Machine learning training methods depend plentifully and intricately on hyperparameters, motivating automated strategies for their optimisation. Many existing algorithms restart training for each new hyperparameter choice, at considerable…

Machine Learning · Computer Science 2022-04-22 Ross M. Clarke , Elre T. Oldewage , José Miguel Hernández-Lobato

Pansharpening fuses a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to produce a high-resolution multispectral (HRMS) image. A key difficulty is that jointly processing PAN and MS features often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Haoyu Zhang , Junhan Luo , Yugang Cao , Jie Huang , Liangjian-Deng

We present a parameter-free variant of the halo model that significantly improves the precision of matter clustering predictions, particularly in the challenging 1-halo to 2-halo transition regime, where standard halo models often fail.…

Cosmology and Nongalactic Astrophysics · Physics 2026-03-18 Samuel Brieden , Florian Beutler , Marcos Pellejero-Ibañez

Parameter-efficient fine-tuning (PEFT) has become a common method for fine-tuning large language models, where a base model can serve multiple users through PEFT module switching. To enhance user experience, base models require periodic…

Computation and Language · Computer Science 2025-06-10 Naibin Gu , Peng Fu , Xiyu Liu , Ke Ma , Zheng Lin , Weiping Wang

The practice of deep learning has shown that neural networks generalize remarkably well even with an extreme number of learned parameters. This appears to contradict traditional statistical wisdom, in which a trade-off between model…

Machine Learning · Computer Science 2023-02-21 Yifei Wang , Yixuan Hua , Emmanuel Candés , Mert Pilanci

Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter blocks; Muon is the motivating example. We study an idealized deterministic, continuous-time,…

Optimization and Control · Mathematics 2026-05-11 Gabriel Peyré

This work introduces an adaptive mesh refinement technique for hierarchical hybrid grids with the goal to reach scalability and maintain excellent performance on massively parallel computer systems. On the block structured hierarchical…

Numerical Analysis · Mathematics 2025-08-11 Benjamin Mann , Ulrich Rüde

Foundation models have revolutionized artificial intelligence by providing robust, versatile architectures pre-trained on large-scale datasets. However, adapting these massive models to specific downstream tasks requires fine-tuning, which…

Machine Learning · Computer Science 2025-05-01 Jieming Bian , Yuanzhe Peng , Lei Wang , Yin Huang , Jie Xu

Signal-background classification is a central problem in High-Energy Physics (HEP), that plays a major role for the discovery of new fundamental particles. A recent method -- the Parametric Neural Network (pNN) -- leverages multiple signal…

High Energy Physics - Experiment · Physics 2022-11-15 Luca Anzalone , Tommaso Diotalevi , Daniele Bonacorsi

We explore the application of concepts developed in High Energy Physics (HEP) for advanced medical data analysis. Our study case is a problem with high social impact: clinically-feasible Magnetic Resonance Fingerprinting (MRF). MRF is a…

Single image super-resolution traditionally assumes spatially-invariant degradation models, yet real-world imaging systems exhibit complex distance-dependent effects including atmospheric scattering, depth-of-field variations, and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Tianhao Guo , Bingjie Lu , Feng Wang , Zhengyang Lu

Full-waveform inversion (FWI) is pivotal for reconstructing high-resolution subsurface velocity models but remains computationally intensive and ill-posed. While deep learning approaches promise efficiency, existing Convolutional Neural…

Machine Learning · Computer Science 2026-05-05 Zhenyu Wang , Peiyuan Li , Yongxiang Shi , Ruoyu Wu , Chenfei Liao , Lei Zhang

We give the first approximation algorithm for mixed packing and covering semidefinite programs (SDPs) with polylogarithmic dependence on width. Mixed packing and covering SDPs constitute a fundamental algorithmic primitive with recent…

Data Structures and Algorithms · Computer Science 2021-07-13 Arun Jambulapati , Yin Tat Lee , Jerry Li , Swati Padmanabhan , Kevin Tian

Despite deep neural networks' powerful representation learning capabilities, theoretical understanding of how networks can simultaneously achieve meaningful feature learning and global convergence remains elusive. Existing approaches like…

Machine Learning · Computer Science 2025-07-23 Zixiang Chen , Greg Yang , Qingyue Zhao , Quanquan Gu
‹ Prev 1 8 9 10 Next ›