English
Related papers

Related papers: MuCon: Clipped Muon Updates for LLM Training

200 papers

In contrast to Part I of this treatise [1] that focuses on the optimization problems associated with single matrix variables, in this paper, we investigate the application of the matrix-monotonic optimization framework in the optimization…

Information Theory · Computer Science 2021-02-24 Chengwen Xing , Shuai Wang , Sheng Chen , Shaodan Ma , H. Vincent Poor , Lajos Hanzo

This study proposes a novel hierarchical prior for inferring possibly low-rank matrices measured with noise. We consider three-component matrix factorization, as in singular value decomposition, and its fully Bayesian inference. The…

Methodology · Statistics 2020-10-09 Masahiro Tanaka

This paper suggests two novel ideas to develop new proximal variable-metric methods for solving a class of composite convex optimization problems. The first idea is a new parameterization of the optimality condition which allows us to…

Optimization and Control · Mathematics 2018-12-14 Quoc Tran-Dinh , Liang Ling , Kim-Chuan Toh

PDE-constrained optimization aims at finding optimal setups for partial differential equations so that relevant quantities are minimized. Including sparsity promoting terms in the formulation of such problems results in more practically…

Numerical Analysis · Mathematics 2016-11-23 Margherita Porcelli , Valeria Simoncini , Martin Stoll

The recently proposed MUonE experiment at CERN aims at providing a novel determination of the leading order hadronic contribution to the muon anomalous magnetic moment through the study of elastic muon-electron scattering at relatively…

High Energy Physics - Phenomenology · Physics 2020-12-02 C. M. Carloni Calame , M. Chiesa , Syed Mehedi Hasan , G. Montagna , O. Nicrosini , F. Piccinini

Accurately estimating the normalization term (also known as the partition function) in the contrastive loss is a central challenge for training Contrastive Language-Image Pre-training (CLIP) models. Conventional methods rely on large…

Machine Learning · Computer Science 2026-03-05 Xiyuan Wei , Chih-Jen Lin , Tianbao Yang

We introduce Sauron, a filter pruning method that eliminates redundant feature maps of convolutional neural networks (CNNs). Sauron optimizes, jointly with the loss function, a regularization term that promotes feature maps clustering at…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Juan Miguel Valverde , Artem Shatillo , Jussi Tohka

Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hyperparameters should scale with network width $N$, expert…

Machine Learning · Computer Science 2026-05-15 Leena Chennuru Vankadara , Moritz Haas , Luke Hayward , Sebastian Bordt , Alessandro Breccia

The Mixture of Experts (MoE) paradigm provides a powerful way to decompose dense layers into smaller, modular computations often more amenable to human interpretation, debugging, and editability. However, a major challenge lies in the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 James Oldfield , Markos Georgopoulos , Grigorios G. Chrysos , Christos Tzelepis , Yannis Panagakis , Mihalis A. Nicolaou , Jiankang Deng , Ioannis Patras

Matrix learning is at the core of many machine learning problems. A number of real-world applications such as collaborative filtering and text mining can be formulated as a low-rank matrix completion problem, which recovers incomplete…

Machine Learning · Computer Science 2021-02-23 Yaqing Wang , Quanming Yao , James T. Kwok

Moving Morphable Component (MMC) based topology optimization approach is an explicit algorithm since the boundary of the entity explicitly described by its functions. Compared with other pixel or node point-based algorithms, it is optimized…

Numerical Analysis · Mathematics 2019-10-17 Xinchao Jiang , Hu Wang , Yu Li , Kangjia Mo

Today's unsupervised image segmentation algorithms often segment suboptimally. Modern graph-cut based approaches rely on high-dimensional attention maps from Transformer-based foundation models, typically employing a relaxed Normalized Cut…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Xiao Zhang , Xiangyu Han , Xiwen Lai , Yao Sun , Pei Zhang , Konrad Kording

An expression for the Dalitz plot of the semileptonic decay of a charged kaon, including radiative corrections to order O[\alpha \pi)(q/M_1)], where q is the four-momentum transfer and M_1 is the mass of the decaying kaon, is obtained.…

High Energy Physics - Phenomenology · Physics 2020-04-14 M. J. Sanchez-Glez , A. Martinez , C. Juarez-Leon , M. Neri , J. J. Torres , Ruben Flores-Mendieta

Finite element model updating is challenging because 1) the problem is oftentimes underdetermined while the measurements are limited and/or incomplete; 2) many combinations of parameters may yield responses that are similar with respect to…

Applications · Statistics 2021-07-28 Kai Zhou , Jiong Tang

Composite minimization is a powerful framework in large-scale convex optimization, based on decoupling of the objective function into terms with structurally different properties and allowing for more flexible algorithmic design. We…

Optimization and Control · Mathematics 2023-02-17 Jelena Diakonikolas , Cristóbal Guzmán

A model-independent expression for the Dalitz plot of the semileptonic decays of a neutral kaon $K_{\mu 3}^0$, including radiative corrections to order $\mathcal{O}[(\alpha / \pi )(q/M_1)]$, where $q$ is the four-momentum transfer and $M_1$…

High Energy Physics - Phenomenology · Physics 2023-11-10 J. Vieyra , A. Martínez , M. Neri , A. Hernández-Galeana

As the share of renewable generation in large power systems continues to increase, the operation of power systems becomes increasingly challenging. The constantly shifting mix of renewable and conventional generation leads to largely…

Systems and Control · Electrical Eng. & Systems 2020-05-11 Amer Mešanović , Ulrich Münz , Rolf Findeisen

Token reduction accelerates Multimodal Large Language Models (MLLMs) by reducing excessive tokens, but overlooks structural redundancy differences, where critical and redundant modules process identical token loads. For fine-grained…

Machine Learning · Computer Science 2025-11-14 Aoming Liu , Reuben Tan , Boqing Gong , Bryan A. Plummer

The past several decades have seen significant advancement in applications using cosmic-ray muons for tomography scanning of unknown objects. One of the most promising developments is the application of this technique in border security for…

Instrumentation and Detectors · Physics 2026-05-05 Z. Zaher , H. Lay , T. Dorigo , A. Giammanco , V. Gulik , C. Hrytsiuk , V. A. Kudryavtsev , M. Lagrange , T. Metspalu , G. C. Strong , C. Turkoglu , P. Vischia

The Pseudo-Marginal (PM) algorithm is a popular Markov chain Monte Carlo (MCMC) method used to sample from a target distribution when its density is inaccessible, but can be estimated with a non-negative unbiased estimator. Its performance…

Computation · Statistics 2025-09-30 Sarra Abaoubida , Mylène Bédard , Florian Maire