English
Related papers

Related papers: Soft-to-Hard Routing in Sparse Mixture-of-Experts …

200 papers

Mixture-of-Experts models rely on learned routers to assign tokens to experts, yet standard softmax gating provides no principled mechanism to control the tradeoff between sparsity and utilization. We propose Grassmannian MoE (GrMoE), a…

Machine Learning · Computer Science 2026-02-23 Ibne Farabi Shihab , Sanjeda Akter , Anuj Sharma

Fluid properties near rough surfaces are crucial in describing fundamental surface phenomena and modern industrial material design implementations. One of the most powerful approaches to model real rough materials is based on the surface…

Disordered Systems and Neural Networks · Physics 2021-02-10 Aleksey Khlyupin , Timur Aslyamov

Understanding the parameter estimation of softmax gating Gaussian mixture of experts has remained a long-standing open problem in the literature. It is mainly due to three fundamental theoretical challenges associated with the softmax…

Machine Learning · Statistics 2023-10-31 Huy Nguyen , TrungTin Nguyen , Nhat Ho

The softmax function is widely used in artificial neural networks for the multiclass classification problems, where the softmax transformation enforces the output to be positive and sum to one, and the corresponding loss function allows to…

Machine Learning · Computer Science 2021-12-24 Shaoshi Sun , Zhenyuan Zhang , BoCheng Huang , Pengbin Lei , Jianlin Su , Shengfeng Pan , Jiarun Cao

We introduce a weighed-loop algorithm that is applicable to any weighed graph network. It is designed to prefer a route of energetically unfavourable bonds in the lattice that can then be flipped without changing the structure inside and…

Statistical Mechanics · Physics 2017-05-19 Rick Keesman , Pepijn Overbeeke

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt,…

Machine Learning · Statistics 2025-11-25 Fanqi Yan , Huy Nguyen , Dung Le , Pedram Akbarian , Nhat Ho , Alessandro Rinaldo

Product search is the most common way for people to satisfy their shopping needs on e-commerce websites. Products are typically annotated with one of several broad categorical tags, such as "Clothing" or "Electronics", as well as…

Machine Learning · Computer Science 2021-03-03 Zhuojian Xiao , Yunjiang jiang , Guoyu Tang , Lin Liu , Sulong Xu , Yun Xiao , Weipeng Yan

Supervised fine-tuning (SFT) is a milestone in aligning large language models with human instructions and adapting them to downstream tasks. In particular, Low-Rank Adaptation (LoRA) has gained widespread attention due to its parameter…

Computation and Language · Computer Science 2025-11-05 Jia-Chen Zhang , Yu-Jie Xiong , Xi-He Qiu , Chun-Ming Xia , Fei Dai , Zheng Zhou

High-speed vehicles experience a highly challenging environment in which the free-stream Mach number and surface temperature greatly influence aerodynamic drag and heat transfer. The interplay of these two parameters strongly affects the…

Fluid Dynamics · Physics 2024-02-29 Michele Cogo , Umberto Baù , Mauro Chinappi , Matteo Bernardini , Francesco Picano

We consider non-convex stochastic optimization using first-order algorithms for which the gradient estimates may have heavy tails. We show that a combination of gradient clipping, momentum, and normalized gradient descent yields convergence…

Machine Learning · Computer Science 2021-11-10 Ashok Cutkosky , Harsh Mehta

Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width. However, we argue that effective SMoE training remains challenging because of the…

Artificial Intelligence · Computer Science 2025-05-20 Nam V. Nguyen , Huy Nguyen , Quang Pham , Van Nguyen , Savitha Ramasamy , Nhat Ho

Sparsely activated neural networks with conditional computation learn to route their inputs through different "expert" subnetworks, providing a form of modularity that densely activated models lack. Despite their possible benefits, models…

Machine Learning · Computer Science 2024-05-14 Mohammed Muqeeth , Haokun Liu , Colin Raffel

Precision cutting of soft-tissue remains a challenging problem in robotics, due to the complex and unpredictable mechanical behaviour of tissue under manipulation. Here, we consider the challenge of cutting along the boundary between two…

Robotics · Computer Science 2019-09-17 Artūras Straižys , Michael Burke , Subramanian Ramamoorthy

This paper systematically diagnoses the training failure modes of Token-Choice sparse Mixture-of-Experts (MoE) on video Diffusion Transformers. Starting from a pretrained dense model of about 5 billion parameters, we convert it into an MoE…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Haiying Sha

The traditional viewpoint on Sparse Mixture of Experts (MoE) models is that instead of training a single large expert, which is computationally expensive, we can train many small experts. The hope is that if the total parameter count of the…

Machine Learning · Computer Science 2024-09-04 Youngseog Chung , Dhruv Malik , Jeff Schneider , Yuanzhi Li , Aarti Singh

Equilibrium statistical physics is applied to layered neural networks with differentiable activation functions. A first analysis of off-line learning in soft-committee machines with a finite number (K) of hidden units learning a perfectly…

Disordered Systems and Neural Networks · Physics 2009-10-31 M. Biehl , E. Schloesser , M. Ahr

We consider the problem of designing an overlay network and routing mechanism that permits finding resources efficiently in a peer-to-peer system. We argue that many existing approaches to this problem can be modeled as the construction of…

Data Structures and Algorithms · Computer Science 2007-05-23 James Aspnes , Zoe Diamadi , Gauri Shah

Consider the classical $(2+1)$-dimensional Solid-On-Solid model above a hard wall on an $L\times L$ box of $\bbZ^2$. The model describes a crystal surface by assigning a non-negative integer height $\eta_x$ to each site $x$ in the box and 0…

Probability · Mathematics 2013-02-28 Pietro Caputo , Eyal Lubetzky , Fabio Martinelli , Allan Sly , Fabio Lucio Toninelli

We present a molecular dynamics based method for computing accurately short-range structural forces resulting from the overlap of spatially diffuse solid-liquid interfaces at wetted grain boundaries close to the melting point. The method is…

Materials Science · Physics 2009-11-13 J. J. Hoyt , David Olmsted , Saryu Jindal , Mark Asta , Alain Karma

Resource-efficient machine learning increasingly uses sparse Mixture-of-Experts (MoE) architectures, where the gate acts as both a learning component and a routing interface controlling computation, communication, and accuracy. Motivated by…

Machine Learning · Computer Science 2026-05-08 Mohammad Reza Deylam Salehi , Ali Khalesi