English
Related papers

Related papers: Soft-to-Hard Routing in Sparse Mixture-of-Experts …

200 papers

Many machine learning systems make constrained decisions by optimizing factorized objectives, but the context-specific objective is often treated as fixed. We study contextual decision-weight learning: from logged decisions and proxy…

Machine Learning · Computer Science 2026-05-04 Renjun Hu , Hyun-Soo Ahn

Mixture-of-Experts (MoE) has emerged as a powerful paradigm for scaling model capacity while preserving computational efficiency. Despite its notable success in large language models (LLMs), existing attempts to apply MoE to Diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yujie Wei , Shiwei Zhang , Hangjie Yuan , Yujin Han , Zhekai Chen , Jiayu Wang , Difan Zou , Xihui Liu , Yingya Zhang , Yu Liu , Hongming Shan

Mixture-of-Experts (MoE) architectures decompose prediction tasks into specialized expert sub-networks selected by a gating mechanism. This letter adopts a communication-theoretic view of MoE gating, modeling the gate as a stochastic…

Machine Learning · Statistics 2026-03-27 Ali Khalesi , Mohammad Reza Deylam Salehi

We present a formal model to represent and solve the unicast/multicast routing problem in networks with Quality of Service (QoS) requirements. To attain this, first we translate the network adapting it to a weighted graph (unicast) or…

Logic in Computer Science · Computer Science 2009-09-29 Stefano Bistarelli , Ugo Montanari , Francesca Rossi , Francesco Santini

When analyzing a dataset, it can be useful to assess how smooth the decision boundaries need to be for a model to better fit the data. This paper addresses this question by proposing the quantification of how much should the 'rigid'…

Machine Learning · Computer Science 2022-10-10 Anthea Mérida , Argyris Kalogeratos , Mathilde Mougeot

Reference [11] investigated the almost sure weak convergence of block-coordinate fixed point algorithms and discussed their applications to nonlinear analysis and optimization. This algorithmic framework features random sweeping rules to…

Optimization and Control · Mathematics 2018-04-17 Patrick L. Combettes , Jean-Christophe Pesquet

Mixture-of-Experts (MoE) models are a promising way to scale up model capacity without significantly increasing computational cost. A key component of MoEs is the router, which decides which subset of parameters (experts) process which…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Tianlin Liu , Mathieu Blondel , Carlos Riquelme , Joan Puigcerver

Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central role in determining expert specialization and overall model performance. However, many modern MoE…

Machine Learning · Computer Science 2026-05-15 Minghao Yang , Ren Togo , Guang Li , Takahiro Ogawa , Miki Haseyama

We describe a distributed randomized algorithm computing approximate distances and routes that approximate shortest paths. Let n denote the number of nodes in the graph, and let HD denote the hop diameter of the graph, i.e., the diameter of…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-11-05 Christoph Lenzen , Boaz Patt-Shamir

This work studies the throughput scaling laws of ad hoc wireless networks in the limit of a large number of nodes. A random connections model is assumed in which the channel connections between the nodes are drawn independently from a…

Information Theory · Computer Science 2016-11-18 Shengshan Cui , Alexander M. Haimovich , Oren Somekh , H. Vincent Poor , Shlomo Shamai

In this paper, we explore statistical versus computational trade-off to address a basic question in the application of a distributed algorithm: what is the minimal computational cost in obtaining statistical optimality? In smoothing spline…

Statistics Theory · Mathematics 2017-07-25 Zuofeng Shang , Guang Cheng

This paper is concerned with the hard thresholding operator which sets all but the $k$ largest absolute elements of a vector to zero. We establish a {\em tight} bound to quantitatively characterize the deviation of the thresholded solution…

Machine Learning · Statistics 2020-08-12 Jie Shen , Ping Li

In this paper we present matrix game-theoretic models for joint routing, network coding, and scheduling problem. First routing and network coding are modeled by using a new approach based on compressed topology matrix that takes into…

Signal Processing · Electrical Eng. & Systems 2018-03-14 Ebrahim Karami , Savo Glisic

Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to expert assignments. In this work, we ask whether a dedicated…

Artificial Intelligence · Computer Science 2026-04-02 Jama Hussein Mohamud , Drew Wagner , Mirco Ravanelli

To account for the randomness of propagation channels and interference levels in hierarchical spectrum sharing, a novel approach to multihop routing is introduced for cognitive random access networks, whereby packets are randomly routed…

Optimization and Control · Mathematics 2012-07-05 Emiliano Dall'Anese , Georgios B. Giannakis

This paper presents a tractable analytical framework for the exact calculation of probability of node isolation and minimum node degree distribution when $N$ sensor nodes are independently and uniformly distributed inside a finite square…

Information Theory · Computer Science 2015-07-09 Zubair Khalid , Salman Durrani , Jing Guo

In this aerothermal study, we performed a two-dimensional steady-state Computational Fluid Dynamics (CFD) and heat conduction simulation at Mach 6. The key to our methodology was a one-way coupling between CFD surface temperature as a…

We consider the scaling limit of a generic ferromagnetic system with a continuous phase transition, on the half plane with boundary conditions leading to the equilibrium of two different phases below criticality. We use general properties…

Statistical Mechanics · Physics 2014-10-09 Gesualdo Delfino , Alessio Squarcini

We study the hard-core model defined on independent sets, where each independent set I in a graph G is weighted proportionally to $\lambda^{|I|}$, for a positive real parameter $\lambda$. For large $\lambda$, computing the partition…

Probability · Mathematics 2011-08-15 Ricardo Restrepo , Jinwoo Shin , Prasad Tetali , Eric Vigoda , Linji Yang

Finite mixture models provide a flexible framework for approximating and estimating multivariate probability densities. We study mixtures formed from translated and rescaled copies of a fixed density kernel and obtain explicit results for…

Statistics Theory · Mathematics 2026-04-24 Hien Duy Nguyen , TrungTin Nguyen , Jacob Westerhout , Xin Guo
‹ Prev 1 4 5 6 7 8 10 Next ›