相关论文: Soft-to-Hard Routing in Sparse Mixture-of-Experts …
The scattering of electromagnetic waves from obstacles with wave-material interaction in thin layers on the surface is described by generalized impedance boundary conditions, which provide effective approximate models. In particular, this…
In this paper, we are concerned with the recovery of the geometric shapes of inhomogeneous inclusions from the associated far field data in electrostatics and acoustic scattering. We present a local resolution analysis and show that the…
We consider the problem of optimally compressing and caching data across a communication network. Given the data generated at edge nodes and a routing path, our goal is to determine the optimal data compression ratios and caching decisions…
Sparse MoE routing fails at domain transitions, where the current token belongs to one distribution and the next to another. In a controlled experiment (4 experts, 5 seeds), standard affinity routing assigns only 0.006 +/- 0.001 probability…
Sparse Mixture-of-Experts (MoE) have been widely adopted in recent large language models since it can efficiently scale up the model capability without increasing the inference cost. However, evaluations on broad downstream tasks reveal a…
We explore the consequences of a freeze-out criterion for heavy-ion collisions, based on pion escape probabilities from the hot and dense but rapidly expanding collision region. The influence of the expansion and the scattering rate on the…
We consider a standard distributed optimisation setting where $N$ machines, each holding a $d$-dimensional function $f_i$, aim to jointly minimise the sum of the functions $\sum_{i = 1}^N f_i (x)$. This problem arises naturally in…
Constraints can affect dramatically the behavior of diffusion processes. Recently, we analyzed a natural and a technological system and reported that they perform diffusion-like discrete steps displaying a peculiar constraint, whereby the…
Policy gradient methods are notorious for having a large variance and high sample complexity. To mitigate this, we introduce SoftTreeMax -- a generalization of softmax that employs planning. In SoftTreeMax, we extend the traditional logits…
Dynamical properties of a Lennard-Jones binary mixture embedded in an off lattice matrix of soft spheres are studied in the direct space upon supercooling by molecular dynamics simulations. On lowering temperature the smaller particles tend…
We introduce a machine-learning framework to warm-start fixed-point optimization algorithms. Our architecture consists of a neural network mapping problem parameters to warm starts, followed by a predefined number of fixed-point iterations.…
We study probabilities of various rare events for the limiting point process that appears at the random matrix hard edge. We also show a transition from hard edge to bulk behavior. Asymptotic events studied include a central limit theorem…
We study the problem of bounding path-dependent expectations (within any finite time horizon $d$) over the class of discrete-time martingales whose marginal distributions lie within a prescribed tolerance of a given collection of benchmark…
Direct numerical simulations are conducted to study the receptivity and transition mechanisms in a solitary wave boundary layer developing over randomly organized wave-like bottom topography. The boundary layer flow shows a selective…
Mixture-of-Experts (MoE) architectures have become the key to scaling modern LLMs, yet little is understood about how their sparse routing dynamics respond to multilingual data. In this work, we analyze expert routing patterns using…
In this work, we initiate the study of fault tolerant Max Cut, where given an edge-weighted undirected graph $G=(V,E)$, the goal is to find a cut $S\subseteq V$ that maximizes the total weight of edges that cross $S$ even after an adversary…
We investigate theoretically the possibility of a wetting transition induced by geometric roughness of a solid substrate for the case where the flat substrate does not show a wetting layer. Our approach makes use of a novel closed-form…
Mixture-of-experts (MoE) architectures have expanded from language modeling to automatic speech recognition (ASR). Traditional MoE methods, such as the Switch Transformer, route experts independently within each layer. Our analysis reveals…
Soft-thresholding has been widely used in neural networks. Its basic network structure is a two-layer convolution neural network with soft-thresholding. Due to the network's nature of nonlinearity and nonconvexity, the training process…
Dynamic behaviour of a WMN imposes stringent constraints on the routing policy of the network. In the shortest path based routing the shortest paths needs to be evaluated within a given time frame allowed by the WMN dynamics. The exact…