English
Related papers

Related papers: Sigmoid Gating is More Sample Efficient than Softm…

200 papers

Large transformer models have achieved state-of-the-art results in numerous natural language processing tasks. Among the pivotal components of the transformer architecture, the attention mechanism plays a crucial role in capturing token…

Computation and Language · Computer Science 2026-03-16 Yichuan Deng , Zhao Song , Kaijun Yuan , Tianyi Zhou

It is well-known that trimmed sample means are robust against heavy tails and data contamination. This paper analyzes the performance of trimmed means and related methods in two novel contexts. The first one consists of estimating…

Statistics Theory · Mathematics 2025-12-03 Roberto I. Oliveira , Lucas Resende

Computations for the softmax function are significantly expensive when the number of output classes is large. In this paper, we present a novel softmax inference speedup method, Doubly Sparse Softmax (DS-Softmax), that leverages sparse…

Machine Learning · Computer Science 2019-07-04 Shun Liao , Ting Chen , Tian Lin , Denny Zhou , Chong Wang

Activation functions are core components of all deep learning architectures. Currently, the most popular activation functions are smooth ReLU variants like GELU and SiLU. These are self-gated activation functions where the range of the…

Neural and Evolutionary Computing · Computer Science 2024-06-03 Allen Hao Huang

In classification tasks, softmax functions are ubiquitously used as output activations to produce predictive probabilities. Such outputs only capture aleatoric uncertainty. To capture epistemic uncertainty, approximate Gaussian inference…

Machine Learning · Computer Science 2026-02-12 Bálint Mucsányi , Nathaël Da Costa , Philipp Hennig

Mixture of Experts (MoE) models constitute a widely utilized class of ensemble learning approaches in statistics and machine learning, known for their flexibility and computational efficiency. They have become integral components in…

Machine Learning · Statistics 2025-05-26 Tuan Thai , TrungTin Nguyen , Dat Do , Nhat Ho , Christopher Drovandi

Softmax routing approaches hard top-1 routing as the temperature tends to zero, but the limiting passage is singular at router ties. This paper develops a boundary-layer calculus for this soft-to-hard limit in population squared-loss…

Machine Learning · Computer Science 2026-05-26 Reza Rastegar

The Softmax loss is one of the most widely employed surrogate objectives for classification and ranking tasks. To elucidate its theoretical properties, the Fenchel-Young framework situates it as a canonical instance within a broad family of…

Machine Learning · Computer Science 2026-02-02 Yuanhao Pu , Defu Lian , Enhong Chen

Feature selection, dimension selection, and embedding compression are fundamental techniques for improving efficiency and generalization in deep recommender systems. Although conceptually related, these problems are typically studied in…

Machine Learning · Computer Science 2026-04-10 Yihong Huang , Chen Chu , Fan Zhang , Liping Wang Fei Chen , Yu Lin , Ruiduan Li , Zhihao Li

In many deployed systems (multilingual ASR, cross-hospital imaging, region-specific perception), multiple pretrained specialist models coexist. Yet, new target domains often require domain expansion: a generalized model that performs well…

Machine Learning · Computer Science 2026-01-30 Jong-Ik Park , Shreyas Chaudhari , Srinivasa Pranav , Carlee Joe-Wong , José M. F. Moura

Existing deep learning approaches for wearable fall detection systems rely on self-attention mechanisms that impose quadratic computational overhead, distributing weights across all time steps. This global weight distribution impairs the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Sana Alamgeer , Ronish Kumar , Awatif Yasmin , Muhammad Irshad , Anne H. H. Ngu

This study uses stacked generalization, which is a two-step process of combining machine learning methods, called meta or super learners, for improving the performance of algorithms in step one (by minimizing the error rate of each…

Machine Learning · Computer Science 2020-04-07 Kathleen Kerwin , Nathaniel D. Bastian

In today's landscape, Mixture of Experts (MoE) is a crucial architecture that has been used by many of the most advanced models. One of the major challenges of MoE models is that they usually require much more memory than their dense…

Machine Learning · Computer Science 2025-11-11 Shuning Lin , Yifan He , Yitong Chen

Intra-class compactness and inter-class separability are crucial indicators to measure the effectiveness of a model to produce discriminative features, where intra-class compactness indicates how close the features with the same label are…

Computer Vision and Pattern Recognition · Computer Science 2019-07-16 Yan Luo , Yongkang Wong , Mohan Kankanhalli , Qi Zhao

Low-light enhancement has wide applications in autonomous driving, 3D reconstruction, remote sensing, surveillance, and so on, which can significantly improve information utilization. However, most existing methods lack generalization and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Minwen Liao , Hao Bo Dong , Xinyi Wang , Kurban Ubul , Yihua Shao , Ziyang Yan

We provide a theoretical treatment of over-specified Gaussian mixtures of experts with covariate-free gating networks. We establish the convergence rates of the maximum likelihood estimation (MLE) for these models. Our proof technique is…

Statistics Theory · Mathematics 2022-03-09 Nhat Ho , Chiao-Yu Yang , Michael I. Jordan

Training large-scale generative models is resource-intensive and relies heavily on heuristic dataset weighting. We address two fundamental questions: Can we train Large Language Models (LLMs) modularly-combining small, domain-specific…

Machine Learning · Computer Science 2026-02-25 Corinna Cortes , Mehryar Mohri , Yutao Zhong

While transformers and their variant conformers show promising performance in speech recognition, the parameterized property leads to much memory cost during training and inference. Some works use cross-layer weight-sharing to reduce the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-20 Ye Bai , Jie Li , Wenjing Han , Hao Ni , Kaituo Xu , Zhuo Zhang , Cheng Yi , Xiaorui Wang

Linear attention methods offer a compelling alternative to softmax attention due to their efficiency in recurrent decoding. Recent research has focused on enhancing standard linear attention by incorporating gating while retaining its…

Machine Learning · Computer Science 2025-04-08 Yingcong Li , Davoud Ataee Tarzanagh , Ankit Singh Rawat , Maryam Fazel , Samet Oymak

While Transformer architecture excel at modeling long-range dependencies contributing to its widespread adoption in vision tasks the quadratic complexity of softmax-based attention mechanisms imposes a major bottleneck, particularly when…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yuan Cao , Dong Wang
‹ Prev 1 3 4 5 6 7 10 Next ›