English
Related papers

Related papers: Sparse Transformer Architectures via Regularized W…

200 papers

Optimal Transport (OT) has attracted significant interest in the machine learning community, not only for its ability to define meaningful distances between probability distributions -- such as the Wasserstein distance -- but also for its…

Machine Learning · Computer Science 2025-11-04 Laetitia Chapel , Romain Tavenard , Samuel Vaiter

Weight pruning is among the most popular approaches for compressing deep convolutional neural networks. Recent work suggests that in a randomly initialized deep neural network, there exist sparse subnetworks that achieve performance…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Vinay Kumar Verma , Nikhil Mehta , Shijing Si , Ricardo Henao , Lawrence Carin

Sampling from high-dimensional distributions is a fundamental problem in statistical research and practice. However, great challenges emerge when the target density function is unnormalized and contains isolated modes. We tackle this…

Methodology · Statistics 2023-04-11 Yixuan Qiu , Xiao Wang

Compressed sensing theory is slowly making its way to solve more and more astronomical inverse problems. We address here the application of sparse representations, convex optimization and proximal theory to radio interferometric imaging.…

Instrumentation and Methods for Astrophysics · Physics 2015-08-28 Julien N. Girard , Hugh Garsden , Jean Luc Starck , Stéphane Corbel , Arnaud Woiselle , Cyril Tasse , John P. McKean , Jérôme Bobin

We investigate overdetermined linear inverse problems for which the forward operator may not be given accurately. We introduce a new tool called the structure, based on the Wasserstein distance, and propose the use of this to diagnose and…

Numerical Analysis · Mathematics 2018-11-01 Michael A. Puthawala , Cory D. Hauck , Stanley J. Osher

Deep neural networks have become very popular in modeling complex nonlinear processes due to their extraordinary ability to fit arbitrary nonlinear functions from data with minimal expert intervention. However, they are almost always…

Chemical Physics · Physics 2023-01-16 Erlend Torje Berg Lundby , Adil Rasheed , Ivar Johan Halvorsen , Jan Tommy Gravdahl

We introduce a prior for the parameters of univariate continuous distributions, based on the Wasserstein information matrix, which is invariant under reparameterisations. We discuss the links between the proposed prior with information…

Statistics Theory · Mathematics 2022-07-28 W. Li , F. J. Rubio

The computation of Wasserstein gradient direction is essential for posterior sampling problems and scientific computing. The approximation of the Wasserstein gradient with finite samples requires solving a variational problem. We study the…

Machine Learning · Computer Science 2022-05-27 Yifei Wang , Peng Chen , Mert Pilanci , Wuchen Li

This paper introduces the sparsifying preconditioner for the pseudospectral approximation of highly indefinite systems on periodic structures, which include the frequency-domain response problems of the Helmholtz equation and the…

Numerical Analysis · Mathematics 2014-09-18 Lexing Ying

This thesis examines self-attention training through the lens of Optimal Transport (OT) and develops an OT-based alternative for tabular classification. The study tracks intermediate projections of the self-attention layer during training…

Machine Learning · Statistics 2026-02-19 Alessandro Quadrio , Antonio Candelieri

Flexible Bayesian models are typically constructed using limits of large parametric models with a multitude of parameters that are often uninterpretable. In this article, we offer a novel alternative by constructing an exponentially tilted…

Methodology · Statistics 2023-03-20 Abhisek Chakraborty , Anirban Bhattacharya , Debdeep Pati

This study proposes a new discrete neural operator for surrogate modeling of transient Darcy flow fields in heterogeneous porous media with random parameters. The new method integrates temporal encoding, operator learning and UNet to…

Numerical Analysis · Mathematics 2025-12-04 Zhenglong Chen , Zhao Zhang , Xia Yan , Jiayu Zhai , Piyang Liu , Kai Zhang

In-network distributed estimation of sparse parameter vectors via diffusion LMS strategies has been studied and investigated in recent years. In all the existing works, some convex regularization approach has been used at each node of the…

Machine Learning · Computer Science 2016-11-15 Bijit Kumar Das , Mrityunjoy Chakraborty , Jerónimo Arenas-García

We address the non-convex optimisation problem of finding a sparse matrix on the Stiefel manifold (matrices with mutually orthogonal columns of unit length) that maximises (or minimises) a quadratic objective function. Optimisation problems…

Optimization and Control · Mathematics 2021-10-04 Florian Bernard , Daniel Cremers , Johan Thunberg

The traffic assignment problem is essential for traffic flow analysis, traditionally solved using mathematical programs under the Equilibrium principle. These methods become computationally prohibitive for large-scale networks due to…

Machine Learning · Computer Science 2026-04-28 Mostafa Ameli , Sulthana Shams , Van Anh Le , Alexander Skabardonis

Covariate shift arises when covariate distributions differ between source and target populations while the conditional distribution of the response remains invariant, and it underlies problems in missing data and causal inference. We…

Methodology · Statistics 2026-01-13 Junjun Lang , Qiong Zhang , Yukun Liu

Model Updating is frequently used in Structural Health Monitoring to determine structures' operating conditions and whether maintenance is required. Data collected by sensors are used to update the values of some initially unknown…

Computation · Statistics 2024-01-23 Felipe Igea , Alice Cicirello

For downlink transmission in massive multi-user multiple-input multiple-output (MU-MIMO) systems, conventional precoding research heavily focuses on reducing the computational complexity of precoding matrix design, while largely overlooking…

Signal Processing · Electrical Eng. & Systems 2026-05-19 Shuai Gao , Fan Xu , Mian Li , Xinzhi Ning , Lei Qiu , Ye Yang , Qingjiang Shi

In this work, we focus on variational Bayesian inference on the sparse Deep Neural Network (DNN) modeled under a class of spike-and-slab priors. Given a pre-specified sparse DNN structure, the corresponding variational posterior contraction…

Statistics Theory · Mathematics 2020-08-04 Jincheng Bai , Qifan Song , Guang Cheng

Optimal transportation theory and the related $p$-Wasserstein distance ($W_p$, $p\geq 1$) are widely-applied in statistics and machine learning. In spite of their popularity, inference based on these tools has some issues. For instance, it…

Statistics Theory · Mathematics 2024-03-01 Yiming Ma , Hang Liu , Davide La Vecchia , Metthieu Lerasle
‹ Prev 1 4 5 6 7 8 10 Next ›