English
Related papers

Related papers: A Measure-Theoretic Analysis of Reasoning: Structu…

200 papers

Recent work on neural algorithmic reasoning has investigated the reasoning capabilities of neural networks, effectively demonstrating they can learn to execute classical algorithms on unseen data coming from the train distribution. However,…

The advent of the Transformer has led to the development of large language models (LLM), which appear to demonstrate human-like capabilities. To assess the generality of this class of models and a variety of other base neural network…

Machine Learning · Computer Science 2024-03-05 Takuya Ito , Soham Dan , Mattia Rigotti , James Kozloski , Murray Campbell

In this essay, we discuss the notion of optimal transport on geodesic measure spaces and the associated (2-)Wasserstein distance. We then examine displacement convexity of the entropy functional on the space of probability measures. In…

Metric Geometry · Mathematics 2012-04-17 Otis Chodosh

In recent years, Wasserstein Distributionally Robust Optimization (DRO) has garnered substantial interest for its efficacy in data-driven decision-making under distributional uncertainty. However, limited research has explored the…

Machine Learning · Computer Science 2025-10-01 Ahmad-Reza Ehyaei , Golnoosh Farnadi , Samira Samadi

Bubeck and Sellke (2021) pose as an open problem the connection between the law of robustness and robust generalization. The law of robustness states that overparameterization is necessary for models to interpolate robustly; in particular,…

Machine Learning · Computer Science 2026-02-26 Himadri Mandal , Vishnu Varadarajan , Jaee Ponde , Aritra Das , Mihir More , Debayan Gupta

Autoregressive (AR) language models enforce a fixed left-to-right generation order, creating a fundamental limitation when the required output structure conflicts with natural reasoning (e.g., producing answers before explanations due to…

Computation and Language · Computer Science 2026-01-30 Longxuan Yu , Yu Fu , Shaorong Zhang , Hui Liu , Mukund Varma T , Greg Ver Steeg , Yue Dong

Unsupervised learning aims to capture the underlying structure of potentially large and high-dimensional datasets. Traditionally, this involves using dimensionality reduction (DR) methods to project data onto lower-dimensional spaces or…

Machine Learning · Computer Science 2025-06-30 Hugues Van Assel , Cédric Vincent-Cuaz , Nicolas Courty , Rémi Flamary , Pascal Frossard , Titouan Vayer

Diffusion generative models have emerged as powerful tools for producing synthetic data from an empirically observed distribution. A common approach involves simulating the time-reversal of an Ornstein-Uhlenbeck (OU) process initialized at…

Machine Learning · Statistics 2025-12-02 Valentin de Bortoli , Romuald Elie , Anna Kazeykina , Zhenjie Ren , Jiacheng Zhang

Many variants of Optimal Transport (OT) have been developed to address its heavy computation. Among them, notably, Sliced Wasserstein (SW) is widely used for application domains by projecting the OT problem onto one-dimensional lines, and…

Machine Learning · Computer Science 2025-06-10 Viet-Hoang Tran , Trang Pham , Tho Tran , Minh Khoi Nguyen Nhat , Thanh Chu , Tam Le , Tan M. Nguyen

Matching a source to a target probability measure is often solved by instantiating a linear optimal transport (OT) problem, parameterized by a ground cost function that quantifies discrepancy between points. When these measures live in the…

Machine Learning · Computer Science 2023-11-27 Othmane Sebbouh , Marco Cuturi , Gabriel Peyré

Suppose we are given two metric spaces and a family of continuous transformations from one to the other. Given a probability distribution on each of these two spaces - namely the source and the target measures - the Wasserstein alignment…

Probability · Mathematics 2025-03-11 Soumik Pal , Bodhisattva Sen , Ting-Kam Leonard Wong

Large language models (LLMs) have achieved remarkable proficiency on solving diverse problems. However, their generalization ability is not always satisfying and the generalization problem is common for generative transformer models in…

Machine Learning · Computer Science 2024-08-20 Xingcheng Xu , Zihao Pan , Haipeng Zhang , Yanqing Yang

Though Transformers have achieved promising results in many computer vision tasks, they tend to be over-confident in predictions, as the standard Dot Product Self-Attention (DPSA) can barely preserve distance for the unbounded input domain.…

Machine Learning · Computer Science 2023-07-19 Wenqian Ye , Yunsheng Ma , Xu Cao , Kun Tang

We propose REpresentation-Aware Distributionally Robust Estimation (READ), a novel framework for Wasserstein distributionally robust learning that accounts for predictive representations when guarding against distributional shifts. Unlike…

Methodology · Statistics 2025-09-12 Zitao Wang , Nian Si , Molei Liu

Tokenization is the first - and often underappreciated - layer of computation in language models. While Chain-of-Thought (CoT) prompting enables transformer models to approximate recurrent computation by externalizing intermediate steps, we…

Computation and Language · Computer Science 2025-05-21 Xiang Zhang , Juntai Cao , Jiaqi Wei , Yiwei Xu , Chenyu You

We study the problem of resource provisioning under stringent reliability or service level requirements, which arise in applications such as power distribution, emergency response, cloud server provisioning, and regulatory risk management.…

Optimization and Control · Mathematics 2025-04-11 Anand Deo , Karthyek Murthy

Dimension reduction (DR) methods provide systematic approaches for analyzing high-dimensional data. A key requirement for DR is to incorporate global dependencies among original and embedded samples while preserving clusters in the…

Machine Learning · Statistics 2023-03-10 Antoine Collas , Titouan Vayer , Rémi Flamary , Arnaud Breloy

This paper advances the theory and practice of Domain Generalization (DG) in machine learning. We consider the typical DG setting where the hypothesis is composed of a representation mapping followed by a labeling function. Within this…

Machine Learning · Computer Science 2023-05-23 Boyang Lyu , Thuan Nguyen , Prakash Ishwar , Matthias Scheutz , Shuchin Aeron

Adversarial examples have pointed out Deep Neural Networks vulnerability to small local noise. It has been shown that constraining their Lipschitz constant should enhance robustness, but make them harder to learn with classical loss…

Generalization to out-of-distribution (OOD) data is a capability natural to humans yet challenging for machines to reproduce. This is because most learning algorithms strongly rely on the i.i.d.~assumption on source/target data, which is…

Machine Learning · Computer Science 2022-08-15 Kaiyang Zhou , Ziwei Liu , Yu Qiao , Tao Xiang , Chen Change Loy
‹ Prev 1 3 4 5 6 7 10 Next ›