English
Related papers

Related papers: FishBack: Pullback Fisher Geometry for Optimal Act…

200 papers

Imitation learning is a control design paradigm that seeks to learn a control policy reproducing demonstrations from expert agents. By substituting expert demonstrations for optimal behaviours, the same paradigm leads to the design of…

Machine Learning · Computer Science 2024-12-20 Dharmesh Tailor , Dario Izzo

Inverted landing is a challenging feat to perform in aerial robots, especially without external positioning. However, it is routinely performed by biological fliers such as bees, flies, and bats. Our previous observations of landing…

Robotics · Computer Science 2022-03-04 Bryan Habas , Bader AlAttar , Brian Davis , Jack W. Langelaan , Bo Cheng

Navigation is crucial for animal behavior and is assumed to require an internal representation of the external environment, termed a cognitive map. The precise form of this representation is often considered to be a metric representation of…

Neurons and Cognition · Quantitative Biology 2020-02-10 Tie Xu , Omri Barak

We propose a novel approach to input design for identification of nonlinear state space models. The optimal input sequence is obtained by maximizing a scalar cost function of the Fisher information matrix. Since the Fisher information…

Optimization and Control · Mathematics 2016-03-18 Patricio E. Valenzuela , Johan Dahlin , Cristian R. Rojas , Thomas B. Schön

Active soft bodies can affect their shape through an internal actuation mechanism that induces a deformation. Similar to recent work, this paper utilizes a differentiable, quasi-static, and physics-based simulation layer to optimize for…

Computer Vision and Pattern Recognition · Computer Science 2024-01-29 Lingchen Yang , Byungsoo Kim , Gaspard Zoss , Baran Gözcü , Markus Gross , Barbara Solenthaler

Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both steer the model and be readable along the unembedding. Function vectors (FVs), extracted as mean…

Machine Learning · Computer Science 2026-05-12 Mohammed Suhail B Nadaf

Riemannian optimization uses local methods to solve optimization problems whose constraint set is a smooth manifold. A linear step along some descent direction usually leaves the constraints, and hence retraction maps are used to…

Statistics Theory · Mathematics 2023-01-19 Alexander Heaton , Matthias Himmelmann

The Fisher information matrix (FIM) is fundamental to understanding the trainability of deep neural nets (DNN), since it describes the parameter space's local metric. We investigate the spectral distribution of the conditional FIM, which is…

Machine Learning · Statistics 2021-03-31 Tomohiro Hayase , Ryo Karakida

Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficient adaptation, and prompting. In parallel, a growing body of work modifies internal…

Computation and Language · Computer Science 2026-04-16 Simon Ostermann , Daniil Gurgurov , Tanja Baeumel , Michael A. Hedderich , Sebastian Lapuschkin , Wojciech Samek , Vera Schmitt

We introduce a new optimization procedure for Euclidean path integrals which compute wave functionals in conformal field theories (CFTs). We optimize the background metric in the space on which the path integration is performed.…

High Energy Physics - Theory · Physics 2017-08-30 Pawel Caputa , Nilay Kundu , Masamichi Miyaji , Tadashi Takayanagi , Kento Watanabe

Large language models sometimes assert falsehoods despite internally representing the correct answer, failures of honesty rather than accuracy, which undermines auditability and safety. Existing approaches largely optimize factual…

Machine Learning · Computer Science 2025-12-09 Gracjan Góral , Marysia Winkels , Steven Basart

Optimization in large language models (LLMs) unfolds over high-dimensional parameter spaces with non-Euclidean structure. Information geometry frames this landscape using the Fisher information metric, enabling more principled learning via…

Computation and Language · Computer Science 2025-12-09 Riccardo Di Sipio

The selection problem of an optimal set of sensors estimating the snapshot of high-dimensional data is considered. The objective functions based on various criteria of optimal design are adopted to the greedy method: D-optimality,…

Signal Processing · Electrical Eng. & Systems 2021-04-09 Kumi Nakai , Keigo Yamada , Takayuki Nagata , Yuji Saito , Taku Nonomura

Expert specialization is fundamental to Mixture-of-Experts (MoE) model success, yet existing metrics (cosine similarity, routing entropy) lack theoretical grounding and yield inconsistent conclusions under reparameterization. We present an…

Artificial Intelligence · Computer Science 2026-04-17 Dongxin Guo , Jikun Wu , Siu Ming Yiu

We consider the problems of clustering, classification, and visualization of high-dimensional data when no straightforward Euclidean representation exists. Typically, these tasks are performed by first reducing the high-dimensional data to…

Machine Learning · Statistics 2009-09-29 Kevin M. Carter , Raviv Raich , William G. Finn , Alfred O. Hero

Probabilistic Latent Variable Models (LVMs) excel at modeling complex, high-dimensional data through lower-dimensional representations. Recent advances show that equipping these latent representations with a Riemannian metric unlocks…

Machine Learning · Computer Science 2025-05-20 Luis Augenstein , Noémie Jaquier , Tamim Asfour , Leonel Rozo

Transformers evaluated in a single, fixed-depth pass are provably limited in expressive power to the constant-depth circuit class TC0. Running a Transformer autoregressively removes that ceiling -- first in next-token prediction and, more…

Machine Learning · Computer Science 2025-07-21 Mrinal Mathur , Mike Doan , Barak Pearlmutter , Sergey Plis

Deep autoregressive models are one of the most powerful models that exist today which achieve state-of-the-art bits per dim. However, they lie at a strict disadvantage when it comes to controlled sample generation compared to latent…

Computer Vision and Pattern Recognition · Computer Science 2020-05-12 Wilson Yan , Jonathan Ho , Pieter Abbeel

This paper presents state estimation and stochastic optimal control gathered in one global optimization problem generating dual effect i.e. the control can improve the future estimation. As the optimal policy is impossible to compute, a…

Optimization and Control · Mathematics 2023-03-27 Emilien Flayac , Karim Dahia , Bruno Hérissé , Frédéric Jean

Deep Reinforcement Learning (DRL) systems often tend to overfit to early experiences, a phenomenon known as the primacy bias (PB). This bias can severely hinder learning efficiency and final performance, particularly in complex…

Machine Learning · Computer Science 2025-02-04 Massimiliano Falzari , Matthia Sabatelli