English
Related papers

Related papers: Target Weight Mechanism doesn't make delta hedge e…

200 papers

Reinforcement learning algorithms typically assume rewards to be sampled from light-tailed distributions, such as Gaussian or bounded. However, a wide variety of real-world systems generate rewards that follow heavy-tailed distributions. We…

Machine Learning · Computer Science 2021-02-26 Vincent Zhuang , Yanan Sui

While large-scale training data is fundamental for developing capable large language models (LLMs), strategically selecting high-quality data has emerged as a critical approach to enhance training efficiency and reduce computational costs.…

Machine Learning · Computer Science 2025-07-23 Yang Yu , Kai Han , Hang Zhou , Yehui Tang , Kaiqi Huang , Yunhe Wang , Dacheng Tao

This paper explores continuous-time control synthesis for target-driven navigation to satisfy complex high-level tasks expressed as linear temporal logic (LTL). We propose a model-free framework using deep reinforcement learning (DRL) where…

Robotics · Computer Science 2023-03-17 Mingyu Cai , Makai Mann , Zachary Serlin , Kevin Leahy , Cristian-Ioan Vasile

Learning to control an unknown dynamical system with respect to high-level temporal specifications is an important problem in control theory. We present the first regret-free online algorithm for learning a controller for linear temporal…

Artificial Intelligence · Computer Science 2025-06-09 Rupak Majumdar , Mahmoud Salamati , Sadegh Soudjani

Learning from human preference is a paradigm used in large-scale language model (LLM) fine-tuning step to better align pretrained LLM to human preference for downstream task. In the past it uses reinforcement learning from human feedback…

Artificial Intelligence · Computer Science 2024-09-02 Shiming Xie , Hong Chen , Fred Yu , Zeye Sun , Xiuyu Wu , Yingfan Hu

Dynamic pricing strategies are crucial for firms to maximize revenue by adjusting prices based on market conditions and customer characteristics. However, designing optimal pricing strategies becomes challenging when historical data are…

Machine Learning · Computer Science 2025-02-03 Fan Wang , Feiyu Jiang , Zifeng Zhao , Yi Yu

The question of pricing and hedging a given contingent claim has a unique solution in a complete market framework. When some incompleteness is introduced, the problem becomes however more difficult. Several approaches have been adopted in…

Probability · Mathematics 2007-08-08 Pauline Barrieu , Nicole El Karoui

Options are contingent claims regarding the value of underlying assets. The Black-Scholes formula provides a road map for pricing these options in a risk-neutral setting, justified by a delta hedging argument in which countervailing…

Mathematical Finance · Quantitative Finance 2026-05-26 Erina Nanyonga , Matt Davison

Reinforcement Learning (RL) has emerged as an efficient method of choice for solving complex sequential decision making problems in automatic control, computer science, economics, and biology. In this paper we present a model-free RL…

Logic in Computer Science · Computer Science 2019-09-13 Mohammadhosein Hasanbeig , Yiannis Kantaros , Alessandro Abate , Daniel Kroening , George J. Pappas , Insup Lee

State-of-the-art efficient model-based Reinforcement Learning (RL) algorithms typically act by iteratively solving empirical models, i.e., by performing \emph{full-planning} on Markov Decision Processes (MDPs) built by the gathered…

Machine Learning · Computer Science 2019-11-01 Yonathan Efroni , Nadav Merlis , Mohammad Ghavamzadeh , Shie Mannor

We present a logic that extends CTL (Computation Tree Logic) with operators that express synchronization properties. A property is synchronized in a system if it holds in all paths of a certain length. The new logic is obtained by using the…

Logic in Computer Science · Computer Science 2016-05-25 Krishnendu Chatterjee , Laurent Doyen

Portfolio management remains a crucial challenge in finance, with traditional methods often falling short in complex and volatile market environments. While deep reinforcement approaches have shown promise, they still face limitations in…

Machine Learning · Computer Science 2025-03-07 Fengchen Gu , Zhengyong Jiang , Ángel F. García-Fernández , Angelos Stefanidis , Jionglong Su , Huakang Li

In this paper we discuss the notion of reducibility for matrix weights and introduce a real vector space $\mathcal C_\mathbb{R}$ which encodes all information about the reducibility of $W$. In particular a weight $W$ reduces if and only if…

Representation Theory · Mathematics 2016-11-02 Juan Tirao , Ignacio Zurrián

The strong Lottery Ticket Hypothesis (LTH) claims the existence of a subnetwork in a sufficiently large, randomly initialized neural network that approximates some target neural network without the need of training. We extend the…

Machine Learning · Computer Science 2022-11-01 Zheyang Xiong , Fangshuo Liao , Anastasios Kyrillidis

When traditional pole-dynamics attacks (TPDAs) are implemented with nominal models, model mismatch between exact and nominal models often affects their stealthiness, or even makes the stealthiness lost. To solve this problem, our current…

Systems and Control · Electrical Eng. & Systems 2022-10-31 Dajun Du , Changda Zhang , Chen Peng , Minrui Fei , Huiyu Zhou

We tackle the problem of online optimization with a general, possibly unbounded, loss function. It is well known that when the loss is bounded, the exponentially weighted aggregation strategy (EWA) leads to a regret in $\sqrt{T}$ after $T$…

Machine Learning · Statistics 2022-01-17 Pierre Alquier

Traditional portfolio management methods can incorporate specific investor preferences but rely on accurate forecasts of asset returns and covariances. Reinforcement learning (RL) methods do not rely on these explicit forecasts and are…

Portfolio Management · Quantitative Finance 2022-03-23 Ruan Pretorius , Terence van Zyl

We study the convergence behavior of the celebrated temporal-difference (TD) learning algorithm. By looking at the algorithm through the lens of optimization, we first argue that TD can be viewed as an iterative optimization algorithm where…

Machine Learning · Computer Science 2023-11-10 Kavosh Asadi , Shoham Sabach , Yao Liu , Omer Gottesman , Rasool Fakoor

The basic concept of multi-dimensional limiting process (MLP) on unstructured grids is inherited and modified for improving shock stabilities and reducing numerical dissipation on smooth regions. A relaxed version of MLP condition, simply…

Numerical Analysis · Mathematics 2017-12-07 Fan Zhang , Jun Liu , Biaosong Chen

Parker's formulation of isotopological plasma relaxation process in magnetohydrodynamics (MHD) is extended to Hall MHD. The torsion coefficient alpha in the Hall MHD Beltrami condition turns out now to be proportional to the "potential…

Plasma Physics · Physics 2015-06-03 B. K. Shivamoggi