English
Related papers

Related papers: Target Weight Mechanism doesn't make delta hedge e…

200 papers

We study a natural generalization of the maximum weight many-to-one matching problem. We are given an undirected bipartite graph $G= (A \cup P, E)$ with weights on the edges in $E$, and with lower and upper quotas on the vertices in $P$. We…

Discrete Mathematics · Computer Science 2016-03-29 Ashwin Arulselvan , Ágnes Cseh , Martin Groß , David F. Manlove , Jannik Matuschke

Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies…

Computation and Language · Computer Science 2024-11-27 Bowen Ping , Shuo Wang , Hanqing Wang , Xu Han , Yuzhuang Xu , Yukun Yan , Yun Chen , Baobao Chang , Zhiyuan Liu , Maosong Sun

The Primal-Dual (PD) algorithm is widely used in convex optimization to determine saddle points. While the stability of the PD algorithm can be easily guaranteed, strict contraction is nontrivial to establish in most cases. This work…

Optimization and Control · Mathematics 2018-11-21 Hung D. Nguyen , Thanh Long Vu , Konstantin Turitsyn , Jean-Jacques Slotine

Serving many task-specialized LLM variants is often limited by the large size of fine-tuned checkpoints and the resulting cold-start latency. Since fine-tuned weights differ from their base model by relatively small structured residuals, a…

Machine Learning · Computer Science 2025-12-24 Stefan Kuyumdzhiev , Radostin Cholakov

In this paper, we tackle the challenging problem of delayed rewards in reinforcement learning (RL). While Proximal Policy Optimization (PPO) has emerged as a leading Policy Gradient method, its performance can degrade under delayed rewards.…

Mechanistic interpretability seeks to localize model behavior to the internal components that causally realize it. Prior work has advanced activation-space localization and causal tracing, but modules that appear important in activation…

Artificial Intelligence · Computer Science 2026-04-16 Chenghao Sun , Chengsheng Zhang , Guanzheng Qin , Rui Dai , Xinmei Tian

Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise is challenged by recent benchmarks where simple majority…

Computation and Language · Computer Science 2026-04-24 Peng Kuang , Yanli Wang , Xiaoyu Han , Yaowenqi Liu , Kaidi Xu , Haohan Wang

Optimizing credit limits, interest rates, and loan terms is crucial for managing borrower risk and lifetime value (LTV) in personal loan platform. However, counterfactual estimation of these continuous, multi-dimensional treatments faces…

Machine Learning · Computer Science 2025-08-12 Kexin Zhao , Bo Wang , Cuiying Zhao , Tongyao Wan

In reinforcement learning, the discount factor $\gamma$ controls the agent's effective planning horizon. Traditionally, this parameter was considered part of the MDP; however, as deep reinforcement learning algorithms tend to become…

Machine Learning · Computer Science 2020-06-24 Chen Tessler , Shie Mannor

Learning the value function of a given policy from data samples is an important problem in Reinforcement Learning. TD($\lambda$) is a popular class of algorithms to solve this problem. However, the weights assigned to different $n$-step…

Machine Learning · Computer Science 2021-11-24 Rohan Deb , Meet Gandhi , Shalabh Bhatnagar

In the total matching problem, one is given a graph $G$ with weights on the vertices and edges. The goal is to find a maximum weight set of vertices and edges that is the non-incident union of a stable set and a matching. We consider the…

Combinatorics · Mathematics 2024-01-01 Luca Ferrarini , Samuel Fiorini , Stefan Kober , Yelena Yuditsky

Delta hedging, which plays a crucial r\^ole in modern financial engineering, is a tracking control design for a "risk-free" management. We utilize the existence of trends in financial time series (Fliess M., Join C.: A mathematical proof of…

Pricing of Securities · Quantitative Finance 2010-05-31 Michel Fliess , Cédric Join

It is well known that mirror descent may diverge or cycle on merely monotone variational inequalities. In this paper, we propose \emph{Target Mirror Descent} (TMD), a unified framework that stabilizes monotone flows via a target point…

Optimization and Control · Mathematics 2026-04-22 Yu-Wen Chen , Can Kizilkale , Murat Arcak

Using a physically motivated stress energy tensor, we prove weak and strong monotonicity formulas for solutions to the semilinear elliptic system $\Delta u=\nabla W(u)$ with $W$ nonnegative. In particular, we extend a recent two dimensional…

Analysis of PDEs · Mathematics 2014-02-27 Christos Sourdis

We consider Targeted Maximum Likelihood Estimation (TMLE) of weighted average treatment effects (WATEs), a class of causal estimands that reweight the covariate distribution using a specified function of the propensity score. This class…

Statistics Theory · Mathematics 2026-04-02 Yang Liu , Patrick Lopatto , Ivana Malenica

In the dynamic linear program (LP) problem, we are given an LP undergoing updates and we need to maintain an approximately optimal solution. Recently, significant attention (e.g., [Gupta et al. STOC'17; Arar et al. ICALP'18, Wajc STOC'20])…

Data Structures and Algorithms · Computer Science 2022-07-18 Sayan Bhattacharya , Peter Kiss , Thatchaphol Saranurak

DeepLLL algorithm (Schnorr, 1994) is a famous variant of LLL lattice basis reduction algorithm, and PotLLL algorithm (Fontein et al., 2014) and $S^2$LLL algorithm (Yasuda and Yamaguchi, 2019) are recent polynomial-time variants of DeepLLL…

Data Structures and Algorithms · Computer Science 2021-06-02 Takuto Odagawa , Koji Nuida

Developing meta-learning algorithms that are un-biased toward a subset of training tasks often requires hand-designed criteria to weight tasks, potentially resulting in sub-optimal solutions. In this paper, we introduce a new principled and…

Machine Learning · Computer Science 2023-01-05 Cuong Nguyen , Thanh-Toan Do , Gustavo Carneiro

Weight averaging has become a standard technique for enhancing model performance. However, methods such as Stochastic Weight Averaging (SWA) and Latest Weight Averaging (LAWA) often require manually designed procedures to sample from the…

Machine Learning · Computer Science 2025-02-17 Peng Wang , Shengchao Hu , Zerui Tao , Guoxia Wang , Dianhai Yu , Li Shen , Quan Zheng , Dacheng Tao

Portfolio management via reinforcement learning is at the forefront of fintech research, which explores how to optimally reallocate a fund into different financial assets over the long term by trial-and-error. Existing methods are…

Artificial Intelligence · Computer Science 2021-02-09 Rundong Wang , Hongxin Wei , Bo An , Zhouyan Feng , Jun Yao
‹ Prev 1 3 4 5 6 7 10 Next ›