中文
相关论文

相关论文: On Risk and Time Pressure: When to Think and When …

200 篇论文

Agents that plan and act in the real world must deal with the fact that time passes as they are planning. When timing is tight, there may be insufficient time to complete the search for a plan before it is time to act. By commencing…

人工智能 · 计算机科学 2023-03-07 Amihay Elboher , Ava Bensoussan , Erez Karpas , Wheeler Ruml , Shahaf S. Shperberg , Solomon E. Shimony

I model a rational agent who experiences endogenous deadline pressure in the face of a fixed future deadline. The agent holds a resource stock, and opportunities to spend resources arise randomly according to a Poisson process. When the…

理论经济学 · 经济学 2025-09-10 Conrad Kosowsky

In reinforcement learning, it is common to let an agent interact for a fixed amount of time with its environment before resetting it and repeating the process in a series of episodes. The task that the agent has to learn can either be to…

机器学习 · 计算机科学 2022-01-28 Fabio Pardo , Arash Tavakoli , Vitaly Levdik , Petar Kormushev

People often deviate from expected utility theory when making risky and intertemporal choices. While the effects of probabilistic risk and time delay have been extensively studied in isolation, their interplay and underlying theoretical…

理论经济学 · 经济学 2025-04-10 Ho Ka Chan , Taro Toyoizumi

Humans and animals have the ability to reason and make predictions about different courses of action at many time scales. In reinforcement learning, option models (Sutton, Precup \& Singh, 1999; Precup, 2000) provide the framework for this…

机器学习 · 计算机科学 2021-08-09 Khimya Khetarpal , Zafarali Ahmed , Gheorghe Comanici , Doina Precup

We discuss representing and reasoning with knowledge about the time-dependent utility of an agent's actions. Time-dependent utility plays a crucial role in the interaction between computation and action under bounded resources. We present a…

人工智能 · 计算机科学 2013-03-26 Eric J. Horvitz , Geoffrey Rutledge

Agents that learn to select optimal actions represent a prominent focus of the sequential decision-making literature. In the face of a complex environment or constraints on time and resources, however, aiming to synthesize such an optimal…

机器学习 · 计算机科学 2021-06-23 Dilip Arumugam , Benjamin Van Roy

An agent choosing between various actions tends to take the one with the lowest cost. But this choice is arguably too rigid (not adaptive) to be useful in complex situations, e.g., where exploration-exploitation trade-off is relevant in…

数据分析、统计与概率 · 物理学 2018-12-04 Armen E. Allahverdyan , Aram Galstyan , Ali E. Abbas , Zbigniew R. Struzik

Planning and reinforcement learning are two key approaches to sequential decision making. Multi-step approximate real-time dynamic programming, a recently successful algorithm class of which AlphaZero [Silver et al., 2018] is an example,…

人工智能 · 计算机科学 2020-05-18 Thomas M. Moerland , Anna Deichler , Simone Baldi , Joost Broekens , Catholijn M. Jonker

An agent acquires information dynamically until her belief about a binary state reaches an upper or lower threshold. She can choose any signal process subject to a constraint on the rate of entropy reduction. Strategies are ordered by "time…

理论经济学 · 经济学 2024-08-23 Daniel Chen , Weijie Zhong

We consider a dynamic moral hazard problem between a principal and an agent, where the sole instrument the principal has to incentivize the agent is the disclosure of information. The principal aims at maximizing the (discounted) number of…

理论经济学 · 经济学 2021-03-09 Wei Zhao , Claudio Mezzetti , Ludovic Renou , Tristan Tomala

Advanced reasoning models with agentic capabilities (AI agents) are deployed to interact with humans and to solve sequential decision-making problems under (approximate) utility functions and internal models. When such problems have…

We introduce a class of learning problems where the agent is presented with a series of tasks. Intuitively, if there is relation among those tasks, then the information gained during execution of one task has value for the execution of…

机器学习 · 计算机科学 2012-09-06 Christos Dimitrakakis

In a dynamic matching market, such as a marriage or job market, how should agents balance accepting a proposed match with the cost of continuing their search? We consider this problem in a discrete setting, in which agents have cardinal…

计算机科学与博弈论 · 计算机科学 2021-06-16 Ishan Agarwal , Richard Cole , Yixin Tao

Time-inconsistency is a characteristic of human behavior in which people plan for long-term benefits but take actions that differ from the plan due to conflicts with short-term benefits. Such time-inconsistent behavior is believed to be…

计算机科学与博弈论 · 计算机科学 2025-01-15 Yasunori Akagi , Naoki Marumo , Takeshi Kurashima

We seek to understand fundamental tradeoffs between the accuracy of prior information that a learner has on a given problem and its learning performance. We introduce the notion of prioritized risk, which differs from traditional notions of…

机器学习 · 计算机科学 2023-04-27 Anirudha Majumdar

We study the problem of option pricing and hedging strategies within the frame-work of risk-return arguments. An economic agent is described by a utility function that depends on profit (an expected value) and risk (a variance). In the…

统计力学 · 物理学 2008-12-02 Erik Aurell , Karol Życzkowski

We show that strategies implemented in automatic theorem proving involve an interesting tradeoff between execution speed, proving speedup/computational time and usefulness of information. We advance formal definitions for these concepts by…

计算机科学中的逻辑 · 计算机科学 2015-06-16 Santiago Hernández-Orozco , Francisco Hernández-Quiroz , Hector Zenil , Wilfried Sieg

We consider reinforcement learning with performance evaluated by a dynamic risk measure. We construct a projected risk-averse dynamic programming equation and study its properties. Then we propose risk-averse counterparts of the methods of…

最优化与控制 · 数学 2020-03-03 Umit Kose , Andrzej Ruszczynski

We consider the problem of reinforcement learning under safety requirements, in which an agent is trained to complete a given task, typically formalized as the maximization of a reward signal over time, while concurrently avoiding…

‹ 上一页 1 2 3 10 下一页 ›