中文
相关论文

相关论文: Game and Reference: Policy Combination Synthesis f…

200 篇论文

Transfer learning can greatly speed up reinforcement learning for a new task by leveraging policies of relevant tasks. Existing works of policy reuse either focus on only selecting a single best source policy for transfer without…

人工智能 · 计算机科学 2019-03-11 Siyuan Li , Fangda Gu , Guangxiang Zhu , Chongjie Zhang

World models have emerged as a unifying paradigm for learning latent dynamics, simulating counterfactual futures, and supporting planning under uncertainty. In this paper, we argue that computational epidemiology is a natural and…

机器学习 · 计算机科学 2026-04-14 Zeeshan Memon , Yiqi Su , Christo Kurisummoottil Thomas , Walid Saad , Liang Zhao , Naren Ramakrishnan

This paper is concerned with the design of control policies from example datasets. The case considered is when just a black box description of the system to be controlled is available and the system is affected by actuation constraints.…

最优化与控制 · 数学 2022-01-11 Davide Gagliardi , Giovanni Russo

We introduce a policy model coupled with the susceptible-infected-recovered (SIR) epidemic model to study interactions between policy-making and the dynamics of epidemics. We consider both single-region policies, as well as game-theoretic…

最优化与控制 · 数学 2024-08-06 Xia Li , Andrea L. Bertozzi , P. Jeffrey Brantingham , Yevgeniy Vorobeychik

In this work we present a framework which may transform research and praxis in epidemic planning. Introduced in the context of the ongoing COVID-19 pandemic, we provide a concrete demonstration of the way algorithms may learn from…

机器学习 · 计算机科学 2022-10-06 Sekou L. Remy , Oliver E. Bent

Accurate epidemic forecasting is crucial for outbreak preparedness, but existing data-driven models are often brittle. Typically trained on a single pathogen, they struggle with data scarcity during new outbreaks and fail under distribution…

机器学习 · 计算机科学 2026-02-25 Zewen Liu , Juntong Ni , Bohan Wang , Max S. Y. Lau , Wei Jin

Modern studies of societal phenomena rely on the availability of large datasets capturing attributes and activities of synthetic, city-level, populations. For instance, in epidemiology, synthetic population datasets are necessary to study…

数据库 · 计算机科学 2016-02-26 Hao Wu , Yue Ning , Prithwish Chakraborty , Jilles Vreeken , Nikolaj Tatti , Naren Ramakrishnan

Epidemiologists model the dynamics of epidemics in order to propose control strategies based on pharmaceutical and non-pharmaceutical interventions (contact limitation, lock down, vaccination, etc). Hand-designing such strategies is not…

Epidemics of infectious diseases posing a serious risk to human health have occurred throughout history. During recent epidemics there has been much debate about policy, including how and when to impose restrictions on behaviour.…

理论经济学 · 经济学 2024-04-08 Simon K. Schnyder , John J. Molina , Ryoichi Yamamoto , Matthew S. Turner

The beneficial effects of treatments vary across individuals in most studies. Treatment heterogeneity motivates practitioners to search for the optimal policy based on personal characteristics. A long-standing common practice in policy…

统计理论 · 数学 2025-01-06 Xuqiao Li , Ying Yan

Predictive models are often introduced to decision-making tasks under the rationale that they improve performance over an existing decision-making policy. However, it is challenging to compare predictive performance against an existing…

机器学习 · 计算机科学 2024-06-13 Luke Guerdan , Amanda Coston , Kenneth Holstein , Zhiwei Steven Wu

Approximating model predictive control (MPC) policy using expert-based supervised learning techniques requires labeled training data sets sampled from the MPC policy. This is typically obtained by sampling the feasible state-space and…

最优化与控制 · 数学 2022-03-16 Dinesh Krishnamoorthy

Infectious disease outbreaks can have a disruptive impact on public health and societal processes. As decision making in the context of epidemic mitigation is hard, reinforcement learning provides a methodology to automatically learn…

When deploying artificial agents in real-world environments where they interact with humans, it is crucial that their behavior is aligned with the values, social norms or other requirements of that environment. However, many environments…

机器学习 · 计算机科学 2023-05-05 Mattijs Baert , Pietro Mazzaglia , Sam Leroux , Pieter Simoens

The problem of retrosynthetic planning can be framed as one player game, in which the chemist (or a computer program) works backwards from a molecular target to simpler starting materials though a series of choices regarding which reactions…

机器学习 · 计算机科学 2019-01-23 John S. Schreck , Connor W. Coley , Kyle J. M. Bishop

This paper proposes a feedback design that effectively copes with uncertainties for reliable epidemic monitoring and control. There are several optimization-based methods to estimate the parameters of an epidemic model by utilizing past…

最优化与控制 · 数学 2023-04-06 Muhammad Umar B. Niazi , Philip E. Paré , Karl H. Johansson

Model Predictive Control has been recently proposed as policy approximation for Reinforcement Learning, offering a path towards safe and explainable Reinforcement Learning. This approach has been investigated for Q-learning and actor-critic…

系统与控制 · 电气工程与系统科学 2020-04-06 Sebastien Gros , Mario Zanon

In the context of epidemiology, policies for disease control are often devised through a mixture of intuition and brute-force, whereby the set of logically conceivable policies is narrowed down to a small family described by a few…

种群与进化 · 定量生物学 2021-10-04 Miguel Navascues , Costantino Budroni , Yelena Guryanova

Modeling policies for sequential clinical decision-making based on observational data is useful for describing treatment practices, standardizing frequent patterns in treatment, and evaluating alternative policies. For each task, it is…

AI agents are increasingly deployed as quasi-autonomous systems for specialized tasks, yet their potential as computational models of decision-making remains underexplored. We develop a generative AI agent to study repetitive policy…

多智能体系统 · 计算机科学 2026-01-09 Goshi Aoki , Navid Ghaffarzadegan