中文
相关论文

相关论文: Demystifying MuZero Planning: Interpreting the Lea…

200 篇论文

Learning or identifying dynamics from a sequence of high-dimensional observations is a difficult challenge in many domains, including reinforcement learning and control. The problem has recently been studied from a generative perspective…

机器人学 · 计算机科学 2022-07-12 Oliver Limoyo , Bryan Chan , Filip Marić , Brandon Wagstaff , Rupam Mahmood , Jonathan Kelly

Learning efficiently from small amounts of data has long been the focus of model-based reinforcement learning, both for the online case when interacting with the environment and the offline case when learning from a fixed dataset. However,…

We examine an important setting for engineered systems in which low-power distributed sensors are each making highly noisy measurements of some unknown target function. A center wants to accurately learn this function by querying a small…

机器学习 · 计算机科学 2014-06-26 Maria-Florina Balcan , Chris Berlind , Avrim Blum , Emma Cohen , Kaushik Patnaik , Le Song

Distributionally robust optimization (DRO) provides a framework for training machine learning models that are able to perform well on a collection of related data distributions (the "uncertainty set"). This is done by solving a min-max…

机器学习 · 计算机科学 2021-04-01 Paul Michel , Tatsunori Hashimoto , Graham Neubig

Planning has been very successful for control tasks with known environment dynamics. To leverage planning in unknown environments, the agent needs to learn the dynamics from interactions with the world. However, learning dynamics models…

机器学习 · 计算机科学 2019-06-06 Danijar Hafner , Timothy Lillicrap , Ian Fischer , Ruben Villegas , David Ha , Honglak Lee , James Davidson

Training agents in multi-agent competitive games presents significant challenges due to their intricate nature. These challenges are exacerbated by dynamics influenced not only by the environment but also by opponents' strategies. Existing…

机器学习 · 计算机科学 2023-08-22 The Viet Bui , Tien Mai , Thanh Hong Nguyen

Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more,…

In this work, the trick-taking game Wizard with a separate bidding and playing phase is modeled by two interleaved partially observable Markov decision processes (POMDP). Deep Q-Networks (DQN) are used to empower self-improving agents,…

机器学习 · 计算机科学 2022-05-30 Jonas Schumacher , Marco Pleines

We show that a neural network originally designed for language processing can learn the dynamical rules of a stochastic system by observation of a single dynamical trajectory of the system, and can accurately predict its emergent behavior…

统计力学 · 物理学 2022-02-18 Corneel Casert , Isaac Tamblyn , Stephen Whitelam

In this paper we explore the performance of deep hidden physics model (M. Raissi 2018) for autonomous systems. These systems are described by set of ordinary differential equations which do not explicitly depend on time. Such systems can be…

机器学习 · 计算机科学 2024-08-08 Vijay Kag

In the last years, the DeepMind algorithm AlphaZero has become the state of the art to efficiently tackle perfect information two-player zero-sum games with a win/lose outcome. However, when the win/lose outcome is decided by a final score…

Pre-training Reinforcement Learning agents in a task-agnostic manner has shown promising results. However, previous works still struggle in learning and discovering meaningful skills in high-dimensional state-spaces, such as pixel-spaces.…

人工智能 · 计算机科学 2021-07-20 Juan José Nieto , Roger Creus , Xavier Giro-i-Nieto

In a competitive game scenario, a set of agents have to learn decisions that maximize their goals and minimize their adversaries' goals at the same time. Besides dealing with the increased dynamics of the scenarios due to the opponents'…

人工智能 · 计算机科学 2023-10-03 Pablo Barros , Alessandra Sciutti

Artificial intelligence for card games has long been a popular topic in AI research. In recent years, complex card games like Mahjong and Texas Hold'em have been solved, with corresponding AI programs reaching the level of human experts.…

人工智能 · 计算机科学 2024-09-16 Chang Lei , Huan Lei

We consider a broad class of stochastic imitation dynamics over networks, encompassing several well known learning models such as the replicator dynamics. In the considered models, players have no global information about the game…

系统与控制 · 计算机科学 2021-03-02 Lorenzo Zino , Giacomo Como , Fabio Fagnani

Traditional approaches to training agents have generally involved a single, deterministic environment of minimal complexity to solve various tasks such as robot locomotion or computer vision. However, agents trained in static environments…

机器人学 · 计算机科学 2025-10-01 Kevin Godin-Dubois , Karine Miras , Anna V. Kononova

Recent developments in the field of model-based RL have proven successful in a range of environments, especially ones where planning is essential. However, such successes have been limited to deterministic fully-observed environments. We…

机器学习 · 计算机科学 2021-06-11 Sherjil Ozair , Yazhe Li , Ali Razavi , Ioannis Antonoglou , Aäron van den Oord , Oriol Vinyals

We study a model of learning on social networks in dynamic environments, describing a group of agents who are each trying to estimate an underlying state that varies over time, given access to weak signals and the estimates of their social…

社会与信息网络 · 计算机科学 2013-07-19 Rafael M. Frongillo , Grant Schoenebeck , Omer Tamuz

World models aim to capture the states and dynamics of an environment in a compact latent space. Moreover, using Boolean state representations is particularly useful for search heuristics and symbolic reasoning and planning. Existing…

机器学习 · 计算机科学 2026-03-03 Davide Bizzaro , Luciano Serafini

We study binary coordination games with random utility played in networks. A typical equilibrium is fuzzy -- it has positive fractions of agents playing each action. The set of average behaviors that may arise in an equilibrium typically…

理论经济学 · 经济学 2021-09-01 Marcin Pęski