English
Related papers

Related papers: Strategic Exploration for Innovation

200 papers

In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (MDP), agents train on a fixed set of contexts and must generalise to new ones. Recent work has argued and demonstrated that increased exploration can…

Machine Learning · Computer Science 2025-05-23 Max Weltevrede , Caroline Horsch , Matthijs T. J. Spaan , Wendelin Böhmer

Unsupervised exploration and representation learning become increasingly important when learning in diverse and sparse environments. The information-theoretic principle of empowerment formalizes an unsupervised exploration objective through…

Machine Learning · Computer Science 2019-05-24 Jonathan Binas , Sherjil Ozair , Yoshua Bengio

Exploration is essential in reinforcement learning, particularly in environments where external rewards are sparse. Here we focus on exploration with intrinsic rewards, where the agent transiently augments the external rewards with…

Machine Learning · Computer Science 2024-01-26 Changmin Yu , Neil Burgess , Maneesh Sahani , Samuel J. Gershman

Solving tasks with sparse rewards is one of the most important challenges in reinforcement learning. In the single-agent setting, this challenge is addressed by introducing intrinsic rewards that motivate agents to explore unseen regions of…

Machine Learning · Computer Science 2021-05-25 Shariq Iqbal , Fei Sha

Search in an environment with an uncertain distribution of resources involves a trade-off between exploitation of past discoveries and further exploration. This extends to information foraging, where a knowledge-seeker shifts between…

Computation and Language · Computer Science 2017-02-03 Jaimie Murdock , Colin Allen , Simon DeDeo

Sample-efficient exploration is crucial not only for discovering rewarding experiences but also for adapting to environment changes in a task-agnostic fashion. A principled treatment of the problem of optimal input synthesis for system…

Machine Learning · Computer Science 2019-10-10 Matthias Schultheis , Boris Belousov , Hany Abdulsamad , Jan Peters

We study the optimal investment policy of a firm facing both technological and cash-flow uncertainty. At any point in time, the firm can decide to invest in a standalone technology or to wait for a technological breakthrough. Breakthroughs…

Optimization and Control · Mathematics 2021-06-10 Jean-Paul Décamps , Fabien Gensbittel , Thomas Mariotti

We consider online learning problems under a partial observability model capturing situations where the information conveyed to the learner is between full information and bandit feedback. In the simplest variant, we assume that in addition…

Machine Learning · Computer Science 2026-04-28 Tomas Kocak , Gergely Neu , Michal Valko , Remi Munos

In this survey we present different approaches that allow an intelligent agent to explore autonomous its environment to gather information and learn multiple tasks. Different communities proposed different solutions, that are in many cases,…

Artificial Intelligence · Computer Science 2014-03-07 Manuel Lopes , Luis Montesano

In this paper, I endeavour to construct a new model, by extending the classic exogenous economic growth model by including a measurement which tries to explain and quantify the size of technological innovation ( A ) endogenously. I do not…

Econometrics · Economics 2018-05-03 Murad Kasim

We provide a theoretical framework to understand when firms may benefit from exploiting previously abandoned technologies and brands. We model for the long run process of innovation, allowing for sustainable diversity and comebacks of old…

Economics · Quantitative Finance 2016-07-28 Shidong Wang , Renaud Foucart , Cheng Wan

The success of research institutions heavily relies upon identifying the right researchers "for the job": researchers may need to identify appropriate collaborators, often from across disciplines; students may need to identify suitable…

Computation and Language · Computer Science 2021-06-01 Oana Cocarascu , Andrew McLean , Paul French , Francesca Toni

Achieving effective test-time scaling requires models to engage in In-Context Exploration -- the intrinsic ability to generate, verify, and refine multiple reasoning hypotheses within a single continuous context. Grounded in State Coverage…

Computation and Language · Computer Science 2026-02-13 Futing Wang , Jianhao Yan , Yun Luo , Ganqu Cui , Zhi Wang , Xiaoye Qu , Yue Zhang , Yu Cheng , Tao Lin

Curiosity is a vital metacognitive skill in educational contexts. Yet, little is known about how social factors influence curiosity in group work. We argue that curiosity is evoked not only through individual, but also interpersonal…

Human-Computer Interaction · Computer Science 2017-10-24 Tanmay Sinha , Zhen Bai , Justine Cassell

When an individual's behavior has rational characteristics, this may lead to irrational collective actions for the group. A wide range of organisms from animals to humans often evolve the social attribute of cooperation to meet this…

Multiagent Systems · Computer Science 2021-11-18 Zhenbo Cheng , Xingguang Liu , Leilei Zhang , Hangcheng Meng , Qin Li , Xiao Gang

All learning algorithms for recommendations face inevitable and critical trade-off between exploiting partial knowledge of a user's preferences for short-term satisfaction and exploring additional user preferences for long-term coverage.…

Information Retrieval · Computer Science 2021-08-13 Kihwan Kim

We analyze the dynamic tradeoff between generating and disclosing evidence. Agents are tempted to delay investing in a new technology in order to learn from information generated by the experiences of others. This informational free-riding…

Theoretical Economics · Economics 2025-11-19 Jan Knoepfle , Julia Salmi

A general theory of innovation and progress in human society is outlined, based on the combat between two opposite forces (conservatism/inertia and speculative herding "bubble" behavior). We contend that human affairs are characterized by…

Physics and Society · Physics 2008-12-02 Didier Sornette

Modern recommendation systems ought to benefit by probing for and learning from delayed feedback. Research has tended to focus on learning from a user's response to a single recommendation. Such work, which leverages methods of supervised…

Information Retrieval · Computer Science 2023-08-01 Zheqing Zhu , Benjamin Van Roy

Incomplete knowledge of the environment leads an agent to make decisions under uncertainty. One of the major dilemmas in Reinforcement Learning (RL) where an autonomous agent has to balance two contrasting needs in making its decisions is:…

Machine Learning · Statistics 2024-02-21 Valentina Zangirolami , Matteo Borrotti