Intrinsic motivation supports adaptive behavior and self-organization in ai systems
Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
Artificial Intelligence
Summary
Many living cells work together without a single shared goal, leading to complex behaviors. The authors explore whether artificial intelligence systems can use internal rewards, like curiosity or learning progress, to adapt and organize themselves without external instructions. They review different ways AI can generate these internal rewards and when they might fail. The authors also suggest new experiments to see if these internal drives can lead to more advanced and coordinated learning over time.
What this means in practice
- •For robotics engineers: Design robots that learn exploration skills autonomously before focusing on specific tasks, improving adaptability in changing environments.
- •For multi-agent system developers: Build networks of AI agents each with internal rewards to study how collective behavior emerges without explicit external goals.
A position paper. It proposes an approach and reports no results.
Authors
Anatoly Belikov
Abstract
Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organisation, from individual cells through tissues to whole multicellular organisms. In this perspective and tutorial article we discuss whether intrinsic rewards in artificial neural systems can support adaptation, functional specialisation and higher-level self-organisation without a shared external objective. We review empowerment, curiosity, learning progress, information gain, unsupervised skill discovery, mutual information estimation and the use of world models for intrinsic reward computation. Particular attention is given to failure modes showing when such objectives do not produce sustained exploration or increasingly complex behaviour. We argue that more capable systems may require complementary objectives, communication, memory, learning at multiple temporal scales and environmental constraints. Based on this perspective, we outline three experimental directions. These include a resource-constrained environment in which otherwise stable behavioural attractors become unsustainable, allowing us to test whether environmental constraints can mitigate characteristic failure modes of intrinsic objectives. The network of recurrent agents with per-agent intrinsic rewards, and a hierarchical world-model agent in which exploratory motor competence develops before goal-directed behaviour. These experiments are intended to test whether intrinsic learning can lead to adaptive organisation at progressively higher levels.