A note on goal-based hierarchical RL

arXiv:2609.14605v1 Announce Type: new
Abstract: The agent-centric general value function (ACGVF) construction of
\citet{tasse2026goal} lets the agent make two decisions that are
normally imposed by the environment or agent designer: which goal to
pursue and when to declare a goal as finished (in addition to
choosing the action). This is a very general framework that
subsumes almost all prior work on reinforcement learning, control
and planning, as well as more general formalisms proposed in the
cognitive sciences. However, it assumes the environment is fully
observed, i.e., that the observation is a sufficient statistic. In
\citet{murphy2025rl}, a general agent design was proposed where the
policy is based on an internal belief state $z_t$ and an internal
goal; however, the goals were assumed to be externally provided. In
this note, we unify and extend these two approaches using the
formalism of hierarchical hidden Markov models (HHMM) \citep{murphy2001hhmm}.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: