Abstract: World models support model-based planning through learned latent dynamics, but imagined rollouts can become unstable as the planning horizon grows or the dynamics distribution shifts. We propose HaM-World, a structured world model that combines history-conditioned selective memory with a Soft-Hamiltonian latent dynamics prior. The latent state is decomposed into a canonical (q,p) subspace and a context subspace c. Mamba selective state-space memory summarizes past observations and actions and conditions the same latent transition used for prediction, reward and value estimation, imagined rollouts, and cross-entropy method planning. The (q,p) subspace follows an energy-derived Hamiltonian vector field augmented with learnable residual and control dynamics, while c represents semantic, dissipative, and other non-conservative factors.
On six DeepMind Control Suite tasks, HaM-World ranks first on four tasks and second on two, achieving the highest average AUC on the four-task core suite (117.9, 9.5% above TD-MPC2). Within the short-to-medium horizons used by the planner, it reduces imagined-rollout error to 45% of a strong baseline and wins 11 of 12 rollout-MSE cells for horizons k in {3,5,7}. At longer open-loop horizons, its error grows faster and is overtaken between k=7 and k=20; we therefore do not claim uniform long-horizon stability. Under 12 out-of-distribution perturbations involving dynamics shifts, action delay, and observation masking, it achieves the highest absolute return in every condition, with average gains of 10.2% on Finger Spin and 13.6% on Reacher Easy. Ablations show that memory accounts for the larger share of the observed gains, while Soft-Hamiltonian geometry provides smaller but consistent complementary improvements.
Read the original article:
