Abstract: Multi-agent reinforcement learning in networked populations is governed by the interaction between individual adaptation, local encounters, and changing environmental conditions. To study this interaction, we formulate a coupled learning–environment model in which agents update stateless $Q$-values on a fixed graph, while their population-average behavior drives an environmental variable that dynamically modifies the payoff matrix. Under a first-order mean-field closure, we derive a deterministic transport equation for the population distribution of $Q$-values and couple it with a projected discrete update for the environmental state. The resulting model is evaluated against finite-network Monte Carlo simulations on random regular, Erd\H{o}s–R\'enyi, Barab\'asi–Albert, and random geometric graphs. Across the tested parameter ranges, the mean-field system reproduces the main macroscopic cooperation and environmental trajectories, and the trajectory-level root-mean-square error generally decreases with population size and average degree. The analysis further shows that environmental feedback reshapes the learned action-value ordering, while reinforcing feedback can produce pronounced dependence on the initial learning bias and resource level. The environmental timescale also plays an important role: a rapid response can drive the resource state to a boundary before learning adapts, whereas a slower response preserves the interaction between behavioral learning and environmental recovery. These results provide a population-level description of coupled reinforcement learning and environmental dynamics and characterize the performance of the mean-field approximation within the tested network and parameter ranges.
Read the original article:
