Abstract: We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare: simulators provide cheap data, but sim-to-real mismatch makes direct transfer unreliable, and collecting real-world data to correct this mismatch must itself be safe. Moreover, deployment objectives may vary across tasks, making it costly to collect new data for each reward function. We therefore formulate safe sim-to-real transfer as a reward-free safe reinforcement learning (RL) problem, in which data are collected once and reused to plan for arbitrary reward functions. We develop a computationally efficient algorithm that identifies where the simulator and real dynamics differ, uses certified simulator transitions where they are reliable, and estimates mismatched transitions from safely collected data. With high probability, every policy deployed during learning is feasible, and the collected data support the computation of a feasible and near-optimal policy for any reward function. When the simulator is uninformative, our algorithm recovers online reward-free safe RL while improving the best-known sample complexity by a factor of \(\widetilde{\Theta}(H/\xi^2)\), where \(\xi\) is the safety margin of a baseline policy. When the simulator is accurate on most transitions, this improvement grows to \(\widetilde{\Theta}(H^2|\mc S||\mc A|/(\xi^2|\mc B|))\), where \(|\mc B|\) denotes the size of the sim-to-real mismatch region.
Read the original article:
