Abstract: Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutation and recombination operators involved are qualitatively different from those studied classically. Mutations are no longer random; an ML algorithm mutates a solution with the goal of improving an objective. Similarly, recombination is not based on random collages of parent solutions. Instead, it is an ML optimization-based operator whose goal is to synthesize improved solutions from its inputs. Thus, these mutation and recombination operators are more likely to improve the objective, but their computational cost is much higher.
We introduce a general model of genetic algorithms and formulate optimization in this model as a query complexity problem, using the language of reinforcement learning. We demonstrate three fundamental phenomena. First, we show that diversity of the solution pool can be necessary: for parity learning, viewed in our framework, we show that with pool size $w$ and vectors of length $n$, the optimal query complexity is $\Theta(w+2^{n-w})$. We further show that this phenomenon persists under general memory constraints: $\Theta(n^2)$ bits of memory are necessary for efficient success. Second, we show that generation, mutation, and recombination can all be simultaneously necessary to reach a nearly optimal solution. Finally, we give a phase transition for Gaussian distributions, showing that a positive {\em drift} of the operators yields exponential speedup.
Read the original article:
