Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression

arXiv:2605.08475v3 Announce Type: replace-cross
Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass. Under bounded-data assumptions, we construct a single-head transformer whose forward pass approximately implements \textit{preconditioned Richardson iteration} on the associated kernel system. The construction uses $O(\log(1/\epsilon))$ blocks and MLP width $O(\sqrt{N/\epsilon})$ to achieve $\epsilon$-accurate prediction for prompts of length $N$. Our construction reveals a functional decomposition within the transformer architecture: softmax attention produces a row-normalized Gaussian-kernel operator needed for \emph{cross-token} interactions, while MLP layers act locally to approximate the \emph{intra-token} scalar arithmetic required by the update. Empirically, we train GPT-2-style transformers on Gaussian-process regression tasks and observe that they progressively align with the exact Gaussian KRR estimator across depth in terms of both \emph{prediction error} and \emph{induced weights}, with ablations further supporting this trend. Comparisons with classical KRR solvers also show that deeper layers align with later solver iterates. Together, we empirically demonstrate that the pretrained transformers exhibit progressive refinement toward exact Gaussian KRR across depth, and theoretically establish inexact preconditioned Richardson iteration as a concrete mechanism for approximating this predictor within an explicitly constructed softmax-attention transformer.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: