Abstract: How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-dimensional coordinate systems fit against a chosen functional of the representation, under the functional's own metric, turning distillation into plain least squares. Across six models from three families, spanning 70m to 7B parameters, next-token prediction needs 70–90% of the residual stream's width to stay within 5% of intact perplexity, a width consumed by the rare tail of language, and the variance profile predicts none of it: two directions carry 90% of GPT-2's activation variance and almost none of its function. Dimension is per-functional: the model's own uncertainty reads from six coordinates where the full predictive distribution needs hundreds; and it grows with depth. The dissociation is exploitable: when only a few dimensions can be kept, charts trained under the functional's metric preserve the model's predictions better than variance-based or optimal linear compression.
Read the original article:
