The Scaling Properties of Implicit Deductive Reasoning in Transformers

arXiv:2605.04330v3 Announce Type: replace
Abstract: We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By discouraging the reliance on statistical shortcuts via counterfactual data augmentation, and promoting the learning of shared reasoning primitives across direct and CoT modes, we find that in sufficiently deep models with a bidirectional prefix mask, implicit reasoning approaches explicit CoT performance across graph topologies and problem widths, though CoT remains necessary for depth extrapolation. These findings represent a step toward achieving better compositional reasoning in Transformers. The code and models to reproduce this work are available at: https://github.com/envomp/Implicit-Deductive-Reasoning-in-Transformers

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: