A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging

arXiv:2610.00126v1 Announce Type: cross
Abstract: Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can nevertheless make a solver comparison vacuous: if its probes or predicates encode the target identity, an exact optimizer may appear effective without resolving any genuine ambiguity. We report such a failure in an aggregate-trace debugger for a closed-loop decision agent. Exact minimum hitting set (MHS) and a propagation-aware greedy method returned identical supports in 12/12 development cases and the same planted-fault recovery in 9/12. A subsequent audit found that exact-anchor predicates produced the planted pair in 9/9 cases. After removing those anchors, overall planted-pair recovery was 8/9; hard-probe singleton pairs nevertheless matched the planted pair in 9/9, and no case retained a nonempty residual conflict family after propagation (0/9). The optimizer was correct, but the verifier had already disclosed the answer. We replace solver-first evaluation with a support-gated verification contract. A clean reference map must first show repeated component exposure; a matched reference/current gate must then establish comparable runtime evidence; only afterward may an independently calibrated signal rule return a detection. In a preregistered heldout comprising 1,440 cases and 21,600 partition rows, 55/72 regime-component units passed the reference gate, 54/55 passed the runtime gate, and stable false admission was 0/20 represented components with a one-sided exact 95% upper bound of 0.1391. Within admitted units, affected clean traffic predicted detection better than nominal fault-cell fraction. The main lesson is structural: verify evidence eligibility and non-revelation before optimizing the component selector. Otherwise a stronger solver can merely certify a stronger verifier artifact.

This article has been indexed from cs.AI updates on arXiv.org

Read the original article: