Abstract: We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in the short-time Fourier transform (STFT) domain and measures attack success and audio-span representations. On 520 AdvBench prompts, the primary judge labels 76.7% of adversarial inputs as jailbreaks. A condition-blind, single-annotator validation yields a Rogan-Gladen sensitivity estimate of 0.87 for this condition (about 0.83-0.95 with validation-rate uncertainty); this correction is not applied to masked conditions. The apparent frequency ranking depends on the partition: energy share alone predicts the standard eight-band ranking (Spearman's rho = 0.95), and equal-Hz and equal-energy partitions show that masking any tested band can sharply reduce attack success. At a finer 16-band equal-energy resolution, however, masking the narrow 7520-7960 Hz band leaves ASR at 0.10, which remains unresolved without a matched control. Matched-energy scattered-removal tests show that contiguous removal is more damaging in lower bands, while both forms approach the floor in upper bands. Global rescaling leaves ASR near baseline but tests amplitude sensitivity rather than frequency location. In a re-optimization pilot (n = 20), tested single- and two-band supports reach ASRs of 0.00, 0.25, and 0.40, while random supports covering about half the STFT bins reach a mean of 0.86. In a prompt- and energy-adjusted model, audio-span divergence is associated with band necessity, with the coefficient rising from +0.63 at the projector output to +0.93 at layer 30 (contrast +0.293, 95% CI [0.11, 0.51]). This association is not a causal localization, and single-layer patching does not establish a causal layer. The results support partition-aware auditing of frequency claims, leaving the fine-resolution top-band result and broader generality open.
Read the original article:
