When a safety-aligned vision-language model refuses to answer a visual question, does it refuse because it stopped processing the image — or because it processed it correctly and then decided to withhold the answer? It's the latter. And a targeted activation nudge is enough to restore the visual answer without any retraining.
Aligned VLMs frequently abstain from visual questions that remain answerable under default (unaligned) instruction, even when the image-question input is identical. The paper analyzes internal decoding dynamics across multiple architectures and multimodal benchmarks, finding that visual evidence consistently and significantly influences the internal representations of abstained outputs throughout decoding — perceptual grounding is retained, not suppressed, by safety alignment. Safety alignment operates as a late-stage representational override: the model sees and encodes the image normally, routes the visual signal through its reasoning pathway, and only at the final generation stage substitutes a refusal. Targeted activation-level interventions that suppress refusal-related representations reliably restore grounded, correct answering behavior across all tested architectures — no retraining, no modification of visual inputs required. The finding has dual implications: it clarifies why safety-aligned VLMs are often recoverable via activation steering, and it suggests that safety circuits and perceptual circuits are anatomically separable at the representation level.