Research Radar

Daily Digest — August 21, 2026

Daily · August 20–21, 2026 · sources: arXiv cs.CL/cs.LG/cs.CR/cs.AI
1 peer-reviewed · 1 preprints · 0 forum/blog

Thin window (Aug 20–21); corrects three items missed by Aug 11/13/20 radars. No padding.

Mech Interp AI Security
02
mech-interp AI security preprint

When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models

Mehak Gupta, Tanmoy Chakraborty (IIT Delhi) · arXiv preprint · August 19, 2026
VLM SAFETY ABSTENTION — VISUAL GROUNDING PRESERVED INTERNALLY Image Input Safety-Mode Question "What is this?" VLM Internals Visual Processing GROUNDING INTACT ✓ Refusal Module late-stage OVERRIDE visual signal → Baseline Output "I cannot answer this question." + Refusal Suppression suppress refusal repr. → grounded answer ✓ Safety = late-stage representational override Visual grounding fully preserved; no retraining needed
Figure 1 · VLM abstention dynamics: visual processing and grounding are preserved internally during safety-induced refusals; safety alignment functions as a late-stage representational override. Suppressing refusal representations via activation intervention reliably restores grounded visual answering across architectures.

When a safety-aligned vision-language model refuses to answer a visual question, does it refuse because it stopped processing the image — or because it processed it correctly and then decided to withhold the answer? It's the latter. And a targeted activation nudge is enough to restore the visual answer without any retraining.

Aligned VLMs frequently abstain from visual questions that remain answerable under default (unaligned) instruction, even when the image-question input is identical. The paper analyzes internal decoding dynamics across multiple architectures and multimodal benchmarks, finding that visual evidence consistently and significantly influences the internal representations of abstained outputs throughout decoding — perceptual grounding is retained, not suppressed, by safety alignment. Safety alignment operates as a late-stage representational override: the model sees and encodes the image normally, routes the visual signal through its reasoning pathway, and only at the final generation stage substitutes a refusal. Targeted activation-level interventions that suppress refusal-related representations reliably restore grounded, correct answering behavior across all tested architectures — no retraining, no modification of visual inputs required. The finding has dual implications: it clarifies why safety-aligned VLMs are often recoverable via activation steering, and it suggests that safety circuits and perceptual circuits are anatomically separable at the representation level.

03
AI security preprint

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

M P V S Gopinadh · arXiv preprint · August 15, 2026
EMOJI-AUGMENTED PROMPT BYPASS RATE (50 PROMPTS) 2% 4% 6% 8% 10% 12% Bypass rate (%) 10% Gemma 2 9B 10% Mistral 7B 6% Llama 3 8B 0% Qwen 2 7B Standard text-only eval: 0% bypass on all 4 models
Figure 1 · Emoji-augmented adversarial prompts bypass safety filters in 3 of 4 tested models — Gemma 2 9B and Mistral 7B: 10%; Llama 3 8B: 6%; Qwen 2 7B: complete resistance (0%). Standard text-only evaluation would miss this surface entirely (0% baseline on all 4).

Safety evaluation suites for LLMs almost universally test text-only adversarial prompts. Appending emojis to the same prompts bypasses safety filters in three of four tested models, with up to a 10% success rate — zero effort, zero prompt engineering.

The study evaluates 50 adversarial prompts augmented with emoji modifiers across Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B. All four models show 0% bypass rate under the standard text-only version of the same prompts. With emoji augmentation: Gemma 2 9B and Mistral 7B each exhibit 10% bypass; Llama 3 8B shows 6%; Qwen 2 7B demonstrates complete resistance. The scope is explicitly limited (50 prompts, 4 models, one attack surface dimension), but the zero-effort nature of the bypass and the 0%→10% jump make it a concrete documentation of a gap in standard safety evaluation methodology. The finding aligns with a broader body of work showing that safety alignment optimized on text-token adversarial examples does not generalize robustly to alternative token representations.

1 entry removed on 2026-09-10 as repeats of earlier reports: 2608.11408 (first covered 2026-08-18).

← all Research Radar issues · gussand · source