Research Radar
Daily · August 5, 2026
0 peer-reviewed · 1 preprints · 0 forum/blog
Mech Interp· AI Security· Text Diffusion LMs



Items 4 – 10 · Also notable



07
mech-interp AI security preprint

9 entries removed on 2026-09-10 as repeats of earlier reports: 2604.09839 (first covered 2026-07-26), 2606.20560 (first covered 2026-07-01), 2607.07368 (first covered 2026-07-17), 2607.01774 (first covered 2026-07-16), 2604.20945 (first covered 2026-07-07), 2607.02514 (first covered 2026-07-17), 2606.22673 (first covered 2026-07-02), 2606.26620 (first covered 2026-07-02), 2606.11998 (first covered 2026-07-08).

Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering

Steering Direction Norm — Chat vs. ReAct Agent 0 0.5 1.0 Direction norm Model layer scaffold zone Chat setting ReAct agent
Figure 7: Steering direction strength across model layers in chat vs. ReAct agent settings. Both reach near-full strength at late layers; the rescaling effect is localized to the ReAct format scaffold (early layers, shaded). Agentic deployment amplifies refusal-bypass by up to 2× on some models.

First systematic study of additive activation steering transfer from single-turn chat to ReAct tool-using agents. Transfer is real but rescaled: the injected direction reaches late layers at near-full strength in every tested model and setting, with the amplitude difference localized to ReAct's format scaffold before any tool observation. The security implication: chat-calibrated steering baselines do not carry over safely to agent deployments — agentic context can amplify steering-based refusal bypass by up to 2.00× depending on the model.





← all Research Radar issues · gussand · source