Research Radar
August 27, 2026
Mech Interp · AI Security · Text Diffusion LMs
0 peer-reviewed · 5 preprints · 0 forum/blog

02
mech-interp preprint
Forward-Looking Localization: Pre-SFT → Post-SFT Pre-SFT Model θ₀ + Target dataset D naive: wrong Taylor Expansion ∂ℒ/∂θ · Δθ bridges pre → post Predicted Post-SFT Critical neurons Accurate LoRA targets ✗ Naive pre-SFT localization → wrong subspace → degraded PEFT ✓ Taylor-predicted localization → correct LoRA targets → stronger performance
Fig. 1 — Pre-SFT parameters + target dataset → Taylor expansion → predicted post-SFT critical neurons. Naive pre-SFT localization biases LoRA toward the wrong subspace; the forward-looking framework corrects this before fine-tuning begins.

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

Hang Chen, Jiaying Zhu, Wenya Wang — arXiv preprint, August 25, 2026

Mechanistic interpretability has always looked backward — you identify what the trained model uses. For novel-task fine-tuning, this timing is not just inconvenient but actively harmful: the neurons the pre-SFT model uses differ sharply from those that end up governing the fine-tuned model, and pointing LoRA at the wrong targets degrades performance.

The authors model SFT as continuous parameter evolution and derive a Taylor expansion of the post-SFT mechanistic objective from pre-SFT weights and the target dataset. This closed-form approximation forecasts which parameters will be task-critical after training, enabling accurate LoRA target selection before any fine-tuning occurs. Empirically, naively using pre-SFT localization introduces active bias for novel tasks — the pre- and post-SFT critical-neuron sets diverge substantially — and the proposed forward-looking framework outperforms standard PEFT baselines. The 25-page paper is the first to unite mechanistic interpretability with pre-training optimization in a locating-then-tuning paradigm grounded in theory.



Items 4–10 · Also notable
04
mech-interp AI security preprint
BLADE: Composite Score Gain over Prior Best Baseline +10% +7% +3% 0 +6% TOFU +9% MUSE Books +7% KnowUndo
Fig. 1 — BLADE composite score improvement over prior best baseline on three standard LLM unlearning benchmarks.

BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning

arXiv, August 2026

Bilevel optimization confined to LoRA adapters: a clamped-entropy forget loss erases target knowledge; an asymmetric augmented Lagrangian protects the retain set; the bilevel structure runs a retain-repair step before each forgetting update. Delivers the strongest systematic gains across all three standard unlearning benchmarks in a single method: +6% on TOFU, +9% on MUSE Books, +7% on KnowUndo composite scores over prior best baselines.




07
mech-interp preprint
Mechanistic Circuit Control for Data Synthesis (SAMS) Learnability Challenge Alignment Circuit Identification causal intervention SAMS Scheduling Stage-Aware Mechanistic circuit-steered data generation
Fig. 1 — Three data utility axes → circuit identification via causal intervention → SAMS scheduling steers synthetic data generation based on the model's current optimization stage.

Mechanistic Circuit Identification for Controllable Data Generation

Nakyung Lee, Sangwoo Hong, Jungwoo Lee (Seoul National University, Konkuk University) — arXiv, August 25, 2026

Identifies model-internal circuits governing three data quality axes (learnability, challenge, alignment) via causal intervention and uses them as controllable interfaces to steer LLM-based data generation. SAMS (Stage-Aware Mechanistic Scheduling) schedules circuit-steered samples according to the model's evolving optimization dynamics rather than static heuristic prompts — giving data synthesis pipelines feedback from the model's actual learning state for the first time.


08
AI security preprint
Tool-Mediated Recovery: A New Unlearning Failure Mode Query about forgotten info 🚫 Parametric recall blocked by unlearning ⚠ Tool calls recover web search · retrieval · DB ATU Solution Stage 1: parametric unlearn Stage 2: RL trajectory penalize target-seeking tool calls
Fig. 1 — Tool-mediated recovery: parametric unlearning blocks direct recall but agents recover the same information via tool calls. ATU adds RL-based trajectory penalization to suppress target-seeking tool use while preserving normal tool access for retained knowledge.

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

Baicheng Chen, Zheyuan Liu et al. — arXiv, August 21, 2026

Identifies tool-mediated recovery as a new unlearning failure mode: suppressing direct parametric recall is insufficient when an agent can route the same query through web search, retrieval, or database tools. ATU (Agentic Tool Unlearning) addresses this with Stage 1 (parametric unlearning) followed by Stage 2 (trajectory-level RL penalizing target-seeking tool calls and final-answer leakage). First work to study unlearning specifically in the agentic tool-use setting.


09
mech-interp AI control preprint
Mechanistic Tomography: Measurement × Control Design Space Circuits Features / SAEs Probes Intervention Monitoring Prediction Safety audit High fidelity ✓ Scalable ✓ Limited scope Best fit ✓ Lightweight ✓ High cost Emerging ✓ Ground truth ✓ Scalable ✓ Fast screen ✓
Fig. 1 — Mechanistic Tomography: interpretability tools (circuits, SAEs, probes) mapped against control objectives (intervention, monitoring, prediction, safety audit). Tool selection is co-determined by the intended control operation.

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Vijay Erramilli — arXiv, August 2026

A 24-page, 13-figure framework arguing that mechanistic interpretability measurements should be co-designed with their control objectives — what interventions are feasible, what robustness is required, what scale is needed — rather than evaluated in isolation. Bridges the mech interp toolbox (circuits, features, probes) with control-theoretic perspectives, providing a structured decision space for tool selection based on the intended downstream control operation.


Notes · August 8–27, 2026

5 entries removed on 2026-09-10 as repeats of earlier reports: 2608.07430 (first covered 2026-08-11), 2608.08168 (first covered 2026-08-13), 2608.11408 (first covered 2026-08-18), 2608.13538 (first covered 2026-08-22), 2608.10530 (first covered 2026-08-13).

← all Research Radar issues · gussand · source