Hang Chen, Jiaying Zhu, Wenya Wang — arXiv preprint, August 25, 2026
Mechanistic interpretability has always looked backward — you identify what the trained model uses. For novel-task fine-tuning, this timing is not just inconvenient but actively harmful: the neurons the pre-SFT model uses differ sharply from those that end up governing the fine-tuned model, and pointing LoRA at the wrong targets degrades performance.
The authors model SFT as continuous parameter evolution and derive a Taylor expansion of the post-SFT mechanistic objective from pre-SFT weights and the target dataset. This closed-form approximation forecasts which parameters will be task-critical after training, enabling accurate LoRA target selection before any fine-tuning occurs. Empirically, naively using pre-SFT localization introduces active bias for novel tasks — the pre- and post-SFT critical-neuron sets diverge substantially — and the proposed forward-looking framework outperforms standard PEFT baselines. The 25-page paper is the first to unite mechanistic interpretability with pre-training optimization in a locating-then-tuning paradigm grounded in theory.
Items 4–10 · Also notable
04
mech-interpAI securitypreprint
Fig. 1 — BLADE composite score improvement over prior best baseline on three standard LLM unlearning benchmarks.
Bilevel optimization confined to LoRA adapters: a clamped-entropy forget loss erases target knowledge; an asymmetric augmented Lagrangian protects the retain set; the bilevel structure runs a retain-repair step before each forgetting update. Delivers the strongest systematic gains across all three standard unlearning benchmarks in a single method: +6% on TOFU, +9% on MUSE Books, +7% on KnowUndo composite scores over prior best baselines.
07
mech-interppreprint
Fig. 1 — Three data utility axes → circuit identification via causal intervention → SAMS scheduling steers synthetic data generation based on the model's current optimization stage.
Nakyung Lee, Sangwoo Hong, Jungwoo Lee (Seoul National University, Konkuk University) — arXiv, August 25, 2026
Identifies model-internal circuits governing three data quality axes (learnability, challenge, alignment) via causal intervention and uses them as controllable interfaces to steer LLM-based data generation. SAMS (Stage-Aware Mechanistic Scheduling) schedules circuit-steered samples according to the model's evolving optimization dynamics rather than static heuristic prompts — giving data synthesis pipelines feedback from the model's actual learning state for the first time.
08
AI securitypreprint
Fig. 1 — Tool-mediated recovery: parametric unlearning blocks direct recall but agents recover the same information via tool calls. ATU adds RL-based trajectory penalization to suppress target-seeking tool use while preserving normal tool access for retained knowledge.
Baicheng Chen, Zheyuan Liu et al. — arXiv, August 21, 2026
Identifies tool-mediated recovery as a new unlearning failure mode: suppressing direct parametric recall is insufficient when an agent can route the same query through web search, retrieval, or database tools. ATU (Agentic Tool Unlearning) addresses this with Stage 1 (parametric unlearning) followed by Stage 2 (trajectory-level RL penalizing target-seeking tool calls and final-answer leakage). First work to study unlearning specifically in the agentic tool-use setting.
09
mech-interpAI controlpreprint
Fig. 1 — Mechanistic Tomography: interpretability tools (circuits, SAEs, probes) mapped against control objectives (intervention, monitoring, prediction, safety audit). Tool selection is co-determined by the intended control operation.
A 24-page, 13-figure framework arguing that mechanistic interpretability measurements should be co-designed with their control objectives — what interventions are feasible, what robustness is required, what scale is needed — rather than evaluated in isolation. Bridges the mech interp toolbox (circuits, features, probes) with control-theoretic perspectives, providing a structured decision space for tool selection based on the intended downstream control operation.
Notes · August 8–27, 2026
All 10 items are genuine; no padding. Window covers August 8–27 (last report: August 7).
#1 (2608.07430) is the week's standout — the only paper spanning all three radar tracks. Safety-neuron pruning achieves 73.8–86.6% ASR on LLaDA/Dream with a ≤3% baseline. Flag for weekly roundup.
Unlearning had an unusually strong week: #4 (BLADE), #5 (J-Access/Measure), #8 (Agentic Tool Unlearning) cover three distinct angles of concept erasure. Flag for weekly theme.
#2 and #7 together suggest a shift from interpretability-as-explanation to interpretability-as-control: using circuits/features to guide SFT targeting (#2) and data synthesis (#7).
Text diffusion LMs: 1 direct entry (#1); dLLM track remains dominated by the safety-vulnerability story.
5 entries removed on 2026-09-10 as repeats of earlier reports: 2608.07430 (first covered 2026-08-11), 2608.08168 (first covered 2026-08-13), 2608.11408 (first covered 2026-08-18), 2608.13538 (first covered 2026-08-22), 2608.10530 (first covered 2026-08-13).