📡 Research Radar · Daily Edition

September 20, 2026

Window: August 8 – September 20, 2026
Sources: OpenReview · ACL Anthology · arXiv cs.CL/cs.LG/cs.CR/cs.AI · EMNLP 2026
3 peer-reviewed 7 preprints 0 forum/blog
safety-tuning EMNLP 2026 Main ✓ Minji Kim, Hyounghun Kim · arXiv:2609.04714 · Sep 4, 2026
Response decomposition: Boilerplate refusal statement "I cannot assist with..." → over-refusal ↑ Rationale component "...because it enables harm" → discrimination ↑
Decomposing safety responses into boilerplate vs. rationale reveals which component drives over-refusal.
Finds that safety-tuning over-refusal is largely driven by the boilerplate refusal statement (not the rationale) because it trains the model to pattern-match on superficial cues. Training only on rationales (removing the "I cannot..." boilerplate) reduces false refusal rates while maintaining comparable genuine safety performance — a cheap, high-leverage fix for aligned models that are overly conservative on benign inputs.
applied mech-interp Son Ha Xuan, Phat T. Tran-Truong, Xuan-Bach Le (RMIT / HCMUT) · arXiv:2609.18961 · Sep 16, 2026
Activation-patching recovery score per head/MLP block Top-K sites selected MechSparse-C: joint LoRA/QLoRA trained on selected sites only
MechSparse: activation-patching scores guide sparse LoRA placement, outperforming heuristic baselines at matched parameter budgets.
Uses normalized activation-patching recovery scores (causal signals from mechanistic interpretability) to select which attention heads and MLP blocks receive LoRA adapters, rather than using standard heuristics (magnitude, activation norm, gradient/Fisher). Evaluated on Ministral-8B/NF4 on Swahili span-JSON information extraction, MechSparse outperforms all heuristic baselines at matched parameter budgets — a concrete applied demonstration that mechanistic causal signals improve practical fine-tuning efficiency.

8 entries removed on 2026-09-29 as repeats of earlier reports: 2609.01455 (first covered 2026-09-03), 2606.19168 (first covered 2026-08-28), 2605.02087 (first covered 2026-09-01), 2609.00051 (first covered 2026-09-03), 2608.25390 (first covered 2026-08-29), 2609.02852 (first covered 2026-09-08), 2609.03026 (first covered 2026-09-08), 2609.15533 (first covered 2026-09-17).

← all Research Radar issues · gussand · source