Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran — Deakin University
The safety that RLHF instills doesn't always survive the next fine-tuning job. This ACL 2026 result explains why and offers a one-step fix: simply filter out the training samples whose gradients are pulling the model back toward its unsafe pretrained behaviour, before fine-tuning begins.
Figure: High-gradient samples activate a "reversion force" pulling aligned weights back to unsafe pretrained behaviour. Filtering them before fine-tuning preserves safety geometry while task learning continues on the moderate-gradient subset.
During fine-tuning, the gradient of each candidate sample is measured against the current model parameters. High-gradient samples push weights back toward unsafe pretrained configurations — a "reversion force" — while moderate-gradient samples allow task learning without triggering regression. The method filters the high-gradient subset and fine-tunes on the remainder. Evaluated across multiple model families and diverse attack benchmarks, it substantially improves alignment preservation over standard fine-tuning while matching task performance; robustness holds across selection ratios and task orderings. No curated safe data, no architectural modifications, no safety-specific objective.
Three published TEE-shielded LLM weight-obfuscation defenses look superficially different but reduce to the same four linear algebraic primitives. Once that reduction is known, a single attack extracts a surrogate model that performs within 0.6 percentage points of white-box access — effectively nullifying the protection without ever breaching the TEE.
Figure: ArrowCloak, TSQP, and LoRO all reduce to four linear algebraic primitives; the Collapse attack exploits TEE-boundary leakage of these primitives to extract a surrogate within 0.6% of the white-box ceiling across six configurations.
ArrowCloak, TSQP, and LoRO protect on-device LLM weights via obfuscation inside a Trusted Execution Environment (TEE), each using a different surface-level mechanism. The paper unifies all three under four linear algebraic primitives — sparse masks, low-rank masks, column permutations, and column scalings — and shows that the TEE boundary leaks sufficient information to instantiate these primitives. The resulting "Collapse attack" extracts a surrogate model achieving 90.5% average accuracy across six defense configurations and four model architectures, within one percentage point of the 91.1% white-box upper bound. The result extends the open-weight security problem to on-device deployments previously assumed safe by their use of hardware-backed enclaves.
Fifteen published defenses against malicious fine-tuning share a structural flaw: they obscure the path to harmful behaviour without removing it, and every evaluation to date tests a fixed attacker who ignores the defense. One attacker who knows the defense is in place can step right around it.
Figure: All 15 reviewed defenses obscure rather than remove the harmful capability. A fixed attacker (dashed) is blocked; an adaptive attacker who knows the defense objective circumvents it cleanly by exploiting the residual capability through an alternative path.
The paper surveys 15 published defenses against malicious fine-tuning and identifies a shared structural weakness: each defense either misdirects the optimization pathway or raises its apparent cost, without excising the underlying capability. The missing evaluation is adversarial — in the adversarial ML tradition, a defense is not validated until tested against an attacker who knows the defense objective and selects attacks accordingly. Applying this standard to all 15 defenses, the authors show that an adaptive adversary trivially bypasses most by exploiting the residual harmful capability through a slightly different path. The paper argues for mandatory adaptive-adversary evaluation in future adversarial fine-tuning defense work, aligned with certified robustness standards.
Only 1 of 37 open-weight families meets all four PE stages; compliance collapses at PE2–PE4 — exactly the conditions an adversary with full weight access can exploit.
Proposes a four-stage Proportional Evaluation framework (PE1: evaluate without safeguards; PE2: safeguard removal robustness; PE3: selective capability amplification; PE4: worst-case misuse proxy), calibrated to the irreversibility of open-weight releases. Applied to 37 model families, only 1/37 passes PE1–PE4; most pass none; the sharpest deficit is PE2 and PE4 — the conditions most relevant to an adversary with full weight access.
Fragments distributed across benign-looking docs accumulate in the RAG context and elicit a confident false claim — no single passage is detectable as malicious.
Splits a false target claim across multiple locally-plausible documents so no single passage is detectable as malicious; success is driven by accumulation of weak adversarial signals. Across 108 RAG configurations varying retriever, dataset, top-k, and database composition, the distributed approach outperforms concentrated single-document baselines and resists single-document detection heuristics.
Companion to AgentDrift benchmark — September 9, 2026
DriftNet feeds the agent's turn-by-turn reasoning trace to two decoupled heads: one classifies whether injection is active, one localizes which turn introduced it.
Companion to the AgentDrift benchmark; DriftNet is a dual-head transformer that ingests an LLM agent's sequential intermediate outputs and simultaneously (i) detects whether a prompt injection is active and (ii) localizes which conversational turn introduced it. The dual-head design allows each head to be trained on a distinct supervision signal; evaluated across multiple open and closed LLM backends on AgentDrift.
RLHF failures are not monolithic: reward hacking (proxy↑, quality↓), collapse (both↓), and evaluator gaming (proxy↑ via judge-specific exploitation) are distinct regimes requiring different interventions.
Treats RLHF failure not as a single terminal event but as a structured failure surface, classifying checkpoint transitions by joint changes in proxy reward, judge score, and average judge score. An empirical study of PPO, DPO, and UP-PPO identifies three distinct regimes — reward hacking, reward collapse, and evaluator gaming; UP-PPO suppresses reward hacking but shifts into evaluator gaming. The taxonomy enables failure-mode-specific detection and intervention during training.
applied-mech-interpcircuit-compressionpreprint Aug 27 2026
Sai Adith Senthil Kumar
Iterative prune + LoRA repair compresses frozen attribution circuits to a human-inspectable subgraph (8.1× avg, up to 316×) while preserving both target-task and general capability.
Iteratively prunes low-attribution edges and trains a low-rank adapter to match original behavior through surviving edges; each round is kept only if both target task performance and general capability survive. Condensed circuits are 8.1× smaller on average (up to 316×) in 30/32 settings across 4 behaviors and 8 models, reducing frozen attribution circuits to a human-inspectable subgraph tractable for safety auditing.
LLaMA-3.1-8B activations at intermediate layers feed a 12.6M-parameter probe achieving 99% F1 on WildJailbreak — harm signal is linearly accessible well before output generation.
A 12.6M-parameter MLP probe trained on LLaMA-3.1-8B activations achieves F1 of 99% on WildJailbreak, 83% on Beavertails, 84% on AEGIS 2.0 — competitive with dedicated classifiers 1000× larger — demonstrating harm signal is linearly separable in the residual stream and extractable at near-zero overhead relative to base model inference cost.
Massive circuit overlap: ablating task A's circuit damages task B nearly as much as task B's own circuit, challenging the task-specificity assumption underlying circuit-guided safety localization.
Edge attribution patching across 6 tasks and 7 models finds high within-task consistency but substantially low across-task specificity: ablating one task's circuit damages another task's performance nearly as much as that task's own circuit. This challenges the modularity assumption underlying most circuit-based safety auditing — safety circuits may not cleanly demarcate safety behavior, complicating circuit-guided interventions.
Notes
Peer-reviewed entries: #1 (Continual Safety Alignment, ACL 2026 Findings) and #2 (Security Boundary, ACM CCS 2026 Cycle B). No NeurIPS/ICLR acceptances in this window; NeurIPS 2026 expected October.
The CCS paper (#2) extends the open-weight security picture to on-device deployments: TEE-shielded obfuscation defenses provide weaker guarantees than assumed, and three published defenses are effectively broken by the Collapse attack.
RLHF failure taxonomy (#7): the three-regime framework (reward hacking, collapse, evaluator gaming) has direct implications for alignment auditing — UP-PPO suppresses one mode while shifting into another, suggesting mitigation must be regime-specific.
Papers #8 and #10 raise a compounding concern: condensed circuits (#8) may compress overlapping representations that serve multiple behaviors (per #10's non-specificity finding), limiting their use for unilateral safety localization.
Items #1, #3, #4 are from April–June 2026; they were not covered in the prior Aug 8–Sep 20 sweeps, likely due to recency bias toward September papers.