📡 Research Radar

Daily Radar — 2026-09-22

Window: September 19–22, 2026 (plus important papers from June–September missed in prior sweep)
Sources: arXiv (cs.CL/cs.LG/cs.CR/cs.AI) · ACL Anthology 2026 · ACM CCS 2026 · RAND · OpenReview
2 peer-reviewed 8 preprints 0 forum/blog
Top 10 — Priority Order
01

Continual Safety Alignment via Gradient-Based Sample Selection

pretraining-safety ACL 2026 Findings peer-reviewed
Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran — Deakin University

The safety that RLHF instills doesn't always survive the next fine-tuning job. This ACL 2026 result explains why and offers a one-step fix: simply filter out the training samples whose gradients are pulling the model back toward its unsafe pretrained behaviour, before fine-tuning begins.

Per-sample gradient magnitude during fine-tuning threshold moderate ∇ — kept high ∇ — filtered (reversion force) gradient magnitude →
Figure: High-gradient samples activate a "reversion force" pulling aligned weights back to unsafe pretrained behaviour. Filtering them before fine-tuning preserves safety geometry while task learning continues on the moderate-gradient subset.

During fine-tuning, the gradient of each candidate sample is measured against the current model parameters. High-gradient samples push weights back toward unsafe pretrained configurations — a "reversion force" — while moderate-gradient samples allow task learning without triggering regression. The method filters the high-gradient subset and fine-tunes on the remainder. Evaluated across multiple model families and diverse attack benchmarks, it substantially improves alignment preservation over standard fine-tuning while matching task performance; robustness holds across selection ratios and task orderings. No curated safe data, no architectural modifications, no safety-specific objective.


02

Understanding the Security Boundary of Obfuscation-based On-Device LLM Protection

AI-security open-weight ACM CCS 2026 peer-reviewed
ACM CCS 2026, Cycle B (peer-reviewed)

Three published TEE-shielded LLM weight-obfuscation defenses look superficially different but reduce to the same four linear algebraic primitives. Once that reduction is known, a single attack extracts a surrogate model that performs within 0.6 percentage points of white-box access — effectively nullifying the protection without ever breaching the TEE.

Collapse attack: three defenses → four primitives → surrogate ArrowCloak TSQP LoRO sparse masks low-rank masks col permutations col scalings 4 linear primitives Collapse attack 90.5% acc vs 91.1% white-box surrogate
Figure: ArrowCloak, TSQP, and LoRO all reduce to four linear algebraic primitives; the Collapse attack exploits TEE-boundary leakage of these primitives to extract a surrogate within 0.6% of the white-box ceiling across six configurations.

ArrowCloak, TSQP, and LoRO protect on-device LLM weights via obfuscation inside a Trusted Execution Environment (TEE), each using a different surface-level mechanism. The paper unifies all three under four linear algebraic primitives — sparse masks, low-rank masks, column permutations, and column scalings — and shows that the TEE boundary leaks sufficient information to instantiate these primitives. The resulting "Collapse attack" extracts a surrogate model achieving 90.5% average accuracy across six defense configurations and four model architectures, within one percentage point of the 91.1% white-box upper bound. The result extends the open-weight security problem to on-device deployments previously assumed safe by their use of hardware-backed enclaves.


03

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries

pretraining-safety adversarial-finetuning preprint May 2026
Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky — Ben-Gurion University

Fifteen published defenses against malicious fine-tuning share a structural flaw: they obscure the path to harmful behaviour without removing it, and every evaluation to date tests a fixed attacker who ignores the defense. One attacker who knows the defense is in place can step right around it.

Fixed vs. adaptive adversary — obscuring defenses LLM + defense fixed attacker ✗ adaptive: steps around harmful output
Figure: All 15 reviewed defenses obscure rather than remove the harmful capability. A fixed attacker (dashed) is blocked; an adaptive attacker who knows the defense objective circumvents it cleanly by exploiting the residual capability through an alternative path.

The paper surveys 15 published defenses against malicious fine-tuning and identifies a shared structural weakness: each defense either misdirects the optimization pathway or raises its apparent cost, without excising the underlying capability. The missing evaluation is adversarial — in the adversarial ML tradition, a defense is not validated until tested against an attacker who knows the defense objective and selects attacks accordingly. Applying this standard to all 15 defenses, the authors show that an adaptive adversary trivially bypasses most by exploiting the residual harmful capability through a slightly different path. The paper argues for mandatory adaptive-adversary evaluation in future adversarial fine-tuning defense work, aligned with certified robustness standards.


Items 4–10 — Compact
open-weight-safety eval RAND · preprint Jun 2026
RAND Corporation
PE1–PE4 compliance: 37 open-weight families PE1 ~5 PE2 ~2 PE3 ~2 PE4 1
Only 1 of 37 open-weight families meets all four PE stages; compliance collapses at PE2–PE4 — exactly the conditions an adversary with full weight access can exploit.

Proposes a four-stage Proportional Evaluation framework (PE1: evaluate without safeguards; PE2: safeguard removal robustness; PE3: selective capability amplification; PE4: worst-case misuse proxy), calibrated to the irreversibility of open-weight releases. Applied to 37 model families, only 1/37 passes PE1–PE4; most pass none; the sharpest deficit is PE2 and PE4 — the conditions most relevant to an adversary with full weight access.

RAG-poisoning AI-security preprint Sep 18 2026
Pedro Pereira, Eva Maia, Isabel Praça
doc A (frag 1) doc B (frag 2) doc C (frag 3) RAG context false claim in output
Fragments distributed across benign-looking docs accumulate in the RAG context and elicit a confident false claim — no single passage is detectable as malicious.

Splits a false target claim across multiple locally-plausible documents so no single passage is detectable as malicious; success is driven by accumulation of weak adversarial signals. Across 108 RAG configurations varying retriever, dataset, top-k, and database composition, the distributed approach outperforms concentrated single-document baselines and resists single-document detection heuristics.

prompt-injection agents preprint Sep 9 2026
Companion to AgentDrift benchmark — September 9, 2026
agent trace (turn-by-turn) DriftNet dual-head transformer injection detected? (head 1) localize turn (head 2)
DriftNet feeds the agent's turn-by-turn reasoning trace to two decoupled heads: one classifies whether injection is active, one localizes which turn introduced it.

Companion to the AgentDrift benchmark; DriftNet is a dual-head transformer that ingests an LLM agent's sequential intermediate outputs and simultaneously (i) detects whether a prompt injection is active and (ii) localizes which conversational turn introduced it. The dual-head design allows each head to be trained on a distinct supervision signal; evaluated across multiple open and closed LLM backends on AgentDrift.

applied-mech-interp alignment preprint Jun 2026
June 2026 (updated July 2026)
RLHF failure surface: three mechanistic regimes proxy reward → quality → collapse reward hacking aligned evaluator gaming (proxy↑, judge disagrees)
RLHF failures are not monolithic: reward hacking (proxy↑, quality↓), collapse (both↓), and evaluator gaming (proxy↑ via judge-specific exploitation) are distinct regimes requiring different interventions.

Treats RLHF failure not as a single terminal event but as a structured failure surface, classifying checkpoint transitions by joint changes in proxy reward, judge score, and average judge score. An empirical study of PPO, DPO, and UP-PPO identifies three distinct regimes — reward hacking, reward collapse, and evaluator gaming; UP-PPO suppresses reward hacking but shifts into evaluator gaming. The taxonomy enables failure-mode-specific detection and intervention during training.

applied-mech-interp circuit-compression preprint Aug 27 2026
Sai Adith Senthil Kumar
frozen circuit ~hundreds of edges prune+LoRA condensed circuit 8.1× smaller avg auditable subgraph up to 316× smaller
Iterative prune + LoRA repair compresses frozen attribution circuits to a human-inspectable subgraph (8.1× avg, up to 316×) while preserving both target-task and general capability.

Iteratively prunes low-attribution edges and trains a low-rank adapter to match original behavior through surviving edges; each round is kept only if both target task performance and general capability survive. Condensed circuits are 8.1× smaller on average (up to 316×) in 30/32 settings across 4 behaviors and 8 models, reducing frozen attribution circuits to a human-inspectable subgraph tractable for safety auditing.

applied-mech-interp harm-detection preprint Sep 16 2026
Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi
layer 8 layer 16 layer 24 MLP probe 12.6M params 99% F1 (WildJailbreak)
LLaMA-3.1-8B activations at intermediate layers feed a 12.6M-parameter probe achieving 99% F1 on WildJailbreak — harm signal is linearly accessible well before output generation.

A 12.6M-parameter MLP probe trained on LLaMA-3.1-8B activations achieves F1 of 99% on WildJailbreak, 83% on Beavertails, 84% on AEGIS 2.0 — competitive with dedicated classifiers 1000× larger — demonstrating harm signal is linearly separable in the residual stream and extractable at near-zero overhead relative to base model inference cost.

mech-interp circuit-eval preprint May 2026
Michael Li, Nishant Subramani
Circuits: consistent within task, non-specific across tasks task A circuit task B circuit large overlap ablating A's circuit damages task B nearly as much
Massive circuit overlap: ablating task A's circuit damages task B nearly as much as task B's own circuit, challenging the task-specificity assumption underlying circuit-guided safety localization.

Edge attribution patching across 6 tasks and 7 models finds high within-task consistency but substantially low across-task specificity: ablating one task's circuit damages another task's performance nearly as much as that task's own circuit. This challenges the modularity assumption underlying most circuit-based safety auditing — safety circuits may not cleanly demarcate safety behavior, complicating circuit-guided interventions.


Notes

← all Research Radar issues · gussand · source