Research Radar
Daily · September 2, 2026
0 peer-reviewed · 1 preprints · 0 forum/blog
Pretraining Safety · AI Security · Mech Interp
01
Top pick pretraining-safety preprint

Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs

O'Brien et al. — arXiv 2508.06601, Aug 2026

Post-training safety measures for open-weight LLMs can be erased in dozens of fine-tuning steps. This paper shows that filtering dual-use knowledge from the pretraining corpus itself produces models that resist adversarial fine-tuning by orders of magnitude longer — and without sacrificing capability.

Raw Corpus web text Topic Filter classifier + heuristic Filtered Data biothreat removed 6.9B Model pretrained fresh Adversarial Fine-tuning Steps to Break Safety Baseline (post-training): ~100 steps Deep Ignorance: ≥10,000 steps · 300M tokens 0 steps ──────────> 10k Order-of-magnitude improvement over post-training baselines
Figure 1: Multi-stage pretraining data filtering pipeline (top) and resistance to adversarial fine-tuning (bottom). Post-training baselines break in ~100 steps; Deep Ignorance-trained models resist ≥10,000 steps and 300M tokens of biothreat fine-tuning.

The authors pretrain multiple 6.9B-parameter models from scratch using a multi-stage pipeline that identifies and removes biothreat-relevant documents via topic classifiers and heuristic contamination filters. Resulting models score near-zero on biothreat proxy benchmarks and sustain that behavior under up to 10,000 steps and 300M tokens of adversarial fine-tuning — exceeding post-training RLHF/DPO baselines by over an order of magnitude. Crucially, no degradation in unrelated capabilities (coding, math, general QA) was observed, resolving the key concern that data filtering degrades model quality.




Items 4 – 10 · Also notable

9 entries removed on 2026-09-10 as repeats of earlier reports: 2608.13482 (first covered 2026-08-30), 2606.19168 (first covered 2026-08-28), 2608.25390 (first covered 2026-08-29), 2509.01631 (first covered 2026-07-12), 2608.14392 (first covered 2026-08-29), 2607.07368 (first covered 2026-07-17), 2607.07903 (first covered 2026-07-12), 2606.26620 (first covered 2026-07-02), 2608.02632 (first covered 2026-08-07).







← all Research Radar issues · gussand · source