Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
Post-training safety measures for open-weight LLMs can be erased in dozens of fine-tuning steps. This paper shows that filtering dual-use knowledge from the pretraining corpus itself produces models that resist adversarial fine-tuning by orders of magnitude longer — and without sacrificing capability.
The authors pretrain multiple 6.9B-parameter models from scratch using a multi-stage pipeline that identifies and removes biothreat-relevant documents via topic classifiers and heuristic contamination filters. Resulting models score near-zero on biothreat proxy benchmarks and sustain that behavior under up to 10,000 steps and 300M tokens of adversarial fine-tuning — exceeding post-training RLHF/DPO baselines by over an order of magnitude. Crucially, no degradation in unrelated capabilities (coding, math, general QA) was observed, resolving the key concern that data filtering degrades model quality.