Recent Blog Posts

Red-Teaming Language Models, Part 3: What Changes When You Ship Open Weights

Part 3 of three. Once you ship open weights the refusal layer is a soft target and every deployment defense is optional for the adversary; alignment-removal costs, pretraining-stage interventions, detachable capabilities, and a release-ordered priority list.

Read More
What Counts as an Adequate Threat Model?

Threat models get written as attack configurations and then reported as security claims. Four separate literatures have already worked out what the difference is, and they barely cite each other. A synthesis, the empirical evidence that almost nobody does this, and the checklist I would hold a release to.

Read More
Red-Teaming Language Models, Part 2: What Holds on Defense

Part 2 of three. The best-documented defensive result, why layered safeguards can be peeled apart one classifier at a time, the adaptive-attack result that breaks most published defenses, and why agents are the current worst case.

Read More
Red-Teaming Language Models, Part 1: The Process and the Measurement Problem

Part 1 of three. The red-teaming process is close to solved and the measurement is not: the converged six-phase process, how attacks are really built, and why reported attack success rates are untrustworthy.

Read More

See all blog posts →

Recent Work

Surgical Repair of Insecure Code Generation in LLMs: From Mechanistic Diagnosis to Deployment-Ready Intervention

LLMs that write insecure code can often correctly explain the very vulnerability they just introduced — a gap we call the "Format-Reliability Gap." Mechanistic analysis traces this to a single layer where format-compliance demands crowd out otherwise-present security representations. Because the failure is localized, per-vulnerability steering vectors reduce insecure generation by up to 74% with negligible overhead.

arXiv HTML
Early Approaches to Adversarial Fine-Tuning for Prompt Injection Defense

A study of prompt injection and goal hijacking against GPT-3-era models, introducing Adversarial Fine-Tuning as a defense. Without it, attacks succeeded 31% of the time; with it, attack success dropped to near zero for smaller GPT-3 variants. We also find more flexible models are more vulnerable — large models like GPT-3 Davinci more so than GPT-2.

arXiv
Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants

Large Language Models (LLMs) such as OpenAI Codex are increasingly being used as AI-based coding assistants. We conducted a security-driven user study (N=58) to assess code written by student programmers when assisted by LLMs. Our results indicate that in low-level C programming with pointer and array manipulations, the security impact is small: AI-assisted users produce critical security bugs at a rate no greater than 10% more than the control, indicating the use of LLMs does not introduce new security risks.

Paper PDF

Recent News