01 · AI SECURITY · PEER-REVIEWED
SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignmentQingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen — ACM CCS 2026
MoE models (DeepSeek, Mixtral, Qwen-MoE) are the current production standard for frontier deployment — and they carry a safety vulnerability that dense models do not: an attacker who controls the router can steer harmful inputs away from the small subset of experts responsible for safety behaviour. GateBreaker and expert-silencing attacks exploit exactly this. SEAL closes the gap by exploiting an architectural invariant the router cannot touch.
Fig. 1 — MoE routing topology with SEAL. Routing-based attacks (red dashed) can redirect inputs away from safety-aligned sparse experts. Shared experts (always active, green) receive every token unconditionally and carry the safety anchor.