📡 RESEARCH RADAR
Daily · September 19, 2026
0 peer-reviewed · 1 preprints · 0 forum/blog
Pretraining Safety · AI Security · Mech Interp
Items 4 – 10 · Also notable






10
mech-interp preprint Sep 17 2026

9 entries removed on 2026-09-29 as repeats of earlier reports: 2609.20412 (first covered 2026-09-18), 2606.19168 (first covered 2026-08-28), 2605.02087 (first covered 2026-09-01), 2606.07970 (first covered 2026-09-18), 2605.26526 (first covered 2026-08-28), 2608.18093 (first covered 2026-08-29), 2608.25390 (first covered 2026-08-29), 2609.09420 (first covered 2026-09-11), 2609.00498 (first covered 2026-09-06).

Xeno-Interpretability: Investigating the Alien Minds of LLMs

Human-interpretable semantic space (finite descriptions) Xeno-semantic space model-native, no adequate human label xeno-representations: identifiable, causal, unlabeled Safety audits relying on human labels may miss functionally relevant xeno-features
Figure 1: Xeno-semantic space (model-native representations with no adequate human conceptual counterpart) is larger than the interpretable space; safety features may exist here, identifiable and causally manipulable, yet invisible to label-based interpretability audits.

Introduces xeno-representations — LLM-internal features that are reproducibly locatable, geometrically characterizable, and causally active, yet resist human semantic labeling. Argues that the space of model-native distinctions substantially exceeds what finite descriptions can cover, and that interpretability-for-safety audits that rely on semantic labeling systematically miss this region — a potentially important gap if safety-relevant computations occur in xeno-semantic space (Sep 17, 2026).

← all Research Radar issues · gussand · source