10
Xeno-Interpretability: Investigating the Alien Minds of LLMs
Introduces xeno-representations — LLM-internal features that are reproducibly locatable, geometrically characterizable, and causally active, yet resist human semantic labeling. Argues that the space of model-native distinctions substantially exceeds what finite descriptions can cover, and that interpretability-for-safety audits that rely on semantic labeling systematically miss this region — a potentially important gap if safety-relevant computations occur in xeno-semantic space (Sep 17, 2026).