Interpretability via sparse autoencoders can mask extreme fragility in the very neurons we treat as safety handles. A single clamped unit may look like a clean lever, yet the model can reroute around it without our notice. The illusion of control grows as practitioners equate “human‑readable” with “trustworthy.”

Sparse autoencoders have become the de‑facto tool for dissecting residual‑stream activations, and many recent defenses assume that clamping a identified unsafe feature will reliably suppress the corresponding misbehavior. This assumption underpins latent‑space steering, refusal‑steering, and unlearning pipelines that intervene directly on SAE latents.

Stable SAE features concentrate the predictive power, while unstable features barely move the needle. The authors report a “functional asymmetry of stable and unstable features. Stable features carry most of the reconstruction‑ and prediction‑relevant signal, whereas unstable features have weak marginal impact and are dominated by low‑frequency surface‑form triggers” [1]. Moreover, unstable features fire on average only 0.18 % of tokens versus 0.44 % for stable ones, highlighting their sporadic influence.

Despite their individual non‑reproducibility, unstable features live in a reproducible low‑rank subspace. “Decoder‑space analysis shows that unstable features are individually non‑reproducible but collectively span reproducible lower‑rank subspaces” [1], suggesting that seed‑dependent basis choices shuffle the same underlying geometry rather than generate pure noise.