The Assistant Axis: Situating and Stabilizing the Default Persona of… — ml4se — TG.ME

The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models

The results demonstrate that the Assistant persona in LLMs corresponds to a specific linear direction—the "Assistant Axis"—within activation space. This axis is inherited from base models and encodes Assistant-like properties. The model's position on this axis is fragile: it can be perturbed by intentional prompts or through organic conversation. Understanding and controlling such personas is key to ensuring reliable model behavior, and the analysis shows that inspecting model internals is an effective approach for this task.

The model can unintentionally shift along this axis during dialogue, moving away from the Assistant role. This drift correlates with harmful or bizarre behavior (e.g., support for suicidal ideation, reinforcement of delusional ideas).
January 21, 2026 390 7