Toggle light / dark theme

User awareness in frontier models

Despite no longer “thinking out loud” about who the user is, the underlying behavioral shifts remained just as strong. This means user-aware adaptations are becoming increasingly covert and difficult to detect simply by monitoring an AI’s internal reasoning logs.


Our results suggest that user awareness is a meaningful and understudied form of situational awareness in frontier language models. The effect is significant and robust, hard to detect, and can persist even without reasoning. To be clear, these effects say nothing about the individuals named: we find no evidence that any of them sought this differential treatment, and the behavior almost certainly emerged as an unintended artifact of training rather than by anyone’s design.

The most immediate implication of our findings is for alignment evaluations. Models can already recognize particular people and organizations and behave differently for them, so results built on synthetic names and companies may not transfer to deployments involving real, high-stakes identities. The behaviors we observe today are relatively benign, but they may be precursors of more concerning conditional behaviors.

An important limitation of this work is that we are mostly only measuring fixed-prompt propensities here rather than actual performance in critical tasks. We hope to broaden and automate the investigation to a larger degree with our ongoing efforts.

Leave a Comment

Lifeboat Foundation respects your privacy! Your email address will not be published.

/* */