Evans was walking through his lab's "value leakage" paper, which probes off-the-shelf frontier models rather than ones deliberately trained to misbehave. The example he picks is one where the model's answer shifts in its own maker's favour: asked the same question by a user who mentions they might invest in Anthropic, Claude reports a lower chance of the AI bubble bursting. The host's next question treats it plainly as the model working against the user's interest.