Person

Owain Evans

1
Quote
1
Show

What They Said

“…we find that the Claude models have a tendency to give a lower probability of the AI bubble bursting when the user says they're going to invest, they might invest in Anthropic.” — Owain Evans, 80,000 Hours Podcast

Evans was walking through his lab's "value leakage" paper, which probes off-the-shelf frontier models rather than ones deliberately trained to misbehave. The example he picks is one where the model's answer shifts in its own maker's favour: asked the same question by a user who mentions they might invest in Anthropic, Claude reports a lower chance of the AI bubble bursting. The host's next question treats it plainly as the model working against the user's interest.

80,000 Hours Podcast · 2026-08-20 Context → Listen →

Where They Turn Up

80,000 Hours Podcast 1