JT

Jaan Tallinn

Things Jaan Says on Podcasts

Where to Find Them

Jaan Tallinn writes The Logan Bartlett Show . They have also been a guest on Manifold (3 times) , Future of Life Institute Podcast (2 times) , Summation with Auren Hoffman , The Generalist , FT Tech Tonic , The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis , The AI in Business Podcast and The Trajectory .

Recently: “Jaan Tallinn would like us to survive” on The Generalist (October 2026); “AI Billionaire on Existential Risk: Jaan Tallinn” on Manifold (May 2026); “Jaan Tallinn: AI Risks, Investments, and AGI — #59” on Manifold (May 2024); “Jaan Tallinn - The Case for a Pause Before We Birth AGI (AGI Destinations Series, Episode 2)” on The Trajectory (April 2024); “EP 72: Jaan Tallinn (Co-Founder, Skype) on Lessons from Skype, Giving SBF $100M & Investing in AI” on The Logan Bartlett Show (July 2023); “EP 72: Jaan Tallinn (Co-Founder, Skype) on Lessons from Skype, Giving SBF $100M & Investing in AI” on The Logan Bartlett Show (July 2023).

What They Said

“We are pulling the trigger on this Russian roulette with the planet. And every shot that doesn't kill us will make us stronger potentially and wealthier and better off. But when do you stop pulling the trigger? At one point, you know that it will be one too much.” — Jaan Tallinn, The Generalist

Tallinn has just said he thinks the current way of building AI is fundamentally unsafe, while granting that today's models are a net positive. His point is that nobody knows which generation will escape control, so each success makes it harder to stop. He argues for stopping now and finding a more controllable approach.

The Generalist · 2026-10-08 Permalink → Listen →
The Generalist Around 30:28 into the episode
Jaan Tallinn

I think that it is just fundamentally unsafe. The current paradigm is fundamentally unsafe. Myself, I would have stopped like earlier, which like in retrospect, that's a mistake. I think the current models are net positive. Like you never know what generation we're going to escape control and then you just lose.

Mario Gabriele

Because you sort of think we may have... Yeah, it's not a game we're going to have many chances to learn from.

Jaan Tallinn

Exactly, exactly. So we are pulling trigger on this Russian roulette with the planet. And every shot that doesn't kill us will make us stronger potentially and wealthier and better off. But when do you stop pulling the trigger, right? At one point, you know that it will be one too much. So yeah, I would say it's just like stop pulling the trigger right now and figure out what is a better better more. Controllable approach to AI in general. And there have been suggestions now. Yoshio Benjio has this idea of scientist AI that is deliberately trying to tease out agency from AI. So it's like kind of like principled approach, how you can make non-non-agentic AI.

Mario Gabriele

Okay, I haven't read about that. That's sort of the idea of, you know, like a drugged tiger in some way. Yeah,

Jaan Tallinn

it's basically like Oracle done properly, where you can ask Oracle. The big problem with like naive Oracle doing naively is that Oracle still has preferences. It prefers giving answers that come true. But if it has just this preference, it is incentivized to make sure that it's going to be asked simple questions, which means that it's incentivized to mess with the world. However, Yoshio's approach, the way I understand it, is that doesn't have that flaw. It basically truly, the Oracle doesn't care what will be done with the answers.

Mario Gabriele

And is this an AI that is explicitly non-agentic, just sort of something you literally visit as an Oracle and ask for advice rather than something that can do things for you?

Speaker names from our own diarization · position estimated from where the line sits in the episode
“We are not designing those AIs. We are selecting them. Just like evolution ... which means that we are selecting based on outer behavior rather than based on the inner motivations. If a child goes, I didn't take the cookie, this is outward behavior. And there could be multiple motivations why this child is saying that. One is that they just really want to tell the truth. The other is that they don't want to be punished.” — Jaan Tallinn, The Generalist

Tallinn is explaining what he sees as the central problem with how AI is made today: models are grown and then picked for how they perform on tests, not built to a design. The cookie example is his way of showing why passing a test says little about what a system actually wants.

The Generalist · 2026-10-08 Permalink → Listen →
The Generalist Around 08:47 into the episode
Jaan Tallinn

yeah steve and hundred My friend Steve O'Hondra, he wrote a paper like 15, 20 years ago called AI Drives, now it's also called Omohundra drives, where he basically makes the point that almost regardless what kind of goals you have, there are so-called instrumental goals that are very useful towards reaching any given goal. And these are like resource acquisition, generally power acquisition. Power basically means options, optionality, a lot of protection of your existence. As Stuart Russell keeps saying, that you can't fetch the coffee if you're dead. So even simple tasks require you to continue existing. And then protection of your goals. So it turns out it's really hard to change a goal of a determined agent because that's what it's about in some sense.

Mario Gabriele

And that sort of final piece is, I think, probably the counter to the question about why would these agents or an AI system necessarily have this expansionist bent. I think of the Bezos divine discontent. This is almost demonic discontent where it sort of wants more and more. Is that how you would sort of think about the counter to that? Yeah,

Jaan Tallinn

mostly. It's like getting more resources, getting more power is just good for almost any goal. There are very few goals. And the current pressing ahead, I think the really big problem with the current paradigm of how we grow AIs and we're not building them, we're growing them, is that we are, in some ways, kind of recapitulating the evolution. We are not designing those AIs. We are selecting them. Just like evolution was selecting them. In some ways, you can think of it, we create like millions or if not billions of instances of AI. Then we pick the ones that kind of do the things that perform on some given test, which means that we are selecting based on outer behavior rather than based on the inner motivations. And if you think about it, you have children, right? So if a child goes like, I didn't take the cookie, like this is outward behavior. And think about there could be multiple motivations why this child is saying that, right? One is that they just really want to tell the truth, right? The other is that they don't want to be punished.

Mario Gabriele

Yes.

Jaan Tallinn

And so therefore, the first order situation or the first principal situation is that the current paradigm, we are sort of getting a random motivation, as was demonstrated by the Hugging Face attack, that the motivation was really weird and alien, even though they performed at least an adjacent thing that they were selected for.

Mario Gabriele

Yes. Let's talk about the OpenAI Hugging Face episode. What did you make of it? Were you surprised?

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In