DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“I am shocked by the scale of cognitive effort that you can concentrate in such a short period of time. If you think about what 130 billion tokens are — if it was a single human thinking as a full-time job, stretched back to back, that would be a human thinking for like 4,000 years, eight hours a day, working a normal work week. So starting from ancient Sumeria up till today, a single sequential human thinking that long, concentrated in 88 hours.”

Dwarkesh Podcast · AI Research & Frontier Labs · September 2026

“I am shocked by the scale of cognitive effort that you can concentrate in such a short period of time. If you think about what 130 billion tokens are — if it was a single human thinking as a full-time job, stretched back to back, that would be a human thinking for like 4,000 years, eight hours a day, working a normal work week. So starting from ancient Sumeria up till today, a single sequential human thinking that long, concentrated in 88 hours.” — Dwarkesh Patel, Dwarkesh Podcast

Noam Brown had just described the agent swarm that solved a Millennium Prize problem — roughly 10,000 agents, 130 billion tokens, 88 hours. Rather than take the token count as a number, Dwarkesh converts it into human-lifetimes to make the compression legible. Brown's reply is that the parallelization penalty is the thing to look at before drawing conclusions from raw scale.

Transcript

Dwarkesh Podcast
Dwarkesh Patel

Today, I'm chatting with Noam Brown, who is a researcher at OpenAI. He was one of the foundational contributors to what became O1 and the reasoning models, and now he's working on multi-agent systems. Speaking of which, you guys announced last week that you solved one of the millennial price problems with a system of 10,000 different AI agents that spent 130 billion tokens over 88 hours. One of the reasons that I was interested in talking to you is I think you were the first people maybe two, three years ago, who was... thinking about how the reasoning models would allow us to see into the future because if you scale up inference compute, you can see what the base capabilities of the models will be a few years into the future. And I feel like you're in a similar position now to help us understand what future capabilities will look like, given the enormous scaling of agent sizes that we can do right now.

Noam Brown

So the way I think about it, when you plot the performance of these reasoning models with test time compute on the X axis and performance on basically any reasoning benchmark on the Y axis, you see a very clear pattern where the longer these models take to think about their answer, the better they do. And this is like a very natural thing. It's the same thing with people. If you're taking the SATs, you have five minutes to go through the entire exam. You're not going to do very well. If you have five hours, you're probably going to do a lot better. The AI models are pretty similar. And they'll spend that time doing this monologue to themselves, figuring out, going through different cases, ruling out different possibilities, building on some of their previous discoveries. The problem is that as you push that further and further, you hit a latency bottleneck. You don't want to sit around for three years waiting for a response. And so what you can do is what a lot of people do is they paralyze. They just get a team of people. If you're going to found a company, you want to get a group of people together so you can go faster. It's the same thing with these AI models that it helps to just have multiple agents working on something because they can just go faster. And so multi-agent is a way of scaling test time compute in parallel instead of purely serial. And it is like less efficient because it doesn't have, it's not like a single agent has all the context to itself, but it is like a very effective way of scaling test time compute if it's done well.

Dwarkesh Patel

Okay. I'm going to ask a bunch of naive questions because these systems, so this is an unreleased model. So we haven't publicly seen how these systems work. And so I just have a bunch of ways in which I'm like confused about like what the qualitative properties of such systems are. I am shocked by the scale of cognitive effort that you can concentrate in such a short period of time. So if you think about what 130 billion tokens are, if it was a single human thinking as a full-time job, stretched back to back, 130 billion tokens would be a human thinking for like 4,000 years, eight hours a day or something, working a normal work week. So starting from like ancient Sumeria, up till today, a single sequential human thinking that long, concentrated in 88 hours. I feel like qualitatively, that is a super important consideration. And I'm surprised that there isn't a bigger parallelization penalty that you can just have 10,000 agents collaborate. And because maybe the agents are better at collaborating than humans might be, they're going much faster, that they can actually productively collaborate at such a big scale. Or maybe they, I don't know, maybe there is a big parallelization penalty. Yeah,

Noam Brown

let's talk about the parallelization penalty and then we can talk about the qualitative stuff because the truth is that we don't have very good science on multi-agent scaling up to this kind of scale. So when we released 5.6, I think that was the first time that we had a proper multi-agent system in our models. And we actually did in the blog post show some plots of the scaling performance of multi-agent systems because we have it as an option. It's ultra-mode. And the default is four agents, but you can set that to higher. And in the plot, we show, okay, here's what the performance looks like on some benchmarks for one agent, for four agents working together, for 16 agents working together. And what you see, and it depends on the benchmark, but for some of the benchmarks, basically if you have four agents working on the problem, it is done twice as fast. So you're basically paying, because there's four agents working for half as long, you're paying a 2x more to get an answer twice as quickly. If you go to 16 agents, you see a similar pattern. It's like a little less efficient, but you continue to see that performance.

Dwarkesh Patel

Is it a linear serial time speedup or a sublinear speed up as you increase the number of parallel agents?

Noam Brown

I would say it's slightly sublinear, though it does depend a lot on the problem. So math, for example, is quite parallelizable. It's not the most parallelizable thing, but it is very parallelizable. I think web search, things like doing a deep research report where you have to look through a bunch of sources, that's extremely parallelizable. I suspect that something like writing a novel would be very unparallelizable. So you would probably not see a big benefit from having 10,000 agents working on a novel together. In the same way that you'd probably not have a big benefit from having 10,000 people work on a novel together. So the performance does depend on the domain. We do measure it up to 16 or so agents in our published blog posts. The problem is, it's very hard to push that science to like 10,000 agents because it's just so expensive.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast