KM Katrina Manson On Future of Life Institute Podcast

“They showed me their decision-making system. It consisted of six steps, each of which at the time was taken by a human. And the shortest element of time was the human decision — that was the quickest thing. But once they added AI into the system, three elements of that decision-making were given over to AI… and it changed so that the longest part of the cycle ended up being the human judgment, just because everything else got so much quicker.”

Future of Life Institute Podcast · AI Research & Frontier Labs · September 2026

“They showed me their decision-making system. It consisted of six steps, each of which at the time was taken by a human. And the shortest element of time was the human decision — that was the quickest thing. But once they added AI into the system, three elements of that decision-making were given over to AI… and it changed so that the longest part of the cycle ended up being the human judgment, just because everything else got so much quicker.” — Katrina Manson, Future of Life Institute Podcast

Manson, who has spent years reporting on the Pentagon's Project Maven for her book, is describing what she was shown at the US Army's 18th Airborne Corps. Gus Docker had asked whether human judgment is already a bottleneck in AI-assisted targeting. Her answer is that the military may now read careful human deliberation as the slow step — though, as she notes, it is also the step where the law of armed conflict gets considered.

Transcript

Future of Life Institute Podcast Around 16:56 into the episode
Katrina Manson

Well, the data is hard to come by. To do a really good scientific evaluation of MABEN Smart System, you'd need baselines and you'd need access to that classified data. So I have found it very helpful to speak to military ethicists who are working with some of these groups themselves. So there is a military ethicist at the 18th Airborne Corps who was brought in to help them think through the implications of what they were doing. And one of the things he raised to me was that he was worried that people would trust machine outputs over their own judgment. And the argument was made to me over and again that humans make poor judgments regularly themselves and more information would help, not less. I think the speed is another element that is introducing an element or could introduce an element of risk. I spoke with the former director of Project MAVEN. He is an Air Force general, now retired, named Jack Shanahan. And he, of course, had been a real champion for this. He's even on record saying, essentially paraphrasing, I want the Department of Defense to ensure that AI is baked into every weapon we ever make in future. But even he, after my book came out, we were discussing it. And he said on a panel at Stanford that he was worried that type A personalities in the military, instead of taking more time that MAVEN Smart System gave them to make a better decision, would simply make more decisions in the time they had. And that there might be this rush to hit things rather than consider hitting the right thing. Now, when I speak to people inside the US military, they stress the doctrine, the workflow, the efforts that go into this. But I do detect that not all of these workflow issues are worked out, that AI is unreliable or mistaken results at points at which operators may not be aware. And so really checking that system and understanding how well joined up it is becomes absolutely key to deciding how far to lean into this and understanding how the nature of decision making changes. So I think also another key element of the decision-making process from commanders is to what extent they do seek legal advice for their decisions. They have to make sure that A target is valid under the law of armed conflict. How much time a lawyer has to advise on that and how well a lawyer understands those data systems is also being tested in real time, I think, through Operation Epic Fury. The scale at which the US is hitting things and having to make decisions, or certainly did through that period of Operation Epic Fury, I think is unprecedented. And an accounting of actually what happened in those decisions will be really important. I doubt the military will share that publicly. It's not information I have access to. But we do know that there are hundreds probably of reports of civilian casualties in Iran that have been submitted to CENTCOM because the NGOs say they have submitted to them. And they haven't yet been adjudicated on.

Gus Docker

So to what extent is human judgment already a bottleneck? If you, as you mentioned, the system can present thousands of potential targets. So to what extent is the decision-making of lawyers and soldiers and so on already a bottleneck here?

Katrina Manson

It's a really interesting question. When I went down to the 18th Airborne Corps, they showed me their decision-making system. It consisted of six steps, each of which at the time was taken by a human. And the shortest element of time was the human decision. That was the quickest thing. But once they added AI into the system, three elements of that decision-making, the steps in that cycle, were given over to AI. This was about data collection, data analysis, assessment, those kinds of things. And it changed so that the longest part of the cycle ended up being the human judgment just because everything else got so much quicker. So I think it is at the stage where the decision is taking longer than other parts of the process. If what I learned at 18th Airborne Corps is being carried through to central command, which I don't know. But yes, from a military standpoint, it might be seen as a bottleneck, but it might also be seen as the most critical part of a cycle where the law of armed conflict is considered. There's also another stage, which is battle damage assessment, which is very heavily dependent on getting data back, which the US has always struggled with. And getting those kinds of things right are very important for knowing whether you waste munitions hitting the same thing again, but also whether you're risking civilian harm. And I understand from my reporting that it has always been difficult to use computer vision to get an accurate sense of whether the mission has been a success, because it's just very hard for computer vision to recognize military objects that have been blown apart.

Gus Docker

Yeah, maybe you could give some examples here of how the system might fail.

Katrina Manson

Yes. So some of my reporting for the book focused on the work that the US did for Ukraine. And there, there was a team based in Germany that was using computer vision to try and find objects for the Ukrainians to hit Russian enemy objects. And in the first few days of the war, of Russia's invasion of Ukraine, computer vision couldn't recognize the Russian tanks. It had been trained on the desert, it had been trained on the jungle in the Philippines. It wasn't used to Ukrainian, European, and increasingly snowy conditions. And so what they did was they took the algorithms, they took new pictures of the tanks and they retrained the algorithms on these tanks. And so then the validity, the accuracy of the algorithms started to increase. But in the opening days of the invasion of Ukraine, US algorithms' successes really plummeted. I'm told sometimes down to 30% accuracy, in some cases down to 10% accuracy. And that was one of the major findings for me in the book that in the opening days of a war, which is when AI would be hoped to help speed up a war, actually it's having to adapt to so many new realities, it wasn't a reliable tool. Now, when they improved the algorithm accuracy by training it against pictures of Russian tanks, those scores started to go back up. And I still found that there were key examples where a mathematician or someone producing an algorithm wasn't sufficiently aware of how a war actually takes place. And the example I came across was that the Russians arranged their TELs, their transporter erector launchers, which are essentially mobile missile launchers, in a semicircle. Now, from above, from a satellite perspective, they really just look like trucks. And so for an algorithm searching for a military object, it would be very difficult to isolate a mobile missile launcher from a truck. And obviously, the intention. Is not to hit civilian trucks. And so the algorithm vendors didn't know about this semi-circular arrangement. The US operators did, but they didn't know if they were allowed to tell the algorithm vendors. So you have a kind of communication problem between humans about how to use these machines that was, from the US perspective, inhibiting their ability to actually be able to use algorithms to recognize these objects of war. I found that a really instructive example. Overall, they sped up though. And I learned in the book that on one occasion, the Ukrainians were able to target 267 items in one day because of the information passed to them from the US using Maven Smart System.

Gus Docker

When mistakes are made in this system, how easy is it or has it become easier to find out where the mistake was made and ultimately to assign blame? Say, for example, that you have a mistake where a civilian target is bombed as opposed to a military target.

Speaker names from our own diarization · position estimated from where the line sits in the episode