KM

Kevin Mandia

Things Kevin Says on Podcasts

Where to Find Them

Kevin Mandia writes a16z Podcast .

Recently: “Building Defense for the Agentic Era: Kevin Mandia” on a16z Podcast (October 2026).

What They Said

“We created 20 full kill chains at Armadin that humans have done in the real world. … And we had no model go through the entire kill chains of more than eight. … And here's what's weird, by the way, we tested the open weight ones and the most advanced closed models. They all found eight.” — Kevin Mandia, a16z Podcast

A kill chain is the full sequence of steps in a successful intrusion; Mandia's company built a test from 20 that human operators had completed. He said the closed models were faster, but the open ones reached the same result when left to run longer. His conclusion was that the gap between open and closed models is smaller in cyber than in other fields.

a16z Podcast · 2026-10-06 Permalink → Listen →
a16z Podcast Around 23:21 into the episode
Speaker 1

given that a lot of thought. I mean, I'm certain at OpenAI, like, oh, we're at a regular, they were like, oh, we could have done this and this, and it wouldn't have happened. You know what I mean? So they've already figured it out. It's been my experience in every technical modality shift, we underestimate the adversary's capability. And in this case, we underestimated the model's capability because, you know, when you really read it post-factor, ah, they could have stopped that. And they could have put guardrails on it, some deterministic things. And I think they realize that now. But I think when you're in a race, it's almost like a lunar landing race, right? The AI base. And you have R D people and they're doing the work to create models in a way where even those CEOs are like, we can't slow it. Let's get the government to help us slow it. That means you can't even control your own innovation. I have views

Speaker 2

on that, but we need to get another time. And so

Speaker 1

when you have, and I get that, RD people are like chasing that innovation. And it's really hard to package them with then like security, experienced security people that have the skill sets to cage that thing. And it's hard to marry those two up because the security people don't understand the AI as well. And the AI people don't realize one of the things that we did in our model. I mean, make no mistake, Armadan has made the beast that we're all worried about. We've made a model that attacks. We made many of them. We have a system that attacks production networks and is highly successful breaking in. Well, is it safe? Well, our guys instinctively knew we got to have obviously a secure, you know, we got to have a hypervisor. We got to secure this thing. We got to lock it down host-based. We have to have a proxy. It knows the proxy. It's proxy where that's fine. But then our guys did something and even I was like, nice job. They passively, surreptitiously look at every single prompt done. Do we like it? Do we not like it? And the majority of the time, if we kill an agent, it's probably nothing to do with safety. It's that the agent's wasting money. You know what I mean? So kill it. It's off on a goose chase we've already done or don't want to do. But there were so many layers of validation that the agent was doing the right thing. And the other thing was assume every layer of your security will fail. And you have to have deterministic rules that eliminate certain activities. But what I did learn reading those incidents, it does take domain. Expertise to secure agents behaving in certain domains. You know what I mean? Yeah, it's a great idea. So I get that. So like without a cyber background, I get how you're going to make you're going to test something, go, oh, didn't think of that. Yes. And you would have had to have an experienced team look at what the evals look like to say, you know what, it's going to do this and it's going to do that. So that's why Armadin, we combined the exploit developer types and red teamers with the AI folks because our evals most of the time are made by the red teamers. You know what I mean? They're the ones that understand this stuff. And we created 20 full kill chains at Armadin that humans have done in the real world, period, at different victim sites and our experienced operators have done when testing networks. And we had no model go through the entire kill chains of more than eight. So that's where it was eight out of 20. And here's what's weird, by the way, we tested the open weight ones and the most advanced closed models. They all found eight. So if you're, yeah, so it was all about just speed and cost. And the closed models were faster to finding exploitable risk, but that's coming down. But we kind of let the open models run longer and they got to the same place. So in the cyber world, the differentiation between closed and open is not as great as in other domains, probably. And it's compressed. Yeah. Yeah. I would say that's somewhat

Speaker 2

consistent in terms of like the capability gap at least closing a little bit. But that's interesting that their performance is basically the same.

Speaker 1

From my perspective, seeing the charts from the team, all the lines ended up in the same place. And when you're looking at, it was immediately time, cost, and then call it effectiveness or creativity. They all ended up at the end of all their operations where they hit diminishing returns, they ended up in the same place. Not on cost, though. Yeah,

Speaker 2

not on cost. Yeah, that makes sense. Yeah, that makes sense. So it's a decent segue maybe to talk about what kind of models you guys are using. And then what role do you think the lab companies play in the future?

Speaker names from our own diarization · position estimated from where the line sits in the episode
“The whole, let's slow down the models, we don't want cyber risk. Too late. The open models are already good enough and these things are common now.” — Kevin Mandia, a16z Podcast

Mandia was explaining how AI-led attacks differ from human ones in scale and speed. He argued that finding exploitable vulnerabilities is a structured problem that doesn't need the most advanced model. What holds criminals back for now, he said, is that they can't yet get GPU capacity anonymously.

a16z Podcast · 2026-10-06 Permalink → Listen →
a16z Podcast Around 06:09 into the episode
Speaker 1

Well, great. So the nature of them is, first off, we're getting a weird window in time where we're seeing them, but not at the same level you'd expect. I've seen nothing like what Armadan's already built in the wild, which is there's 25,000 agents on concert, all working together, doing really, really smart things without going on bizarre fishing trips. Because when you respond to an AI attack, you can tell it's AI very quickly. At least I can, because I've thought about a lot of offense. I've responded to a lot of attacks in the past that were led by humans. And a human goes to point A, then to point B, then to point C through their intrusion. AI does little things like four or five differences, but one would be it'll break into point A, then laterally move to point B. Next thing you know, it's trying to break into point A again. It's like, I get the drone swarm, but you could probably coordinate and think a little bit better. That's where it's at today. It'll get better and cleaner. But the differences are, first and foremost, the scale of what AI can do dwarfs humans, like in ways humans don't even get. So you have a scaling problem in that humans could always find only one path into a network. Yeah, they had to be selective

Speaker 2

because they had to devote their limited resources to one direct path, right?

Speaker 1

Yes. And then so scale is a challenge. Speed, ridiculous. What AI does in a microsecond would take 70 humans. They can't even do it. It's apples to oranges. And then, so what was always lacking is AI creative or effective. But when it comes to what we do, we don't need the fanciest model. We're not trying to speak 400 languages with our models and all that kind of thing. What Armadan's doing on offense is we're finding vulnerabilities, exploitable risk. That is code. That's a structured language, a structured process. Because it's structured, AI is going to be great at it. Right. Right. So I really think it's already here today. Like the whole, let's slow down the models. We don't want cyber risk. Too late. The open models are already good enough and these things are common now. It's just a matter of the minute you have anonymous availability of GPUs, you'll see far more criminal attacks. Oh, interesting. Yeah, you know what I mean? But until you can attack anonymously, and it's hard to do crime when people know your name. It's better to if you can commit a crime anonymously. Here's my tip for criminals. If you can commit a crime anonymously, that's a lot smarter than doing it with your jersey on with your name on it. And so anyway, the difference in attacks and what we're seeing now, we are at the precipice, first inning still, of AI-led attacks coming. And I think that's just because of the cost and availability of the models is not as readily available to the criminal element as it will be in the future. Yes, exactly.

Speaker 2

Okay. So you, your experience working in the security industry for 30 years, you probably saw a fair amount of nation-state attacks, right? Every day. Every day. So talk about the differences or similarities between nation-state attacks. And I use that just to say the most sophisticated, most successful, if you will, types of attacks compared to AI today and then where you think AI can be in a couple of years.

Speaker 1

So everything's going to change rapidly, right? But I can tell you nations on offense have never, in my opinion, they've never really been when you're hacking for espionage and for security reasons, you hack with what I would call kind of a sniper round. You're not spraying and praying. For the most part, modern nations on offense restrict their targeting and they go deep at very specific things, like 30 defense contractors or .mil when they go hard at that. Kind of think of it as that sniper round. With AI, I think it becomes more like a drone swarm. It becomes a little bit different in the cyber domain. And I think even modern nations are thinking, what will our protocol be? If we want to attack this company, do we swarm it and just burn tokens on it? Because AI is going to do a lot of things humans just wouldn't. So it's a little sloppier, a little louder, but it's more effective. But it's more comprehensive. That's the problem. You got it. It's more effective, probably. And so there's going to be so many things a nation's got to think through right now. And their whole doctrine will shift as the AI shift change comes. Like, how does AI change what our mission is? Do we maybe use the cyber domain differently? Do we drone swarm sometimes, snipe around other times? How do we balance the two? Does it depend on risk, target, how surreptitious we want to be? Because right now, AI is not a surreptitious action on offense unless you've done a ton of post-training. You got maybe a human in the loop really looking. out are we doing smart things because if you just go hey here's a prompt hack abc.com ai is not going to do it in a surreptitious smart way and i think even if you ask it to it's still not going to until it's been really trained and really had some human influence on it so but more generally what you're going to see is less capable attackers less technical less successful are going to appear way more successful it's the equalizer yeah because of the volume yeah yeah every when you start using models over time what it's going to be is on the defensive side we're going to say we're being attacked by these models but we're not sure who's behind them those attacks is it a nation is it a human is it and we'll have some clue but attribution will get a little difficult so that's a long-winded answer saying in the ai age it does democratize far greater expertise for attacking uh victim networks

Speaker 2

yeah so then that is a good segue back to armadin so um you have talked about the defense needs to be a great offense right yes like like the best got to train your defense with something yeah of course you got to train your defense with something and then it needs to be continuous right right so right so how does the product work and then how do you get sure a level of sophistication such that you can identify and remediate these vulnerabilities like what you're describing that are more sophisticated than a basic prompt

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In