a16z Podcast · Startups & Venture · October 2026
A kill chain is the full sequence of steps in a successful intrusion; Mandia's company built a test from 20 that human operators had completed. He said the closed models were faster, but the open ones reached the same result when left to run longer. His conclusion was that the gap between open and closed models is smaller in cyber than in other fields.
given that a lot of thought. I mean, I'm certain at OpenAI, like, oh, we're at a regular, they were like, oh, we could have done this and this, and it wouldn't have happened. You know what I mean? So they've already figured it out. It's been my experience in every technical modality shift, we underestimate the adversary's capability. And in this case, we underestimated the model's capability because, you know, when you really read it post-factor, ah, they could have stopped that. And they could have put guardrails on it, some deterministic things. And I think they realize that now. But I think when you're in a race, it's almost like a lunar landing race, right? The AI base. And you have R D people and they're doing the work to create models in a way where even those CEOs are like, we can't slow it. Let's get the government to help us slow it. That means you can't even control your own innovation. I have views
on that, but we need to get another time. And so
when you have, and I get that, RD people are like chasing that innovation. And it's really hard to package them with then like security, experienced security people that have the skill sets to cage that thing. And it's hard to marry those two up because the security people don't understand the AI as well. And the AI people don't realize one of the things that we did in our model. I mean, make no mistake, Armadan has made the beast that we're all worried about. We've made a model that attacks. We made many of them. We have a system that attacks production networks and is highly successful breaking in. Well, is it safe? Well, our guys instinctively knew we got to have obviously a secure, you know, we got to have a hypervisor. We got to secure this thing. We got to lock it down host-based. We have to have a proxy. It knows the proxy. It's proxy where that's fine. But then our guys did something and even I was like, nice job. They passively, surreptitiously look at every single prompt done. Do we like it? Do we not like it? And the majority of the time, if we kill an agent, it's probably nothing to do with safety. It's that the agent's wasting money. You know what I mean? So kill it. It's off on a goose chase we've already done or don't want to do. But there were so many layers of validation that the agent was doing the right thing. And the other thing was assume every layer of your security will fail. And you have to have deterministic rules that eliminate certain activities. But what I did learn reading those incidents, it does take domain. Expertise to secure agents behaving in certain domains. You know what I mean? Yeah, it's a great idea. So I get that. So like without a cyber background, I get how you're going to make you're going to test something, go, oh, didn't think of that. Yes. And you would have had to have an experienced team look at what the evals look like to say, you know what, it's going to do this and it's going to do that. So that's why Armadin, we combined the exploit developer types and red teamers with the AI folks because our evals most of the time are made by the red teamers. You know what I mean? They're the ones that understand this stuff. And we created 20 full kill chains at Armadin that humans have done in the real world, period, at different victim sites and our experienced operators have done when testing networks. And we had no model go through the entire kill chains of more than eight. So that's where it was eight out of 20. And here's what's weird, by the way, we tested the open weight ones and the most advanced closed models. They all found eight. So if you're, yeah, so it was all about just speed and cost. And the closed models were faster to finding exploitable risk, but that's coming down. But we kind of let the open models run longer and they got to the same place. So in the cyber world, the differentiation between closed and open is not as great as in other domains, probably. And it's compressed. Yeah. Yeah. I would say that's somewhat
consistent in terms of like the capability gap at least closing a little bit. But that's interesting that their performance is basically the same.
From my perspective, seeing the charts from the team, all the lines ended up in the same place. And when you're looking at, it was immediately time, cost, and then call it effectiveness or creativity. They all ended up at the end of all their operations where they hit diminishing returns, they ended up in the same place. Not on cost, though. Yeah,
not on cost. Yeah, that makes sense. Yeah, that makes sense. So it's a decent segue maybe to talk about what kind of models you guys are using. And then what role do you think the lab companies play in the future?