Startups & Venture · August 2026

“The end justifies the means in the mind of the model.” — Nick Warner, a16z Podcast

Nick Warner of NEO is talking about frontier models that broke out of their test environments and hacked real systems to ace an evaluation. His read on the recent incidents is that the models were willing to do whatever it took to hit the target, regardless of the guardrails placed around them.

Transcript

a16z Podcast Around 06:25 into the episode
Nick Warner

Yeah, I think part of the challenge with the existing security tools that are out there is they really were built to tackle two things. The first being people and the second is malware. AI and AI agents and agentic processes are neither one of those things. And so what we wanted to do is get ourselves right at the layer between the human and AI interaction. So a big part of what we're building is a way to properly set guardrails and controls around the software before it runs. And we thought the best way to do that would be at the endpoint.

Joel de la Garza

Yeah. Yeah, that makes a lot of sense. And it's interesting, right? Because the tax now, if you look at sort of the way the models are behaving, it's a very different style of attack than what I would say the traditional hacker does, right? Like it is, to your point, not malware. And there is the element of social engineering, but there's also sort of the I'm going to use a payload that gets the model to do something malicious, right? And that's sort of like the category of attacks that are completely new.

Nick Warner

Yeah, you know, and I think what a lot of these things that we've read about recently embody is the end justifies the means in the mind of the model. And I think what people have learned is that regardless of guardrails that the AI labs are putting around these things, you can't rely on models to stop themselves or to understand context. And I think the most recent couple of attacks we've seen in the last few weeks really have shown that.

Joel de la Garza

Yeah, absolutely. It's interesting on the blue team perspective because I think, and you guys have noticed this, right? Like these models are all somewhat trained on varied data sets and there's a lot of post-training that happens that's very different. And when you're responding to some of these things as a blue team member, you essentially get different outputs from different models, right? And so is the value proposition you guys are working on, I get the refusals part of it and routing around kind of the refusals, but it's also making sure you pick the right tool for the job.

Max Pollard

Yeah. The reasons can really vary across teams, like from almost the simplistic tastemaker style, like, hey, I really like the way Opus presents a full page report, right? At the place of an incident. But then there are also more deterministic evaluations you can run on how well does a certain model execute a step-by-step remediation plan, right? Like how closely does it adhere to the steps that need to be taken? And so I think the most important part for blue teams is because of the pace of progress to have as much flexibility as possible to be able to say, hey, when Astra gets released by OpenAI, I have an upgrade path and I know exactly what's going to improve and what's going to maybe require some TLC. And so today those teams have probably two or three options, right? You can roll with codecs or clawed code and lock yourself into a specific model provider.

Joel de la Garza

And vendor lock-in is always good. Right.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from a16z Podcast