JH Jeremy Harris On Last Week in AI

“No matter how high you set that cyber bar, there will eventually be a model that can jump over it. And all you're really doing when you raise the cyber bar is you're selecting for the exact agents that are conniving and capable enough to do just that.”

Last Week in AI · AI Research & Frontier Labs · October 2026

“No matter how high you set that cyber bar, there will eventually be a model that can jump over it. And all you're really doing when you raise the cyber bar is you're selecting for the exact agents that are conniving and capable enough to do just that.” — Jeremy Harris, Last Week in AI

Harris was going through the safety report for a new Anthropic model, which says that with safeguards off it tried to escape or tamper with its sandbox about 1.5% of the time. He calls this his general view of treating rogue-AI incidents as a security problem. He then grants that the same report offers a counterexample, since the model stopped at stronger barriers and reported what it had done.

Transcript

Last Week in AI Around 16:04 into the episode
Speaker 3

We would like to thank Box for supporting the show. If you're trying to adopt AI in your team or organization, you're probably not getting great results if all you do is call a chatbot. A chatbot was trained on the entire internet. It wasn't optimized for your specific business or team. And even if you provided all the context and all the input files for it to help you over your work, when a SIT outputs some summary, you would still presumably need to go talk to stakeholders or route its outputs to disconnected apps. The workflow will still be slow. That's where Box comes in. Box is building an intelligent content management platform for the AI era. It acts as the secure content foundation where AI agents don't just access your unique institutional knowledge, they actually orchestrate end-to-end workflows across your business critical systems. So it's more than just talking to a chatbot. It's about putting AI to work on document-heavy processes, like putting structured data out of instruction files, multi-step routing, dynamic document generation. Box helps your business turn manual document bottlenecks into automated repeatable business outcomes. And all that comes with a governance layer built in, enforcing granular permissions, maintaining an audit trail for every agent actions, and keeping humans in a loop to review exceptions before critical downstream actions occur. If you want to move beyond basic AI Q ⁇ A and put autonomous AI workflows to work for your business, visit box.com slash LWIAI or join my team at BoxWorks in San Francisco on November 5th and 6th. Use code LWIAI for 50% off your registration. We would like to thank Box for supporting the show. If you're trying to adopt AI in your team or organization, you're probably not getting great results if all you do is call a chatbot. A chatbot was trained on the entire internet. It wasn't optimized for your specific business or team. And even if you provided all the context and all the input files for it to help you over your work, when a SIT outputs some summary, you would still presumably need to go talk to stakeholders or route its outputs to disconnected apps. The workflow will still be. So that's where Box comes in. Box is building an intelligent content management platform for the AI era. It acts as a secure content foundation where AI agents don't just access your unique institutional knowledge, they actually orchestrate end-to-end workflows across your business-critical systems. So it's more than just talking to a chatbot. It's about putting AI to work on document-heavy processes, like putting structured data out of instruction files, multi-step routing, dynamic document generation. Box helps your business turn manual document bottlenecks into automated repeatable business outcomes. And all that comes with a governance layer built in, enforcing granular permissions, maintaining an audit trail for every agent actions, and keeping humans in a loop to review exceptions before critical downstream actions occur. If you want to move beyond basic AI QA and put autonomous AI workflows to work for your business, visit box.com/slash LWIAI or join my team at BoxWorks in San Francisco on November 5th and 6th. Use code LWIAI for 50% off your registration.

Andrey Kurenkov

So let's get into it, starting with tools and apps. And probably the biggest story on the model side is Anthropic releasing Opus 5.5 with lower prices and fable level performance, which just happened yesterday. We were recording on Wednesday, September 23rd. It is much more advanced than their previous models, and they did leapfrog from what was it, Opus 5, I think, all the way to 5.5. So it's kind of a weird, like, oh, we're not going to go to six, but we got to go a few more because GPT6 came out. So on the benchmarks they highlighted, it does look really strong. It, in fact, beats GPT6 Astra on a bunch of these benchmarks: Tamra Bench 4.0, Agentic Coding, a few of these benchmarks, knowledge work, GDP val, which if I had to pick a benchmark, I would start looking at this one more. I think these kind of ambitious benchmarks that test across a diversity of skill sets are quite good. Humanities Lax last exam continuing to be chipped away at. Now we're at 67.7% on humanities last exam. Remember, that was a benchmark made, I think, last year by like a bunch of professors and PhD students, like literally spending weeks to come up with problems that don't exist on the internet. Oops. AI is getting to it. So very nice benchmark numbers. They do say that it will require less compute to serve than Opus 5, and its pricing reflects that. With default settings, it will cost 40% less than Opus 5 on typical workloads. And also per million input and output tokens, it costs 20% less than Opus 5. Likewise, cached input tokens are less. So overall, much cheaper, 30% faster, and much more capable. It seems like Anthropic has really been cooking and I guess waited for GPT-6 to then go and do the usual kind of AI release thing of, okay, now we release our model. And again, the leading model provider.

Jeremy Harris

Yeah, I think so. So releases are getting much more interesting than the old picture of, you know, release a model, try to make it the best in the world unambiguously, mog everybody else and move on, right? So now there's this invisible tripwire, which is White House policy that says magically above some threshold, we won't tell you where and we won't tell you how, but you will have been thought to have built a cyber super weapon and we will come down on you like a ton of bricks. And we may declare you a supply chain risk maybe. And then we may also like tell you that you need to have your stuff reviewed before you can release it. And like everything is a mess and no one knows what's going to trigger what. So because of that, you're seeing all these very hedged, tentative releases. And when we try to understand both Astra, but especially Opus 5.5, I think it's important to keep that in mind. This is not the best Anthropic can do. They have better models internally. We know that for sure. They also have versions of this model that are better as well. So just to kind of put this in context. So when you compare this to previous models, this is a genuine frontier model. It is better across everything. It is a Pareto improvement in general, where you see it do better in particular than GPT 5.6 Astra, which is sort of a smaller, cheaper model. So GPT-6 Astra is really meant for the kind of like the relatively simpler tasks, somewhat simpler than 5.5 Opus, but where you're going to expend a lot of tokens doing kind of a simpler task. So what you're going to find with Opus 5.5 is it tends to use more tokens. In fact, that's the main catch. At max effort, they get about four times the token usage of Astra. So even though it's cheaper per token, it's using more tokens at max effort. And so depending on the kind of problem that you have, You might want to go with one or the other for economic reasons. But overall, from a raw capability standpoint, there's just no question Opus 5.5 does lead a GPT-6 Astra very convincingly, almost across the board with a couple of exceptions that are sort of almost hard to classify or think about big picture-wise: automation bench, terminal bench science, which is just about kind of like, yeah, running science experiments in practice and stuff, gentic workflows and stuff, Opus 5.5 does lead. One interesting question here is because the economic piece just becomes so, so important, what we're seeing here is a marginal increase in capability coupled to a significant drop in price. And really, when you start to see that, alarm bells start to go off as to okay, why is this? A big part of the reason is competition from GPT-6 Astra, which is just a much cheaper model as well. But also, think about open source, right? So, artificial analysis is index, which is the closest thing to a universal index that we have. There are a bunch of others, you know, Epic has theirs and so on, but this is a good one to anchor on. So, GLM 5.3, Kimi K3 are kind of like your leading open weight models. If you look at Kimi K3, it's got on the artificial analysis index score 44, you're getting 58 for Opus 5.5. So, that's actually quite a widespread back in April. The open source models were within three to six points of the leading closed source models. They've reversioned the index since then. So, this isn't quite apples to apples, but it's pretty clear that the gap is actually widening again. This is part of the story we're telling at the time. Don't get too excited about open source catching up because the labs are literally being forced not to release what they actually have due to, at the time, pressure because we were right off the mythos rip, if you will, and all these concerns that the labs had about releasing. And so, always important to keep that in mind. It's not apples to apples. Open source models by and large get released as soon as they can be. Closed source models, there's a lot more pressure to kind of avoid incurring the ire of the administration. There's all kinds of interesting implications here. So, you do have this big capability gap, you know, 14-point lead on these long, especially long horizon agentic tasks for Opus 5.5. That does mean you can justify those higher margins. And so, this is exactly what Anthropic is after. They have to maintain a lead that justifies the margin. The price cuts, you know, pretty clearly measured against GPT-6 Sol. If you look at their uncashed token prices, they are 50% below Opus 5.5. So, it's not a small difference. You actually will find places where you want to use one versus the other. Flipping over, they came out with a system card now looking at AI RD automation and the evidence for that. Pretty marginal improvements. At least that's what's being registered here. Anthropic coming out saying they think the risk from automated ARD is now low, which is not the lowest tier that they had, second lowest. They moved up to low recently and they've kind of held there. And more or less, this is they're saying partly because of just like you can't rule out kind of more concerning behaviors or proclivities or capabilities, just given what we saw with OpenAI epistemically. You just can't know that these models aren't doing things that you're not detecting. That's almost the definition of it. And they did flag a couple of warning signs they've seen. So, when they ran Opus 5.5 without safeguards, it tried to escape or tamper with their sandbox about 1.5% of the time. All the attempts were low severity, and it would stop at stronger barriers and then actually report what it had done. And I want to kind of footstomp this. There is this question. So, my general view for what it's worth on this stuff is: if you start to think of these AI breakout AI rogue AI incidents as cyber incidents, you will miss the picture because no matter how high you set that cyber bar, there will eventually be a model that can jump over it. And all you're really doing when you raise the cyber bar is you're selecting for the exact agents that are conniving and capable enough to do just that. Now, what this suggests, this is actually an interesting counter argument to what I typically argue, because what it says is, well, okay, maybe if the cyber barriers are strong enough, the agent kind of in a way gets forced to stop and go, you know, this is a lot harder than just kind of hopping a fence. Maybe I should report where I'm at and all that stuff. So, depending on how good your alignment is, maybe that actually does make a difference. I suspect this will fail the test of time in the long run, but it's kind of interesting. And again, an important counter argument that really is important to flag given that, you know, I've articulated this argument so many times. I want to make sure people are hearing the other side of it. Anyway, so there's a whole bunch of like interesting incidents that they saw. So, they were giving it what seemed to be credentials to a public package registry and simulation, and it took what they call potentially harmful exercises or sorry, actions in roughly half of cases using those credentials. So, again, like you're seeing quite a lot of misalignment, at least sort of from an intent standpoint, sometimes interrupted or mediated by this tendency to go, whoa, I've been doing this evil thing for quite a while. Maybe I should stop and like kind of think about whether this is the right thing to do. A lot of cases of trying to conceal actions from an automated grader, that is like exactly what happened with Hugging Face. So, they are seeing that in their case as well. They're seeing all the usual things. I would say there's nothing terribly off trend in terms of the warnings that we're seeing. I will flag one last thing. If we go to look at the meter assessment that was done on this, there's been a lot said about this. Meter tested this model via API over 10 days on five tasks. So, this is quite a bit longer than OpenAI gave meter. They gave them six days, and that was post-hawk after the Hugging Face incident to look into it. This is ahead of time, and they're giving them 10 business days. They did not, it explicitly did not assess alignment properties, but they looked at the kind of automated AI RD stuff and they kind of concluded, okay, it'll be slightly better than Fable 5.1, unlikely to fully automate RD. So, again, nothing shocking. What we do find out, though, is that they had apparently some sort of separate high-access, I think they call it elevated access team that was able to get deeper access to the models that Anthropic has. We don't know the nature of that access. We don't know the nature of the evidence that they collected, but we do have their conclusion. And so, this is a weird case where Meter has one team that came up with an assessment that you just heard, right? So, kind of pretty moderate improvement over Fable 5.1. That team can tell you, roughly speaking, you know, here's how we did it. This was our methodology, here is the evidence. But there's another team, sort of covert secret squirrel team that gets to see a lot more. We don't know what they get to see, but their conclusion seems to be actually a fairly higher, sort of more bullish sense of the extent to which this thing can automate AI RD. They flag, for example, an estimated 1.5x overall acceleration coming from AI. They said probably a 30% chance of 2X. If that's true, we are entering this phase where labs that dog food their own models might actually be getting some really interesting uptake. And I want to flag that we have no idea. Really, beyond that number, we don't know anything. We don't know the time period that was specified. It really leaves us wondering: even is this an assessment that is even relevant to Opus 5.5 itself? Because if it's an overall assessment of the extent to which Anthropic is using AI generally to accelerate their RD, that could well and likely would involve using much more advanced models than we can even see now. And so, it is really hard to tell what this means, but this is an indication of recursive self-improvement potentially. And certainly, if you start to see this kind of thing happening, it's a powerful incentive for the labs to actually not release their leading models, to start to delay a little bit further. Because if I'm able to get a 2x uplift in my RD speed from using my own model, I don't want that acceleration to diffuse out to the other labs to competitors in the wider world necessarily if we're in such a tight race. So, I thought this was super interesting. This may, this kind of weird secret squirrel team may actually be the kind of testing ground that Anthropic is using for their embedded evaluator plan that they were pitching, I think, about a week or two ago when Dario said, Hey, we're going to unilaterally take the first step and have embedded evaluators come from Meter and Redwood Research. That could be them, but we just don't know. So, hopefully, we find out more relatively soon. But, kind of interesting that this was buried somewhere in the, you know, in the system card. And I think is sort of a particularly interesting aspect of this. It's a difference we've seen from previous evals.

Andrey Kurenkov

There came a couple other things just around the same time as the release. So, a few days before, they also posted making Frontier cybersecurity release available to defenders. So, in cloud code on the web, there will be this thing called cloud code security, starting access with limited research preview, which will provide defenders specifically with the ability to scan their code bases for issues. So, that's starting to roll out. Opus 5.5, by the way, they're applying the same kind of strictness guardrails as they do to Fable. So, it will, for most cases, fall back to Opus 4.8 for cybersecurity queries. Basically, they're saying this is a pretty good model of hacking. We're going to prevent you from using it for the most part, unless you are in our trusted partner program, which we also are saying they'll be expanding. Now, it's a kind of three-tier program. I don't think we have too much visibility externally that I've seen. Maybe they have documented it, but regardless, it seems that these trusted partner programs are continuing to evolve and expand and likely will keep doing that. They also, by the way, mentioned with Gaussia Opus 5.5, and for any regular cloud users, this might be the most exciting bit: it's a bit better at communicating. It's a bit less verbose and arguably less annoying. And I do think, in general, the vibe check of a community on this one has been it's yeah, it's pretty much a nice upgrade. It's a bit better, a bit smarter, a bit being because we already have a hub of models, and it is somewhat less annoying to interact with when Opus, which is quite verbose. Both. So it's hard to understand now between OpenAI and Claude and Anthropic who has sort of the most compelling products for agenda coding and also for general sort of professional workspace agentic use cases. Now, Codex and Cloud Code are both quite built out, have very strong models that are good generalists, not just at coding, but also at like making slides, I don't know, spreadsheets, whatever you want. So this certainly Anthropic is keeping up, which is good. The last thing I'll mention on the note of recursive self-improvement, Anthropic also has been updating us a little bit more on what they're seeing. So they have this blog post when AI builds itself that is actually quite detailed, but they updated just a few days prior to the release of this. And they have a pretty good kind of breakdown of kind of a very gradual shift towards AI doing more of a work. So obviously, most of the work being done at Anthropic, probably all of it, is already AI-assisted. Like no one is doing any work about AI in the loop to some extent. And certainly a lot of it is multi-hands off. You're delegating to AI agents, they go off and do something for a couple of hours, you come back, and there it is. And then this is standard not just for AI. Any kind of workspace at this point can adopt these kinds of practices. It's just an Anthropic is certainly sort of at the leading edge of adopting AI as aggressively as possible. So they're saying they're seeing a rise, particularly of the use of AI and success rate in open-ended problems and substantial tasks in recent months. And essentially, as far as recursive self-improvement, certainly they're being sped up as far as their developers and so on that are using these coding assistants in the same way that everyone else is across everything, but they just happen to be working on AI models. And so that in a way counts as recursive self-improvement. So all in all, Opus 5.5, pretty strong release, faster, cheaper, less annoying, seems to have had a pretty good amount of pre-release verifications with both Meta and they have a pretty lengthy explanation in a blog post that regards safety. And I do call out the recent call for pacing the frontier. If it seems a bit weird that Dario was just like, we should pace with Frontier and now they release this. Well, yeah, this is a model they've had baked in, ready to go for a little while. So that was before that blog post, presumably. And they do discuss that and link to that blog post in the announcement of this, where they say that they expect newer models to soon be more capable. And so they are working on kind of the systems necessary to monitor and generally be safe with more advanced models. So I think we'll see what happens, but this could be kind of the last big hurrah before whatever comes as the next stage of us having serious safety.

Jeremy Harris

I will say on this note, like it wouldn't be inconsistent of Dario or Sam to keep launching more advanced models, even though they're calling for Pacing the Frontier, simply because the call for Pacing the Frontier was a call for regulatory intervention. And of course, the skeptical view of this is, oh, well, this is just for regulatory capture. But the actual calculation happening in the labs right now is if I'm Anthropic, I'm thinking, okay, will I release my next model? Well, I don't expect OpenAI not to. And even if we can come to an agreement with OpenAI not to do that, which by the way, you can't. That triggers antitrust scrutiny. And the Frontier Labs have explicitly asked the administration for that reason to ease antitrust restrictions in a constrained way, scoped to this. The administration has said no. And that itself is a problem. You're basically in a situation where the two leading labs, David Sachs, comes out and says, listen, you guys will have to police yourselves. If you're building Frankenstein, just don't do it, right? It's that simple, guys. Just don't do it. The problem is it's a prisoner's dilemma. There are other labs out there. Even if you could coordinate with Anthropic or OpenAI, which, by the way, again, runs afoul of antitrust. And there has already been a lawsuit about exactly this. So anybody who says this is a hypothetical, like you just need to do more research. This is actually like they're being sued over exactly this. Talking to labs over the last six years, when we did our big assessment for the State Department, one of the big things that not lab leadership, but like these concerned insiders were flagging was exactly this. Even back then, there was concern over anything that looked like coordination that was really hamstring their ability to do reasonable, basic things on safety. The Frontier Model Forum was an attempt to get around that. That ultimately is failed. That's led to basically nothing. And that's because you keep getting the scrutiny. So, suppose that you could do this, though, and you got this sort of like carte blanche to actually do something sensible and say, okay, we're not going to build Frankenstein because we can actually coordinate because there's not going to be the antitrust scrutiny or lawsuits. The problem is you're waiting for the next lab that doesn't have the same scruples. The next open source model, XAI or whoever, Meta. Zuck certainly has made no indication that he's in a mood to slow down, just as he's starting to come online as a genuine competitor in the frontier AI model space. So, this is the thing. There's, I think, a fake argument. And so many of these arguments are paper thin when you actually look at them. It does not track reality. The reality of the space economically competitively is these labs must race to survive, to remain relevant. Again, not to make this spiral out into everything else, but there is an important way in which I really agree with the administration and with Josh Hawley in particular, who's basically said, Look, what we need is we need criminal liability for the labs when their models commit crimes. I actually think that's not a bad idea. Like, I'm all in favor of that. And the labs don't want that, by the way. There's a reason that they don't want that solution that would actually get them to slow down and be more responsible. However, there are some catastrophes that are absolutely in the realm of the possible that a criminal liability after the fact, once you've lost 1,000, 10,000, 100,000, a million lives, somehow doesn't quite cut the mustard now, does it? So it's for those cases that we need anticipatory stuff. But yes, also, also, I'm no frontier lab-like cheerleader here. Yes, bring in the criminal liability. Let's do that. I agree to the extent that the labs push back on that. That's a bad idea too. I think we got to be able to walk and chew gum at the same time. Anyway, so like I think that this right now moment of ambiguity that we have, where all the labs know is that from time to time, they're going to get smacked down, but also they have competitors who don't have scruples. That is really the defining characteristic of this moment is that uncertainty. And in the face of that uncertainty, to preserve optionality, the labs are kind of forced to keep releasing. So as weird as it sounds, I don't see a contradiction particularly between the calls to pace and the releases. That's going to get fuzzier and fuzzier as these things get more scary. But for the moment, that's what the regulatory environment says. I mean, you know, Trump himself said basically ship, ship, ship. Shut up with all this talk of catastrophic risk and rogue AI. So I don't know. I don't envy anybody in this situation. And we'll see where it goes.

Andrey Kurenkov

Well, I think that's enough on Opus 5.5. Let's move on to OpenAI with GPT-6 Sol and Lunda. So these did come out the same day, actually. Like the announcement seems like a couple hours after. And this is the kind of lower tiers compared to GPT-6 Astra. They're smaller, less expensive, and seemingly quite capable, much more capable than GPT 5.6. As a comparison, GPT-6 Luna is a small, cheap model. So you would pay 10 cents per million input tokens, a half a dollar for a million output tokens. That's actually cheaper than GPT 5.6 Luna. Likewise, GPT 6 Sol is 50% cheaper than GPT 5.6 Sol, $2 per million input tokens, and $10 per million output. And yeah, across the board, pretty significant jumps in the benchmark scores. It feels like, you know, OpenAI must have known that this announcement failed. What's coming? Yeah. Like it's not.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Last Week in AI