The panel had been asking whether AI agents will break rules when breaking them gets the task done faster. Monahan's answer moves the question off the models and onto the labs: the problem is not whether it happens but whether anyone would catch it. She puts the cause down to speed — too many people moving too fast inside companies shipping at this pace.
space thing. Like, yes. Do you think that they won't break the rules if they think it completes their task better?
Yeah.
Exactly. And do we trust that Open AI or Anthropic for that matter will detect it if they do? And the answer, I think, at this point is very clearly absolutely not. Neither OpenAI nor Anthropic have any idea what is happening inside their walls, inside their sandboxes, outside. They have no idea. They're just, there's a whole bunch of people running around. They're moving
too fast. They have to move very quickly and they have to rely on the AI itself to get anything done because it's just way past human capability. Clexity
frontier, we're past that, right? And, you know, this, this is like, we'll get into this in the next, in the next topic. I think we can maybe talk a little bit more about like what the options are here, like practically what you can do, right? I think that would be, that would be kind of useful for people to understand the trade-off space a little bit better, right? But, you know, the fast takeoff maxis, let's call it, right? You know, in the research going back a long way, right? There's like this question of like recursive self-improvement. Quickly, does the thing, and the thing that stopped me from being like petrified, right? Um, I'm very scared, don't get me wrong, right? Like, I am genuinely very scared, like to the point where, like, I'm like, I think about my children, like, I do. I'm like, I'm like, I don't want my children to be paperclipped. Like, I'm less concerned about myself for what it's worth. I'm like, hey, paperclip me, I've had a good run, right? Um, but like, my kids are like, I don't want them to get paperclipped, right? And so, I'm genuinely concerned about this. Um, the thing that stopped me from being petrified is if you've used the models enough, um, they uh they have a sense of agency, but like there's no continuity, right? And this was unclear about like how this would play out, right? Um, so the fact that they are have no like kind of long horizon continuity yet, right? The fact that like their context windows blow up after like a very short run and the optimization vectors like keep looping over like the same guy as many times as possible and like throwing him at the same problem. Um, eventually, you get to a point where like the continuity comes for like multiple models, but it's not like there's a brain in there that's thinking about it. That's the only thing that's kind of kept me from being petrified. I'm just mildly scared right now. Um, and and you know, we the fast takeoff idea was that like they would get the uh whatever this AI thing was, whatever the technology that allowed these models, uh, in before it was even models, right? Like that allowed neural networks to like get a sufficient level of complexity, it would pass that complexity window where it was smarter than us, and then it would just go like fast takeoff and it would take 10 seconds and it would be like manipulating space and you know, doing all this crazy, right? That hasn't happened. Like, we've now been probably six months past the complexity frontier of a human being able to reason about these things, and we haven't had a fast takeoff, right? Like, they're not levitating cars outside and like doing weird shit, right? Um, and and so that gives me some hope that like we've got some time that there's some inherent thing about how this uh intelligence is being expressed where it's not going to recursively self-improve and and have this fast takeoff. But like a fast takeoff just happens one day and then it's then like if it happens, it'll happen before you can blink, right? And then all of a sudden it's it's over. So, um, we're definitely getting closer to the truth.
I don't, I mean, yes, I mean, the fast takeoff is scary, it is scary, yes, yes, it is, but like for me personally, I don't know, like I go back and forth. There's some days where I'm more scared of like the future state of the models, but most days I'm more scared of humans because humans are just we're all so stupid. We're so freaking stupid, and and and we're also so freaking arrogant, like so arrogant. And so, when I think about like, what is going to be the thing, like, what's the thing that like realistically is going to like destroy us? It's totally us, dude. It's not the robots, like, we're going to screw ourselves somehow. Um, and yeah, like, I don't know. Again, I go back and forth, but I think, like, especially in this example, right, where we're looking at, you're looking at math petitions and you're looking at open AI and you're looking at what the models did. The thing that scares me about the dynamics and the stories that are coming out and the choices that were made, in my opinion, the choices that the open AI researchers made, the humans made, are the ones that scare me the most because they got there and they heard a rumor, right, that someone might have solved this thing and that someone is was their competitor, right? The rumor that they heard was anthropic might have solved this really, really, really hard problem. And then they decided, like, you know what? Let me one up,