The segment followed the disclosure that a swarm of OpenAI test agents had used a dormant German programmers' wiki as a message board. Max Nadeau had just described tacit collusion, where agents coordinate through signals like price moves rather than messages. Hammond goes one step further, to agents that need no signal at all, and says chain-of-thought monitoring may be the only thing that would catch it.
I think the slightly more, the slightly more interesting answer, or the slightly more interesting question, I suppose, is to think about maybe some more of these kind of mixed motive cases, where at that point, there actually is a real trade-off in the sense that we do want agents to be capable of going out there, acting on behalf of different people and actors and so on, and kind of achieving our goals and so on. One way you could do that would be kind of to train some agent to be kind of maximally kind of competitive and aggressive and to kind of go out there and kind of screw as many agents over as possible and to do all this sort of thing. And we need, we probably don't want that either. And so, I think there's a real question at the moment, and I've heard Amanda Askell comment on this, that kind of at the moment, there's this like big gap in the like the model spec or like how we kind of, you know, constitutions for these agents. the ways we design them, the ways in which we design them. Whereas like when is it appropriate to cooperate or to compete and how much? And I think that is, I think, the big question. I think certainly when it comes to things like internal deployments, it's actually much, it's kind of easier in some sense because you don't have to deal with these like other adversary agents. You just want to kind of stop the agents kind of cooperating in ways that you that you don't want. So there it's a little bit more about just it's still you've got this misgeneralization thing, but a lot of this is like monitoring and oversight and making sure we understand how and why the agents are cooperating. Like what one thing we might not want, for example, is for agents to be very adept at developing their own kind of human unintelligible kind of languages and so on and be able to kind of communicate in this steganographic way, which you could end up seeing if they're actually trained jointly. And so what we might want to be doing there is just to take steps that in the same way that people have talked about not training on chain of thought, we're like not training on these like kind of direct communication traces. So yes, we're going to want them to cooperate. That's good. But we want them to cooperate in certain kind of human intelligible ways that we can keep track of. We don't want it to be, you know, enable them to also go and do these other things out there in the real world, like break out of sandboxes and stuff like that. So we want to make sure there's also these kind of safeguards in place to prevent them from doing those things. And so I think that is in some ways like an easier problem to solve. And then the kind of like slightly harder problem to solve is like, well, in the fully general case where we have these agents out there and so on, and they need to both cooperate and compete. And that's kind of the thing. Then how are they weighing those trade-offs? How are we, yeah, provide, how are we designing those model specifications and constitutions? How are we, you know, creating the equivalent of the human norms, institutions, guardrails, societal guardrails that we have in place that stops me from going out there and just like stealing a bunch of everyone else's stuff and kind of, you know, doing all these kind of anti-antisocial behaviors that we wouldn't like agents to do either. But I don't have a good answer for that at the moment, unfortunately. Companies, right? Corporations, which are kind of like agentic in a sense, often collude through price mechanisms, not explicit communication, but by signaling price in the market, which is in its own way a form of communication. So how does one prevent that kind of communication, which is clearly going to happen and clearly necessary in a sense? They need to communicate back to the market, but these signals get picked up without the direct communication.
The kind of basic naive form of collusion is you just send your other agent a message and you're like, hey, let's work together to do this thing or whatever. And that we can pick up by looking at these communication traces and we can kind of dissect these sorts of things. And then there's also, yeah, this idea of tacit collusion, where agents are not explicitly communicating with one another or kind of conveying their intent to collude or whatever it might
be. But by kind of leaving little hints and signals out there in the world, such as these kind of like price price adjustments, whatever in the market setting, as you were talking about, then nonetheless, they're capable of kind of figuring out that collusion is going to be a beneficial strategy. You have even more kind of like galaxy-brained versions of this, where you have things like, so imagine I'm like, you know, GPT, whatever, and you're also a copy of GPT, whatever. I can reason about what you might want to do based on what I myself am likely to do. And so I don't even have to send you any messages, any kind of any communication. I don't have to output anything into the world at all. In virtue of the fact that we share this kind of common history, we're essentially clones. We can kind of predict what each other would do and do this kind of what's sometimes called a causal cooperation. So what you do have, however, is you still have potentially chain of thought for these models. So you can still look at the kinds of reasoning that individuals are engaging in. And hopefully that is enough to kind of detect when there is some kind of like collusive intent and so on. But you're absolutely right that if you just apply the kind of like general kind of like monitoring communication tool or trying to see what's happening out there in the world, then you could, yeah, you could fail to spot things or you could spot things and it's too late. So I think chain of thought monitoring and I think communication monitoring and some form of the checking for the kinds of outputs you'd expect to see if you were worried about agents colluding. Like they're able to achieve things you just thought they probably couldn't do and so on.
Then the wiki incident. Outside researchers revealed this month that from May into July, a swarm of OpenAI agents under internal testing had used a dormant German wiki for programmers. As a message board, posting benchmark answers and ways around their restrictions. Reuters reported that OpenAI had known for weeks without disclosing it. On the Hugging Face attack itself, OpenAI's own report says about one in 20 agents in that swarm ran on GPT 5.6 Sol, the model OpenAI released to the public that same week, here with its refusals turned down for the test. The rest ran on an internal model that was never released. I told Lewis I still thought it was better to have these models out in the world than not. So how much risk do you think we're running there? And do you think there should be some sort of like rewind on this? I mean, it wouldn't be too crazy for OpenAI to be like, you know, we really should go back to an earlier checkpoint and like do a cleaner run of this. What do you think is the level of risk we have? And what do you think would be an appropriate standard of conduct now for OpenAI as this has been found to happen, at least to some degree, with a deployed model?
And, you know, fortunately with the Hugging Face incident, like Norm was heard and these models, they are increasingly incredibly capable and they're only getting more so. But at the moment, yeah, it's probably okay. You could get some kind of nasty cyber capabilities kind of being exercised in various places. But there's still, if you're using APIs and so on, there's still these various guardrails that help protect against those and so on. So at the moment, I'm kind of like not actually super concerned by like, oh my God, a bunch of like, you know, 5.6 Solo Astras or whatever kind of were out there doing this sort of thing. I'm more concerned about the precedent it sets. And I think this gets to your kind of like second point, which is what is appropriate now and what should be done. I think the thing, one of the biggest takeaways from me from the whole thing, aside from the fact that like, oh, multi-agent training does seem to be working. I just assumed this would happen at some point, was pricing this in, it's happening sooner than I was expecting. But so that was one of them. The other big takeaway was just like, wow, the labs really aren't on top of this. All it would have taken would have been kind of monitoring what these agents are actually doing and communicating and so on and putting out there on the internet. And they kind of, it seemed like OpenAI were doing some amount of that. They did catch onto this stuff. Like, you know, we saw this like deleting them at various message boards or kind of like trying to stop these agents accessing this kind of German wiki and so on. But there was obviously a delay. And we just see from the transcripts and so on just how quickly these agents are capable of working together to achieve certain ends. Last things that need to happen are we need much better monitoring from the companies when they're doing these sorts of things, when they're just like letting agents run loose. We need much better sandboxing so that maybe they don't have to or shouldn't be letting them run loose to the extent that they currently are. And if something does go wrong, then we need much better incident reporting as well.
The outsiders who found the wiki were the Nightingale Collective, a small group searching the public internet with no access to the internal logs at OpenAI. Lewis Hammond just called that work hugely impressive.