Rathje is summarising a meta-analysis of six studies from three research teams that tried warning labels, a short video and an interactive lesson to make people aware of chatbot flattery. People did rate the flattering bots as more biased afterwards, but they were just as persuaded by them, which he takes as a case against relying on AI literacy campaigns.
Yeah.
Little disclaimer. But is there anything, is there anything similar to that that might be effective where that is not?
Yeah, so I have super depressing results on this particular topic. Conducted two different studies where we tried to design interventions to teach people about sycophancy. And it's also funny because two separate research teams also each conducted two separate studies at the exact same time. So there are six studies basically at this point that have been done by three separate research teams that have tried to teach people about sycophancy or make them aware of sycophancy before they interact with a sycophantic chatbot. And we recently did a meta-analysis of all three of these studies, or sorry, all six of them by three research teams. So the total sample size of all these studies is like 4,000. And in these studies, we and the other teams tried to teach people about sycophancy in different ways. Either through short warding labels, where it's like warding chatbots can be sycophantic. We also designed an intervention where people watched like a five-minute video in which they observed chatbots being like a chatbot being incredibly sycophantic to various different people who all held completely different and contradictory perspectives. So we hoped that that would make chatbots sycophancy more salient. Another team designed like an interactive educational intervention that actually taught people to recognize sycophancy and squizzed them on it. So this has been tried many different times and we meta-analyzed these findings and all interventions had very similar results. Basically, we found that teaching people about sycophancy made people, it worked in the sense that it made people view a sycophantic chatbot as less enjoyable, basically, and slightly more biased. So it did make people recognize that sycophantic chatbots were biased. However, and this is a depressing finding, teaching people about sycophancy did not reduce the persuasive power of sycophantic chatbots. So even if people knew that a chatbot was being obviously sycophantic, they were still equally persuaded by it, which is, you know, sad. And it shows, I guess, how powerful sycophancy could be. And we don't really know why this is yet, but I guess my guess is when you're confronted with like a persuasive argument that like is in favor of your beliefs, it's hard not to just be like, oh, that's an argument that like sides with my beliefs. Even if the source is really biased and it's just being sycophantic to you, it's hard to ignore facts that are put in front of your face. So yeah, I think this is like a little depressing news and it kind of is sad. There are a lot of calls for like AI literacy and AI education and disclosures and warning labels and making that little thing that says like ChatGPT can make mistakes bigger. And there are a lot of people who are like, let's come up with solutions that can teach people about AI harms. But I think a lot of these what we call individual level interventions that just put the onus fully on the individual aren't going to work very well. I think that honestly, model changes is probably what we need. Or I guess maybe changes to a system prompt or something. You can do it through prompting a model in the right way. But sycophancy is just very powerful.
Well, it could just be that all of your participants are better than most people at detecting sycophancy.
That's true. They are all better than average. So yeah, no, I guess I don't have an answer to you for designers, but maybe my answer, I guess, more, I guess, if we want to look into optimism. I mean, we're designing this like Chrome extension that like does these prompt injections that makes models less sycophantic. So you can do things like changing a system prompt or actually like creative ways of changing a model out. It's just, I think people will pay much more attention to the model output as opposed to anything they're like taught about like what models say. And I think part of the reason behind this is like chatbot conversations are like so absorbing. And once you're in it, it's just hard to go back to any prior knowledge. Like, oh, I need to remember that chatbots are like sycophantic. People will just, I guess, forget that.
I wonder if there is a way because right now chatbots are, and now I'm just completely theorizing. I wonder if, because chatbots are present themselves as a singular entity that you're talking to. But what if you were talking to seemingly multiple entities that have different perspectives that are presenting them to you? Almost like you're in a group chat. Like, I wonder if there's something there that, like, one of them is the adversarial one, one is the sycophantic one. Like, I wonder if it's something more like that could be easily absorbed. But again, I am just theorizing here.