The 404 Media Podcast · September 2026
Koebler reading aloud from an internal Microsoft document unsealed in the New York Times' copyright suit against OpenAI and Microsoft, as the 404 Media hosts go through the summary-judgment filings. The same batch of documents has an OpenAI policy director describing the company as creating systems that substitute for the labor of the people who define the culture of society. The point Koebler draws out: the companies are describing model collapse in their own words.
Yeah, so they say this in a couple different places. There's this quote, basically, like the training of LLMs is quote, an astonishing theft of unprecedented proportions. And they also said that it's the largest theft of labor in human history. This quote came from Microsoft's director of applied science, who it's basically like a Microsoft executive that came from a deposition. And kind of like throughout the case or throughout these quotes, the executives at these AI companies are admitting that one, they were trained on copyrighted material, which we've known forever, but it's still notable that they are kind of saying it in court. And then the other thing that they're saying is that this is a competitive product with the New York Times and with journalists in general. They're trying to make a fair use argument. And the fair use argument, like some of the things that you consider when something is potentially fair use under copyright law is whether it's transformative, meaning whether it's like substantially different from the original source, which these AI companies have argued that LLM training is. And then also whether it competes with that original source, like whether it destroys the market for that original thing that you're stealing or taking or using or remixing. And what these executives are saying is that, yes, ChatGPT and LLMs and AI in general is making work that is substantially similar to what the New York Times is doing, that people are using ChatGPT in a way that they used to use the New York Times, meaning they're getting their news from it. And then interestingly, like there's a ChatGPT engineer in here saying that they have tried adding links to the output from ChatGPT and people don't click. Like people don't go to the original source. Satya Nadella was in here. The CEO of Microsoft was in here saying that it's a competitive product. And also Microsoft did this study that after they started adding AI to their search results, clicks to news sites and the New York Times fell by more than 90%. So, I mean, just like a huge, huge decrease. And all this aligns with what we've seen in real life. News sites are not getting clicks from Google anymore because of AI overviews, like that traffic has created, so on and so forth.
As you allude to, it's different to hear them say and they're milling it. And of course, this isn't a deposition of, you know, Google, but the same would apply most likely to Google AI summaries as well. And we've seen other reporting on that as well. So they definitely know that they are sabotaging the news industry, which ironically is the thing that they are pulling all the reporting from to then present to their own users, right? I guess just on that briefly, because you brought up sort of the media stuff, there's a bit where OpenAI admits to getting around the New York Times paywall. Now, I don't think this is probably going to be the most sophisticated thing in the world. I think anybody on the internet probably knows how to get around, you know, certain paywalls and that sort of thing. But what did they say and sort of what is the significance of that?
Yeah, they didn't explain how they did it. I mean, it would have theoretically been possible just to make a bunch of accounts, like to buy a handful of subscriptions. Obviously, there's like various ways to get around paywalls in general, but it basically quotes these OpenAI internal chats where someone at the company told OpenAI co-founder Greg Brockman that they had created, quote, a hack to get around New York Times paywall. And then Greg Brockman said, oh, nice. So it's just like evidence in there that they knew that they were doing something wrong. There's also parts in this document that I didn't quote from where Satya Nadella was like, I would have hoped that we would have understood that scraping from paywalled content is like not a good thing. I didn't include that quote because there was like a lot of context to kind of explain how he was saying that. But basically, there was various executives being like, hmm, maybe we shouldn't have taken that paywalled content. I mean, the document that I thought was most interesting and which we've kind of mentioned, but haven't gotten into what it is, is this one that's in the headline about Doom Loop. It is an internal Microsoft document that basically assesses what LLMs are doing to the industry and specifically what their strategy is doing. And so the quote is: quote, our AI content strategy has started a doom loop that will hurt the performance of our models and the entire web at the same time. It is highly unusual that an end product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain. It's basically like they're saying we are making something that is siphoning from another thing, these suppliers, these people who write articles and publish articles and books and art and all that. And what we're doing is destroying that. And there's an open AI document that also kind of touches on this from their policy director, Jack Clark, where he said that the company was, quote, creating systems that substitute for the labor of the people who define the culture of society. And then Microsoft says that it could significantly disrupt the employment. Of the very people who generate the data on which the foundation model was trained. LLMs are a product that destroys its supply chain. OpenAI called itself an existential threat to news publishers. And an OpenAI software engineer testified that no matter how prominently we show the links, users won't click.
Is there a date for those?
There's not a date. I mean, a lot of this stuff was happening around like GPT-3. GPT-3 was like many at this point, generations ago, but that's kind of the moment that ChatGPT was like became a thing and where it had this amazing growth and people started using it. And then the depositions have taken place over the last like year. So a lot of this was said within the last year, but then a lot of the court case has to do with the initial training of these models like years ago. But this sort of like nods at the idea of like, I guess, model collapse or the idea that like you destroy all the journalists, you destroy all the writers, you destroy all the human labor. Then you just have AI training on AI or AI training on really low quality stuff. And, you know, that hasn't happened yet, but it sort of is. As in, if you go search for stuff on Chat GPT or on some like Gemini, et cetera, like a lot of high-quality websites won't provide information to those models. And so you end up with like really shitty content and shitty, low-quality websites being cited in these models. And that's part of the whole like generative engine optimization game where it's like, we'll just create our own websites that will then be cited by AI and so on and so forth. And it's like, that is definitely starting to happen. The entire collapse of the system isn't happening, although, you know, there's like all these papers about how it might and how in like more controlled situations it is, like that, that sort of thing. But it's pretty, pretty bleak.
Yeah. So, I mean, listeners and viewers of this podcast probably kind of already know of this, right? And we obviously know it and other people in the media industry do, other people who just follow AI are like none of these are new ideas to them. But it is obviously significant that we're hearing this from Microsoft and Open AI people just around us out. Like, why do you think it is significant that we're now learning that they actually had internalized a lot of these criticisms as well? It's only now that we're learning them, even though they tried to not have it come out.