Newly unredacted material in the copyright lawsuit brought by The New York Times and other publishers against OpenAI and Microsoft has exposed unusually blunt internal assessments of how generative AI training affects journalism. The material is described in the publishers’ own court brief, meaning the allegations have not been adjudicated and some underlying exhibits remain sealed.

According to TechCrunch’s account of the filing, a January 2023 memo by Microsoft applied-science director Brent Hecht called large-scale copying of news material an extraordinary theft and described it as the largest theft of labor in human history. The brief says OpenAI’s mid-training data included more than 91,000 copies of works from the Times, New York Daily News and Center for Investigative Reporting, while a Common Crawl-derived collection included more than two million documents from nytimes.com.

The filing also alleges that OpenAI and Microsoft assembled training data through multiple channels, including the Bing index and projects identified as Taxi and Mango. Project Mango allegedly contained copies of at least 160,903 works from the plaintiff publishers. The publishers further claim that OpenAI personnel discussed ways to get around the Times paywall and removed copyright notices from material before model training. Those are plaintiffs’ characterizations, and the report cautions that several quotations are presented without their full original context.

Internal discussions quoted in the brief extend beyond data collection. OpenAI executive Nick Turley reportedly warned that chatbot products could become increasingly substitutive for publishers. Microsoft documents reportedly described a risk that generative AI could disrupt the jobs of people who created the material used for training. Another internal presentation said Microsoft’s Copilot answer engine reduced click-through traffic to the Times domain by as much as 93 percent compared with conventional Bing search, framing the feedback effect on the web and model quality as a potential “doom loop.”

Microsoft chief executive Satya Nadella testified that paywalled material should be licensed for grounding or training, according to the filing. He also said that, had he known OpenAI trained on paywalled information, he would have considered invoking Microsoft’s contractual rights to demand retraining.

The disclosures sharpen the central dispute over whether using copyrighted works to train models qualifies as fair use. Market substitution and harm to the original work are among the factors courts weigh. OpenAI and Microsoft did not respond to TechCrunch’s requests for comment. The new material therefore adds evidence to the publishers’ case, but it does not itself resolve the legal questions or establish liability.