Two storied American newspapers, the Seattle Times and Newsday, have brought separate lawsuits against OpenAI and Microsoft, alleging that decades of original journalism were quietly harvested to train commercial artificial intelligence systems without consent or compensation. The cases arrive at a moment when the economics of news publishing are already fragile, and when the question of who owns the raw material of machine intelligence has become one of the defining tensions of the digital age. At stake is not merely money, but a principle: whether the labor of reporting — the interviews, the
Seattle Times, Newsday sue OpenAI and Microsoft over alleged copyright infringement
Who bears the cost of training the next generation of AI?
Why does it matter that these two newspapers specifically are suing? Are they bigger than others, or is there something about their situation that makes them test cases?
They're both substantial regional papers with significant archives and ongoing newsrooms. The Seattle Times in particular has done major investigative work. But honestly, the suit could have come from almost any major publisher. What matters is that someone finally pushed the button.
Do we know if OpenAI or Microsoft actually used their content, or is this based on the assumption that they scraped everything public online? There's a difference between "we found our articles in your training data" and "you probably used our articles."
That's the core of the lawsuit—proving the actual use. The companies don't publish detailed breakdowns of what's in their training sets, so the newspapers will have to make that case in discovery.
What happens if the newspapers win? Does OpenAI have to pay damages, or would they have to change how they build models going forward?
Both, potentially. Damages for past use, and injunctions that could force changes to future practices. But the real leverage is precedent—if a court says this is infringement, it opens the door for dozens of other suits.
The companies will argue fair use. They'll say training an AI model is transformative, that it doesn't substitute for the original articles. That's not a frivolous argument. We shouldn't assume the newspapers will win just because they filed suit.
True. But the newspapers have something the companies don't: they created the original work. The question is whether the law protects that creation when it's used at scale for commercial purposes.
What's the middle ground here? Is there a world where this gets settled?
Licensing agreements, probably. Some publishers are already negotiating with OpenAI. You pay per article or per archive access, the way news organizations license content to each other. It's not revolutionary, but it's a model that exists.
The problem is scale. Licensing millions of articles from hundreds of publishers is administratively complex and expensive. That's why the companies scraped in the first place. A court order might force them to do it anyway.
Il Polso
- The Seattle Times and Newsday have filed lawsuits against OpenAI and Microsoft, accusing both companies of systematically using copyrighted journalism to build AI products without permission or payment.
- The suits land as newsrooms already hemorrhaging advertising revenue now face the prospect that their most valuable asset — original reporting — is being used to train systems that may directly compete with them.
- OpenAI and Microsoft have defended their data practices under fair use doctrine, but have declined to disclose which outlets' content appears in their training sets or in what volume.
- The legal hinge point is whether courts will view AI training as a transformative use of copyrighted material or as straightforward commercial appropriation — a question with no settled precedent.
- Other major publishers, including The New York Times, are pursuing parallel legal or licensing strategies, signaling that this conflict is widening into an industry-wide reckoning.
- A ruling in favor of the newspapers could force AI developers to negotiate explicit licensing agreements, fundamentally altering how the next generation of language models is built and funded.
Two storied American newspapers, the Seattle Times and Newsday, have brought separate lawsuits against OpenAI and Microsoft, alleging that decades of original journalism were quietly harvested to train commercial artificial intelligence systems without consent or compensation. The cases arrive at a moment when the economics of news publishing are already fragile, and when the question of who owns the raw material of machine intelligence has become one of the defining tensions of the digital age. At stake is not merely money, but a principle: whether the labor of reporting — the interviews, the investigations, the editorial judgment — can be absorbed into a new industry without acknowledgment or return.
Two of America's established regional newspapers — the Seattle Times and Newsday — have filed separate lawsuits against OpenAI and Microsoft, alleging that their published journalism was used without permission to train commercial AI systems. Both papers are seeking damages for what they describe as systematic copyright infringement, arguing that their archives were scraped to build products from which they have seen no benefit.
The suits reflect a deepening conflict between news publishers and AI developers over training data. As companies raced to build large language models, they drew on vast quantities of internet content, including millions of newspaper articles. Publishers contend they never consented to this use and have received nothing in return for the value their work provides to these systems.
The frustration is sharpened by context: newsrooms have spent years absorbing the losses of declining print advertising, and the idea that expensive investigative and beat reporting could now fuel competing AI products has become a flashpoint for the industry. OpenAI and Microsoft have defended their practices as consistent with fair use, though neither has disclosed the precise composition of their training datasets.
The cases will likely turn on whether courts view AI training as a transformative use of copyrighted material — a key fair use consideration — or as direct commercial appropriation. The newspapers must also demonstrate measurable harm, whether through lost licensing revenue or eroded competitive position.
They are not alone. The New York Times and other outlets have pursued their own legal or licensing strategies, and some publishers have erected technical barriers against scraping. No clear precedent yet governs how courts will treat journalism in the context of AI development. The outcome of these cases may determine not only whether AI companies must pay for the content they use, but who ultimately bears the cost of building the systems that are reshaping how information itself is produced and consumed.
Two of America's most established newspapers have filed separate lawsuits against OpenAI and Microsoft, accusing the companies of harvesting their published journalism without permission to train artificial intelligence systems. The Seattle Times and Newsday, both major regional papers with decades of reporting archives, are seeking damages for what they characterize as systematic copyright infringement—the unauthorized use of their original reporting to build commercial AI products.
The legal action marks an escalation in a broader conflict between news organizations and AI developers over the question of who owns the right to use journalistic content for machine learning. As AI companies have raced to build large language models capable of generating human-like text, they have relied on vast quantities of training data scraped from the internet, including millions of articles from newspapers large and small. Publishers argue they never consented to this use and have received no compensation for the value their work provides to these systems.
The timing of these suits reflects growing frustration among media outlets watching their content fuel commercial ventures while their own business models continue to erode. Newsrooms have already faced years of declining print advertising revenue and audience fragmentation. The prospect that their reporting—the product of expensive investigative work, beat reporting, and editorial oversight—could be used to train systems that compete with their own digital products has become a focal point of industry concern.
OpenAI and Microsoft have previously defended their use of publicly available internet content as consistent with fair use doctrine, arguing that training data sourcing is a standard practice in machine learning development. The companies have also suggested that AI systems trained on diverse sources ultimately benefit from exposure to high-quality journalism. Neither company has publicly detailed exactly which news outlets' content appears in their training datasets or in what proportion.
The Seattle Times and Newsday cases will likely hinge on whether courts view the use of copyrighted articles in AI training as transformative—a key test under fair use law—or whether they consider it a direct commercial appropriation of intellectual property. The newspapers will need to demonstrate both that their content was used without permission and that this use caused them measurable harm, whether through lost licensing revenue, diminished competitive advantage, or other economic injury.
These lawsuits arrive as other media organizations, including The New York Times, have also begun exploring legal remedies or negotiating licensing agreements with AI companies. Some publishers have attempted to reach commercial arrangements with OpenAI and other developers, seeking payment for the use of their archives. Others have implemented technical barriers to prevent their content from being scraped for training purposes. The legal landscape remains unsettled, with no clear precedent yet established for how courts will treat journalistic content in the context of AI model development.
The outcome of the Seattle Times and Newsday cases could reshape how AI companies source training data and whether they must obtain explicit permission or negotiate licensing fees with publishers. It may also influence whether other news organizations pursue similar litigation or whether the industry moves toward standardized compensation models. For now, the question of who bears the cost of training the next generation of AI systems—and whether that cost should be borne by the companies building them or absorbed by the creators of the content they use—remains unresolved.
Citazioni salienti
OpenAI and Microsoft have defended their use of publicly available internet content as consistent with fair use doctrine— Company positions in prior public statements