🚨NYT Sues OpenAI Over AI Scraping, Calling It Theft
AI scraping is now officially theft
TL;DR
The New York Times has filed a lawsuit against OpenAI and Microsoft, alleging that AI scraping is theft. The case reveals internal admissions that scraping copyrighted content for AI training is illegal.
The New York Times has filed a lawsuit against OpenAI and Microsoft, alleging that AI scraping is theft. The lawsuit reveals that OpenAI and Microsoft allegedly obtained and used content by bypassing paywalls undetected, scraping millions of articles from news publishers. This case highlights the legal grey area surrounding AI training practices and the potential threat to publishers. The companies allegedly stripped copyright notices from training data, and OpenAI's mid-training datasets alone contain over 91,692 copies of works from major publications. The implications are significant for the future of AI and copyright law.

Key Points
NYT lawsuit alleges AI scraping is theft, with over 91,692 copies of works from major publications in OpenAI's datasets.
Internal Microsoft memo describes AI scraping as 'the largest theft of labor in human history'.
Microsoft's Copilot caused click-through rates for NYT's domain to drop up to 93%, creating a 'doom loop'.
OpenAI allegedly circumvented paywalls undetected, scraping millions of articles from Common Crawl.
The companies allegedly assembled Project Mango data into a training dataset containing copies of at least 160,903 unique works.
Why It Matters
The lawsuit challenges the legality of AI training practices, specifically the use of copyrighted content. If publishers win, it could reshape how AI companies operate and train their models. This affects anyone relying on AI tools that depend on large datasets, including developers and businesses using AI for content creation and analysis.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,483 builders reading daily.