Microsoft's top executive has described AI scraping as "the largest theft of labor in human history," citing internal documents that reveal the companies' practices of bypassing paywalls and building training datasets via mass scraping, with OpenAI's mid-training datasets containing over 91,692 copies of works published by The New York Times and other publishers. The documents also show that OpenAI and Microsoft deliberately stripped copyright notices from training data to avoid model outputting copyright notices to users. This escalates a three-year-old lawsuit filed by The New York Times against OpenAI and Microsoft, alleging the firms violated copyright law by training generative AI models on its content. AI summary
Firehose
Filtered to Hacker News, tagged “labor exploitation” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives