Microsoft Executive Calls AI Training Data Use ‘Largest Theft of Labour’
Unsealed court filings reveal Microsoft executive Brent Hecht’s concerns about AI companies using internet content to train large language models.
AI companies’ use of internet content for training large language models has come under renewed scrutiny after unsealed court filings revealed strong internal concerns at Microsoft about the practice. Brent Hecht, Microsoft’s director of applied science, described the large-scale collection of online content as an “astonishing theft” and referred to it as the largest theft of labour in human history.
Hecht’s comments emerged in documents connected to The New York Times’ copyright lawsuit against Microsoft and OpenAI. He argued that millions of people had not intended for their work to be used to train AI systems and were generally not compensated for that use.
Microsoft and OpenAI have defended their use of copyrighted material by arguing that training AI models falls under fair use. They maintain that their systems transform the material rather than simply reproducing it. News organisations involved in the lawsuit dispute this argument and contend that AI-generated answers can compete directly with original journalism.
The filings also highlight concerns about declining traffic to news websites as users increasingly obtain information through chatbots. Microsoft data cited in the case showed significantly lower click-through rates for some news domains on Copilot’s answer engine compared with traditional Bing Search.
The New York Times filed its lawsuit against Microsoft and OpenAI in 2023. Other news organisations subsequently joined the case, while Microsoft and OpenAI continue to defend their position on copyright and AI training.
