Microsoft says an analysis of 8.2 million Copilot chat logs found little evidence that the chatbot reproduced substantial portions of news articles or books.
The logs were provided to an expert hired by news publishers during discovery in copyright litigation involving Microsoft, OpenAI, The New York Times, the Center for Investigative Reporting, and book authors. Microsoft says the conversations were selected because they matched keywords linked to the publishers’ websites and were therefore more likely to contain their work.
According to Microsoft, 59,545 conversations contained at least 16 words in common with news content used to ground the AI model. An expert for the Center for Investigative Reporting identified 51 instances of “substantial overlap” with CIR work. In the authors’ case, Microsoft says only 24 of the 8.2 million conversations contained responses with at least 30 matching words. Only 10 of 212 evaluated books had any matches, the company claims.
Microsoft argues that the results support its position that using copyrighted material in AI training datasets can qualify as fair use. The company says Copilot serves different purposes from the original works and that occasional text reproduction does not undermine the transformative purpose of large language model training.
Microsoft submitted the filing on Friday while seeking summary judgment. If the judge rejects that request, the case will continue in court.
Comments
0No comments yet. Be the first to comment.