tech

OpenAI desperate to avoid explaining why it deleted pirated book datasets

OpenAI risks increased fines after deleting pirated books datasets.

OpenAI desperate to avoid explaining why it deleted pirated book datasets

TL;DR

  • Authors allege ChatGPT was illegally trained on pirated books from datasets "Books 1" and "Books 2."
  • OpenAI deleted these datasets prior to ChatGPT's release in 2022, claiming they fell out of use.
  • The authors suspect OpenAI deleted the datasets to conceal their illegal training data.
  • US magistrate judge Ona Wang ordered OpenAI to disclose communications with in-house lawyers about the dataset deletion.
  • Judge Wang rejected OpenAI's claims of attorney-client privilege regarding the "non-use" of the datasets.
  • The ruling allows authors to explore OpenAI's "good faith and state of mind," which could lead to increased damages if willful infringement is proven.
  • OpenAI plans to appeal the ruling.
  • The case draws parallels to a settlement involving Anthropic, where similar concerns about pirated book data were raised.