tech
OpenAI desperate to avoid explaining why it deleted pirated book datasets
OpenAI risks increased fines after deleting pirated books datasets.

TL;DR
- Authors allege ChatGPT was illegally trained on pirated books from datasets "Books 1" and "Books 2."
- OpenAI deleted these datasets prior to ChatGPT's release in 2022, claiming they fell out of use.
- The authors suspect OpenAI deleted the datasets to conceal their illegal training data.
- US magistrate judge Ona Wang ordered OpenAI to disclose communications with in-house lawyers about the dataset deletion.
- Judge Wang rejected OpenAI's claims of attorney-client privilege regarding the "non-use" of the datasets.
- The ruling allows authors to explore OpenAI's "good faith and state of mind," which could lead to increased damages if willful infringement is proven.
- OpenAI plans to appeal the ruling.
- The case draws parallels to a settlement involving Anthropic, where similar concerns about pirated book data were raised.