tech
OpenAI desperate to avoid explaining why it deleted pirated book datasets
OpenAI risks increased fines after deleting pirated books datasets.

TL;DR
- Authors are suing OpenAI, alleging that ChatGPT was illegally trained on their copyrighted works.
- OpenAI deleted two datasets, "Books 1" and "Books 2," created by former employees using data from Library Genesis (LibGen) before ChatGPT's release.
- US district judge Ona Wang ordered OpenAI to share internal communications with lawyers about the dataset deletion and any withheld references to LibGen.
- The judge ruled that OpenAI's shifting claims about the 'non-use' of the datasets and its assertion of attorney-client privilege were contradictory, potentially waiving the privilege.
- OpenAI intends to appeal the ruling, arguing confusion over phrasing rather than a flip-flop.
- The authors believe that revealing OpenAI's rationale for deletion could prove willful copyright infringement and potentially increase damages.
- Judge Wang noted that many communications in a Slack channel discussing the deletion were not privileged as they lacked requests for legal advice.
- OpenAI's defense of good faith is being challenged by its attempts to block discovery into its state of mind.
- The judge criticized OpenAI for misrepresenting a previous ruling in its defense against discovery requests.
- The outcome of this dispute could impact OpenAI's decision on whether to settle the lawsuit.