tech

OpenAI desperate to avoid explaining why it deleted pirated book datasets

OpenAI risks increased fines after deleting pirated books datasets.

OpenAI desperate to avoid explaining why it deleted pirated book datasets

TL;DR

  • Authors are suing OpenAI, alleging that ChatGPT was illegally trained on their copyrighted works.
  • OpenAI deleted two datasets, "Books 1" and "Books 2," created by former employees using data from Library Genesis (LibGen) before ChatGPT's release.
  • US district judge Ona Wang ordered OpenAI to share internal communications with lawyers about the dataset deletion and any withheld references to LibGen.
  • The judge ruled that OpenAI's shifting claims about the 'non-use' of the datasets and its assertion of attorney-client privilege were contradictory, potentially waiving the privilege.
  • OpenAI intends to appeal the ruling, arguing confusion over phrasing rather than a flip-flop.
  • The authors believe that revealing OpenAI's rationale for deletion could prove willful copyright infringement and potentially increase damages.
  • Judge Wang noted that many communications in a Slack channel discussing the deletion were not privileged as they lacked requests for legal advice.
  • OpenAI's defense of good faith is being challenged by its attempts to block discovery into its state of mind.
  • The judge criticized OpenAI for misrepresenting a previous ruling in its defense against discovery requests.
  • The outcome of this dispute could impact OpenAI's decision on whether to settle the lawsuit.