Fine-tuning memorization claims challenged due to flawed measurements

Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright

Computers and Society

Summary

Some researchers claimed that fine-tuning AI models lets people copy big parts of copyrighted books by memorizing the text. This paper reviews that claim and finds big problems: the tests used to detect copying were too weak, and the way they asked the AI to show memorized text might have accidentally given away the answers. The authors say the original results don't prove that copying really happened, and important control experiments are missing. Also, the costs of these leaking risks weren't reported, which matters for legal discussions on copyright.

What this means in practice

  • For legal teams: Evaluate the reliability of claims about AI model memorization when assessing copyright infringement cases.
  • For ai safety auditors: Review and improve methods for detecting memorization and data leakage in AI model fine-tuning.

A position paper. It proposes an approach and reports no results.

Authors

A. Feder Cooper

Abstract

After careful review, I'm confident the headline fine-tuning memorization results in Alignment Whack-a-Mole use an invalid measurement procedure. The book memorization coverage metric these headline results depend on counts sequence matches far shorter than what field standards consider valid evidence of memorization, and the prompting procedure used to elicit memorization runs the risk of leaking the text being "extracted" in the prompt. The paper doesn't include the negative-control experiments needed to see how much the results are inflated by false positives: claiming extraction success (and therefore memorization of training data) when matches between generations and training data may be due to other factors. Given these validity issues, the paper's claims that fine-tuning lets users extract substantial portions of copyrighted books, in a form that could substitute for the originals, aren't supported by the reported results. The failure to report the experiments' cost (an important component of the threat model) further compromises the copyright claims. I'm writing this note because, in the last month, (prospective) plaintiffs have reached out to me to ask about this paper. They're looking to cite this work as valid evidence in support of claims in ongoing and potential future copyright litigation.