Legal cases in the US and UK reflect how the fight to protect creators is evolving.
Two recent cases have shed more light on how US courts are beginning to approach the unlicensed training of large language models (LLMs).
The first case involved a claim by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson against Anthropic PBC for copyright infringement. The authors alleged that their books were used without permission to train the LLM behind Anthropic’s generative AI tool, Claude.
The second case, brought on similar grounds, involved 13 published authors, including Sarah Silverman, who sued Meta over the training of its generative AI model, LLaMA.
In both cases, the courts ruled in favour of the developers (Anthropic and Meta), finding that the training of LLMs on published materials without a licence could be justified under the US doctrine of "fair use".
Under US copyright law, copying protected works may be considered fair use if it serves purposes including criticism, comment, news reporting, teaching, scholarship or research. Courts also consider factors such as the commercial nature of the use, the nature of the copyrighted work, the amount used and the effect on the market value of the original work. Whether the fair use defence applies is highly fact-dependent, as emphasised by the judges in both cases.
The judges found that the use of the copyrighted books was sufficiently "transformative" to qualify for fair use. Judge Alsup, ruling in favour of Anthropic, stated: "Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different. If this training process reasonably required making copies within the LLM or otherwise, those copies were engaged in a transformative use."
Another key issue in both cases was the extent to which LLM training impacted the market value of the copyrighted works. The judges noted that most of the works in question were fiction by well-known authors. They acknowledged that AI-generated content could affect different markets in different ways.
It is arguable that UK authors currently enjoy stronger protections against the unauthorised use of their works for LLM training than their US counterparts
Continues...
In the Meta case Judge Chhabria observed: "An LLM that could generate accurate information about current events might be expected to greatly harm the print-news market. The market for certain nonfiction works – for example, books about how to take care of your garden – could be greatly diminished by the ability of LLMs to produce books on that topic. For fiction works, it might be more dependent on the author or the genre in which that author operates."
These rulings appear to be a win for developers: training LLMs on copyrighted works may be exempt under the fair-use doctrine. However, the burden remains on developers to demonstrate that their use does not harm the market for the original work. This suggests that, absent legislative change, developers may face a growing number of similar claims and will need to weigh the cost of litigation against the cost of obtaining licences.
It is also important to note the specific facts of each case. In the Anthropic case, the judge declined to rule on whether ancillary uses – such as storage and access to the books during training – also fall under fair use. This issue will proceed to a full trial.
So, how do UK fair use rules differ from the US?
While there are many similarities between UK and US copyright law, the UK takes a more restrictive approach. In the UK, developers may rely on a "fair dealing" defence for text and data mining but only for non-commercial research purposes. Fair dealing does not apply to commercial use.
A key case to watch is the ongoing claim brought by Getty Images against Stability AI, which will address whether the training of Stability AI’s model on Getty’s image library without a licence constitutes copyright infringement.
Against this backdrop, the UK government launched a consultation in December 2024 proposing that text and data mining of copyrighted works be permitted for commercial purposes – provided the works are freely accessible online or via subscription. Authors would retain the right to withdraw consent. However, this proposal faced significant pushback, particularly from prominent figures in the creative industries. The government has yet to announce a timeline for its response or for drafting new legislation.
Despite the proposed changes, it is arguable that UK authors currently enjoy stronger protections against the unauthorised use of their works for LLM training than their US counterparts. We are undoubtedly at a crossroads: developers are advocating for reform to support innovation and growth, while rights holders are calling for the current framework to be preserved, along with greater transparency obligations for developers.
