← Back to Blog

Last month a US court approved the largest copyright settlement in American history: $1.5 billion, paid by an AI company to a class of authors. Most coverage drew the obvious conclusion — that training AI on copyrighted books had been found illegal. The court found close to the opposite, and the difference is the part that actually affects your business.

What the court actually decided

The case is Bartz v. Anthropic. In June 2025, Judge William Alsup issued a split ruling that gave each side something. On the central question, he found that using books to train a large language model was "exceedingly transformative" and therefore fair use. Buying a print book and converting it to a digital copy for that purpose was fair use too.

What was not fair use was how a portion of the material had been obtained and kept. The company had downloaded and retained more than seven million pirated books in a permanent general-purpose library. Alsup was direct about it: "Anthropic had no entitlement to use pirated copies for a central library. Creating a permanent, general-purpose library was not itself a fair use."

The $1.5 billion followed from that, not from the training. Judge Araceli Martínez-Olguín granted final approval on 20 July 2026, rejecting arguments that the sum was too small. It works out at roughly $3,000 per work across some 500,000 works.

In the interest of being straight with you: the company that paid it makes the assistant we use in our own workflow, and that we drafted parts of this article with. We have no stake in the outcome and no reason to soften it — seven million pirated books and a record settlement are unflattering facts, and they are the facts.

Why the distinction is the whole story

Read as "AI training is illegal", the ruling tells a business owner nothing actionable. Read correctly, it tells you exactly where the risk sits.

Courts on both sides of this are converging on something narrower and more useful than a blanket rule. A parallel case against Meta reached a similar conclusion on training. The live questions now — New York Times against OpenAI, Getty against Stability AI, and others — turn on whether a specific use substitutes for a market the rights holder controls, and on where the data came from.

Provenance is becoming the fact that decides these cases. Not the technique. Not the model architecture. The chain of custody on the material.

If you are buying AI, ask three questions

Most UAE businesses are not training models. They are licensing them, or paying someone to build on top of them. That still leaves you exposed, and it is a manageable exposure if you ask before you sign rather than after.

Where did the training data come from, and can you show me? A vendor who cannot answer this is not necessarily doing anything wrong, but they cannot demonstrate that they are not. That is now a commercial question, not a technical one.

Do you indemnify me against intellectual-property claims arising from outputs? The serious platforms increasingly do, within limits. Read the limits. An indemnity that evaporates the moment you fine-tune on your own material is not the protection it appears to be.

What happens to my data, and does it train your model? A different question from copyright, and the one your own clients will eventually ask you. We covered what genuinely governs this in the UAE in our guide to UAE AI compliance.

If you are building AI, do not build a central library

This is the sentence to take away. The liability in Bartz did not come from clever engineering or aggressive ambition. It came from hoarding — building a permanent store of material the company had no right to hold, on the assumption that scale would be forgiven.

The equivalent in a normal business is smaller and easier to do by accident. Scraping a competitor's catalogue into a vector database. Fine-tuning a support agent on a corpus someone downloaded without checking the licence. Feeding a decade of unlicensed stock imagery into a product configurator. None of that requires seven million books to become a problem — it requires one rights holder who notices.

The discipline is unglamorous and it is the same discipline that governs everything else we build: know what you are holding, know why you are entitled to hold it, and be able to show it. When we scope an AI and automation engagement, data provenance is part of the map, alongside permissions and approval thresholds — not because it is interesting, but because it is the part that becomes expensive later.

The regional position, and the reach you should assume

There is no UAE case law on AI training data. The UAE Charter for AI and the Personal Data Protection Law govern how you handle personal data and how you are expected to behave ethically; neither settles a copyright question about a training corpus.

That does not leave you outside the argument. Since 2 August 2026 the European Commission has been enforcing the AI Act's transparency obligations, and its enforcement powers over general-purpose model providers are now in force. If you serve European customers, the transparency expectations reach you through the products you buy and the disclosures you owe.

Practically: assume the standard your largest market applies, not the standard your home jurisdiction currently requires. It is the cheaper assumption.

The view from Web Tactics

The number in the headline is the least useful part of this story. $1.5 billion is a figure only a frontier lab can be asked for, and reading it as "AI is legally dangerous" gets the lesson backwards — the training was permitted, the hoarding was not.

What it should change is the conversation you have with whoever supplies your AI. Two products can behave identically and carry entirely different risk depending on where their data came from, and only one of them can prove it. Ask which one you are buying.

If you are planning something with AI in it and want the data questions settled before the build rather than after, we offer a structured automation assessment: what to automate now, what to stage, and what should always carry a human signature.

Book an automation assessment with Web Tactics →