On 1 September 2026, the US Department of Justice filed a statement arguing that training language models on copyrighted works is, in general, fair use. It was filed in the federal court in Manhattan, in the multidistrict copyright litigation against OpenAI — the case that gathers The New York Times's suit together with those of authors and publishers. According to Reuters, it is the first time the US government has taken a position in this wave of lawsuits.
The headline that travelled was “the US government said you can train AI on copyrighted content”. That is true, and it is not enough — because the filing itself draws three limits that considerably change what a company should conclude.
The three limits the headline leaves out
1. It is not a ruling, it is an opinion. A statement of interest is not binding: the government is not a party to the case and the judge is not obliged to follow it. It is the instrument a federal agency uses to put its position on the record in someone else's lawsuit. The litigation stands exactly where it stood.
2. It covers training only. This is the point almost every summary dropped. The filing addresses the use of works during training and expressly leaves two other questions open: how the data was acquired and stored, and outputs that reproduce protected expression. The second is a good part of what The New York Times alleges. In other words: even if the training argument prevails, the case does not end.
3. It does not apply outside the United States. Fair use is a feature of US law. Most other jurisdictions work with a closed list of exceptions and no general equivalent. What gets decided in Manhattan does not transplant.
The argument is economic, not moral
The reasoning is worth reading, because it is honest about what it is defending. The government argues that requiring licences to train would create an entry barrier only the largest technology companies could absorb, concentrating the language-model market. And that those fees would disproportionately benefit legacy media organisations, precisely because of the size of the archives they have accumulated. The filing rests on a January 2025 executive order on removing barriers to American leadership in AI, and adds a competitiveness and national-security argument.
On the other side, The New York Times replied that the government had sided “with a handful of trillion-dollar AI companies, at the expense of the countless American creators whose work they stole”.
The week the two biggest markets moved in opposite directions
Here is what strikes us as the most useful data point, and it only appears when you put two consecutive days side by side.
On 31 August, the European Commission classified ChatGPT as a search engine under the DSA — meaning more obligations: risk assessment, external audits, data access, fines up to 6% of global turnover. We covered it in ChatGPT is officially a search engine.
On 1 September, the US government moved the other way: less friction, in the name of competitiveness.
This is not a contradiction, they are different projects. Europe is regulating AI as information infrastructure; the United States is treating it as an industrial race. For anyone publishing content and selling services in both markets, it means the same company will operate under two logics at once, and that a content decision made by looking at only one side will get the other wrong.
What changes for anyone publishing content
There will be no cheque. That is the hardest practical conclusion and the most likely one. If the fair-use argument prevails in the United States, your text will keep becoming training material and you will not be paid for it. Planning content around licensing revenue is planning around something that does not exist today.
So the return has to come from where it always came. Being found, being cited, converting. That is exactly the AEO job: preparing the material to be the answer the AI recommends, with your name attached. The difference is that you can now say it without illusion: content is not worth what AI would pay for it, it is worth the customer who arrives afterwards.
Blocking is possible, and it has a price. You can block AI crawlers in robots.txt and signal through llms.txt. But the same block that prevents training tends to take you out of the answer — and for most small businesses, disappearing from the answer costs more than serving as training data. Anyone with a large archive negotiating licences is doing a different sum.
What AI does not have is the specific. It trained on the generic version of your sector. It did not train on your case, your measured number, your named client. That is why content with proprietary data keeps working while generic text disappears: the generic has already been absorbed.
What is still undecided
Three things, so nobody leaves here with false certainty. A filing is not a judgment and the judge may disagree. The two open questions — how the data was acquired and what the models output — may be decided against the AI companies even if training is held to be fair use. And the debate elsewhere is different, under different law and without a general fair-use clause.
What you can already do depends on none of this: finding out whether AI cites your company today, and fixing what stops it. That is what we do at our SEO agency: getting the site into a state where it can be cited, measuring whether it is, and treating content as a commercial asset rather than donated raw material.
Frequently asked questions
Did the US government rule that training AI on copyrighted content is legal?
It did not rule — it gave an opinion. On 1 September 2026 the Department of Justice filed a statement of interest in the federal court in Manhattan arguing that training language models on copyrighted works is generally fair use. A statement of interest is not binding: the government is not a party to the case and the judge is not required to follow it. The court decides, and the case is still open.
Does that mean AI can use any content from my website?
That is not what the filing says. It addresses only the use of works during training and expressly leaves two separate questions open: how the data was acquired and stored, and outputs that reproduce protected expression. Those two points are precisely what supports a large part of The New York Times's case.
Does this position apply outside the United States?
No. Fair use is a doctrine of US law. Most other jurisdictions operate with a closed list of limitations and exceptions and no general equivalent. A filing by the US Department of Justice has no effect elsewhere, and each country's debate on AI training follows its own path.
Will I be paid if AI trains on my content?
Under the argument the US government put forward, no. The explicit reasoning is that requiring licences would create a barrier only the largest companies could pay, concentrating the market. Planning content around licensing revenue is betting on something that does not exist today for anyone without an archive the size of a major newspaper.
Should I block AI crawlers on my site?
It depends on what your content does for you. You can block in robots.txt and signal through llms.txt, but the same block that prevents training tends to remove the company from AI answers. For most small businesses, disappearing from the answer costs more than serving as training data. Anyone with a large archive negotiating licences is doing a different sum.
If AI already knows everything, is publishing content still worth it?
Yes, if you change the kind of content. What AI absorbed was the generic version of your sector — and that is exactly the text that stopped working. What it does not have is the specific: your case, your measured number, your client's result. Content with proprietary data keeps bringing people in because it exists nowhere else.


