Australia's push to block AI training data may backfire—and won't work anyway
Australia wants to protect creators from AI companies mining their work — but the policy may burden local developers far more than offshore giants.
The Australian government has decided to draw a line in the sand over AI and copyright, refusing to introduce a Text and Data Mining (TDM) exception that would have allowed AI companies to train on Australian creative works without seeking permission. The intention is clear and, on its surface, defensible: protect Australian creators from having their work vacuumed up by offshore technology companies without compensation. The problem is that the mechanism chosen to deliver that protection is one the government cannot actually enforce, and the economic logic it rests on is shakier than its advocates would like to admit.
The enforcement problem is structural, not incidental
The mechanics matter here. A TDM exception would have explicitly permitted AI systems to mine text and data for training purposes, the approach adopted by the EU, Japan, Singapore, and a growing list of other jurisdictions. Australia's rejection of that path means the default position is that existing copyright law applies: if an AI company trains on copyrighted Australian material without a licence, they are, in principle, infringing. That sounds like a meaningful protection until you ask the obvious follow-up question. How do you know what data a model was trained on?
The honest answer is that you mostly do not. Training datasets for large language models are assembled at enormous scale, often from web crawls, and are rarely disclosed in detail. The government's proposed AI guardrails do require disclosure of data provenance for high-risk AI systems, which is a step toward transparency. But disclosure requirements only apply to systems operating under Australian regulatory jurisdiction. A model trained in the United States or elsewhere, fine-tuned and deployed globally, is beyond the reach of those guardrails entirely. The fence, to extend the obvious metaphor, has no gates on the sides facing out of the country.
A domestic legal position that conflicts with technical reality tends to impose costs on the compliant without constraining the non-compliant.
This is not a hypothetical gap. The dominant AI systems Australians use every day were trained before any of these regulations existed, on datasets that included material from everywhere, and there is no mechanism in the current framework to reach back and address that. Going forward, the restriction bites hardest on entities that are already operating within Australian legal reach, which means domestic AI developers and researchers face more friction than their offshore competitors. That is a notable inversion of the stated goal.
Licensing leverage requires a credible legal threat
The industry bodies backing the decision, particularly the Copyright Agency, argue that licensing is the practical path forward: AI companies should negotiate agreements with rights holders, as has happened in some sectors of the music industry. The Australian Recording Industry Association has pointed to revenue-sharing models as a viable template. That is not an unreasonable aspiration, and there is genuine precedent for it. But licensing works best when there is a credible legal threat to motivate the licensee. If enforcement is practically impossible across borders, the negotiating leverage of Australian rights holders depends entirely on the goodwill or commercial self-interest of companies headquartered elsewhere.
Copyright law has always regulated outputs, not reading
There is also a distinction worth preserving between training and output. Copyright law has always been primarily concerned with what is produced and distributed, not with the reading that precedes it. A human researcher can read a thousand books and synthesise ideas from them without infringing copyright; the infringement comes if they reproduce substantial parts of those books verbatim. The same principle, adapted for AI systems, would place enforcement pressure where it belongs: on outputs that reproduce or closely derive from protected works, not on the training process itself. That is a harder line to hold politically, because the training pipeline is visible and the output infringements are scattered and diffuse. But it is the line that actually maps to how copyright law has worked for two centuries.
The government's broader AI regulatory framework, flagged in the mandatory guardrails proposals, gestures toward transparency and accountability without yet specifying how those goals translate into enforceable rules. As The Bearing has noted, the government has moved to set the scaffolding before the detail is filled in, which is not inherently wrong but does leave significant questions unanswered about what compliance actually looks like and who bears the cost of demonstrating it.
Australia is being watched as a test case. If the licensing model produces fair deals for creators, it will be cited as proof that a hard copyright line creates negotiating leverage. If AI companies simply route their training through more permissive jurisdictions and deploy the resulting models into the Australian market regardless, it will be a demonstration of something the global experience with digital regulation has shown repeatedly: a domestic legal position that conflicts with technical reality tends to impose costs on the compliant without constraining the non-compliant. That outcome would be bad for Australian creators and bad for Australian AI development at the same time, which is a poor result from a policy meant to protect one and enable the other.
Sources
Kluwer Copyright Blog — Australia's proposed Guardrails for High-Risk AI and Copyright Law
IFRRO — Australia confirms there will be no TDM exception for AI
Semafor — Inside Australia's Experiment On The Future Of AI Copyright Laws
Frequently Asked Questions
What is a Text and Data Mining exception and why does it matter for AI?
A Text and Data Mining (TDM) exception is a carve-out in copyright law that permits AI systems to train on protected works without seeking a licence. Without one, training on copyrighted material is potentially infringing — but proving that infringement happened is extremely difficult because companies rarely disclose exactly what data their models were trained on.
Why can't Australia just enforce copyright against AI companies that use Australian content without permission?
Enforcement depends on knowing what data a model was trained on, and training datasets for large AI systems are rarely disclosed in detail. When the training itself happens offshore — in the United States or elsewhere — Australian regulatory jurisdiction does not reach it, making the legal prohibition largely unenforceable in practice.
Does blocking a TDM exception actually protect Australian creators?
The protection is weaker than it appears. The dominant AI systems already in use were trained before these rules existed, and future offshore models will be trained beyond Australian legal reach regardless. The restriction falls most heavily on domestic AI developers, who operate within Australian jurisdiction, giving offshore competitors a structural advantage.
Can licensing deals replace a TDM exception as a way to compensate creators?
Licensing is the model the Copyright Agency and industry bodies are pursuing, but it works best when rights holders have a credible legal threat to motivate the other side. If enforcement across borders is practically impossible, Australian creators' negotiating leverage depends on the commercial goodwill of companies headquartered overseas rather than on law.
How does Australia's approach compare to what other countries are doing on AI and copyright?
The EU, Japan, and Singapore have all introduced TDM exceptions that explicitly permit AI training on copyrighted works, in some cases with opt-out provisions for rights holders. Australia's refusal to follow that path leaves it with a stricter position on paper and a weaker one in practice, since the restriction cannot be applied extraterritorially.