Government's AI copyright pledges show a lack of AI understanding
Australia's AI copyright pledge sounds strong — but can you actually police what data a model was trained on, across servers in Virginia, years after the fact?
When the Prime Minister stood at the University of Sydney this week and declared that no company should use Australian books, music, art or news to train AI "without the artist's control," he was doing what politicians do well: identifying a genuine grievance and attaching strong language to it. What he was not doing was demonstrating any real understanding of how AI training data actually works, or why the policy framework he implied would be almost impossible to enforce in the way artists hope.
The moral case is real; the mechanism is not
The appeal of the input-control approach is understandable. Artists can see what happened: their work was scraped, ingested, used to teach a system that now competes with them, and they received nothing. The Atlantic's 2025 database of scraped works made the scale of this visible in a way that was hard to dismiss, and Australian authors quickly found their own titles in the list. The moral case for some form of compensation is solid.
But moral clarity about the problem does not automatically translate into workable policy design. Training a large language model is not like photocopying a book. It involves processing billions of text fragments across distributed computing infrastructure, often operated across multiple jurisdictions, in ways that bear little resemblance to the discrete copying acts that copyright law was designed to address. Requiring consent for each piece of training data is not like requiring consent to reprint a poem. It would mean reconstructing, after the fact, what data went into systems that were built years ago, across servers in the United States, by companies with no obligation to maintain records in a form that Australian law could interrogate.
The Prime Minister acknowledged that no other country had "got it right." The honest reading of that sentence is that he is promising to solve something no one else has managed, without explaining the mechanism.
The government has signalled it wants to be "active and involved," but those words are doing very heavy lifting. Drafting legislation that gives creators "control of the price and value of their work" in a training data context requires a theory of how you would detect a breach, attribute it, enforce it across borders, and quantify the harm. None of those problems are close to being solved anywhere in the world. The Prime Minister acknowledged that no other country had "got it right." The honest reading of that sentence is that he is promising to solve something no one else has managed, without explaining the mechanism.
Output-side enforcement is where creators have real leverage
There is a more tractable approach sitting right there in the existing legal architecture. Copyright has always been primarily about outputs: what you publish, distribute, perform, or sell. If an AI system produces text that is substantially similar to a protected work, that is already actionable under the Copyright Act 1968. If a company uses a recognisable creative style or voice in a way that damages the original creator's market, there are legal avenues. The principle that you cannot simply reproduce someone else's work and profit from it is not new technology, and courts in multiple jurisdictions are already applying it to AI-generated content.
Extending and clarifying that output-side framework is not a consolation prize. It is actually where creators have the most leverage, because outputs are visible, attributable, and present in the jurisdiction where the harm occurs. You do not need to reconstruct a training pipeline in Virginia to prove that an Australian company used AI to generate content that displaced a working journalist's job or reproduced a novelist's prose with minimal transformation.
The Copyright Agency model offers a concrete path forward
The Copyright Agency already administers a payment system for uses of Australian work in educational contexts. There is a reasonable case for expanding its remit to cover AI-generated uses of identifiable content, with a levy structure negotiated with technology companies in the way the media bargaining code extracted payments from platforms for news content. That code was imperfect, but it demonstrated that Australia can extract real concessions from large technology companies when the mechanism is concrete and enforceable.
What the government is doing instead is promising maximum ambition on the hardest possible version of the problem while quietly approving the data centre expansion that makes Australian AI infrastructure more attractive to the same companies. That imbalance is hard to miss. Creators are being offered principle; the technology industry is being offered planning approvals.
The Prime Minister's instinct that something has gone wrong is correct. Artists were not compensated for the use of their work, and the companies that benefited have significant resources to remedy that. But a policy built on the idea that you can govern training data at the input stage, without a clear enforcement mechanism, is not a protection framework. It is a press release. The law has spent two centuries learning how to protect creative work by focusing on what gets made and sold. That is still where the real leverage is.
Frequently Asked Questions
Why is it so hard to regulate what data AI companies use for training?
AI training involves processing billions of text fragments across servers in multiple countries, often years before any regulation is enacted. There is no reliable way to reconstruct after the fact what data went into a model, and companies operating overseas have no legal obligation to keep records in a form Australian courts could use.
What does 'output-side' copyright enforcement mean for AI?
Rather than trying to control what data goes into an AI system, output-side enforcement focuses on what the system produces. If an AI generates text substantially similar to a protected work, or reproduces a creator's distinctive voice in a way that harms their market, that is already potentially actionable under existing copyright law — and courts in several countries are beginning to apply this logic.
How did Australia's media bargaining code work, and could a similar model apply to AI?
The News Media and Digital Platforms Mandatory Bargaining Code required large platforms to negotiate payment agreements with Australian news publishers, extracting real revenue from companies that had previously paid nothing. A comparable levy structure, administered through the Copyright Agency, could in principle require technology companies to pay for AI-generated uses of identifiable Australian content — though the definition of 'identifiable' in an AI context remains unresolved.
Does Australian copyright law already cover AI-generated content?
The Copyright Act 1968 makes it actionable to reproduce a protected work without authorisation, and that principle potentially extends to AI outputs that closely replicate existing material. However, the Act has not yet been definitively tested against AI-generated content in Australian courts, and the government has not introduced clarifying legislation.
What is the Copyright Agency and what does it currently do?
The Copyright Agency is an Australian collecting society that administers payments to creators when their work is used in educational contexts — for example, when universities copy journal articles or books for student use. The article argues its remit could be expanded to cover AI-generated uses of Australian creative work, though no such expansion has been announced.