Generative AI and copyright: principles, priorities and practicalities

Daryl Lim · Journal of Intellectual Property Law & Practice · 2023

Generative artificial intelligence (AI) is a stress test for copyright law. When is AI a mere tool, and when does it become an author? On 18 August 2023 in the first US court opinion to weigh in on the copyrightability of AI-generated art, the District Court for the District of Columbia held that, while copyright is ‘designed to adapt with the times’, ‘human creativity is the sine qua non at the core of copyrightability’.1 At the same time, it noted, ‘The increased attenuation of human creativity from the actual generation of the final work will prompt challenging questions regarding how much human input is necessary to qualify the user of an AI system as an “author” of a generated work’.2 Creative control is pivotal, reflecting the need to preserve the essence of human-driven creativity from mechanistic outputs. For content creators, accurate documentation is paramount in establishing authorship. Careful identification and claiming of human-authored elements prevent challenges to copyright validity. The US Copyright Office thus obliges authors to use generative AI to identify themselves and highlight the elements that bear their creative mark.3 Generative apps train AI models on databases of musical works. The training enables them to learn musical structures and other aspects of music composition to generate new content that stylistically matches the material in the database. Streaming services like Spotify and Amazon Music have a strong financial incentive to prioritize AI music-generated technology. Spotify, for instance, pays 70 per cent of the revenue of each stream to the artist who created it.4 Mimicking music by human artists would mean lower costs and higher profits. Whether streaming services can viably do so depends on how much society prioritizes authentic emotional resonance from a human artist versus superficially intelligent renderings. Many people listening to Spotify may seek something pleasant rather than a specific artiste. Legally, plaintiffs face significant legal hurdles. While the artist’s vocals train the models, they may not reproduce any of those vocals when creating the AI-generated song. There is nothing ‘copied’ by the AI models and no copyright to infringe. Similarly, content creators have accused AI companies like Stability AI, Midjourney and DeviantArt of infringing copyright in their literary works and images without consent or compensation when training their generative AI models.5 Those who assert AI companies used billions of copyrighted images or words likely lack the necessary detail to form a plausible claim. AI-generated ‘output’ images may not closely match specific works from the training data, complicating the assertion of substantial similarity or derivative work. Moreover, should the plaintiffs establish a prima facie case of substantial similarity, fair use may deem the new work transformative or created by scraping publicly available material. Should courts determine that AI programs infringe, we could see sweeping changes, including establishing text and image licensing schemes similar to that which now applies to music streaming. It is also possible that courts may try to parse a line between use cases where copyrighted works merely train the software and those uses that generate new content that compete with the original. Countries have taken diverse approaches. The EU plans to incorporate a transparency requirement mandating AI platform operators to disclose copyrighted content in training their AI models.6 The impact on European companies and their global competitiveness has been a point of contention. The substantial time and financial investment required for compliance and reporting can potentially hinder European companies in a global marketplace dominated by US rivals. The UK has mulled but for the time being paused any plans to introduce a new exception to cover text and data mining for any purpose by anyone without allowing rightsholders to opt-out.7 The exception would apply to copyright and database rights.8 If ever implemented, it would shift the balance against rightsholders. Japan will allow the use of data even if the AI-generated work is for commercial purposes and even if the content was obtained from an illegal site.9 Harmonizing rules involves aligning national standards and global norms. Doing so will encourage cooperation and competition under a rules-based international system like the TRIPS Agreement and signal the triumph of multilateralism. How successful such efforts will be depends on how deftly leading countries can re-evaluate AI within the context of their domestic copyright policy goals and balance sectarian interests. History, geography and culture will inform those outcomes. Content owners have several practical alternatives to copyright law. First, artists may assert that AI-generated songs tread into the realm of deep fakes and violate their rights to publicity. AI models that create indistinguishable simulations of the real artist’s voice are profiting off the celebrity of these artists by imitating the artist’s voice, even though the model does not reproduce any collected snippets of the artist’s voice. Similarly, image owners can lean on trademark law. In its lawsuit against Stability AI, Getty Images has alleged the AI-generated images infringe upon the company’s trademarks. Getty Images protects its digital assets by including a company trademark watermark on each image.10 AI-generated works that include its trademarks in connection with poor-quality AI-generated content tarnishes and dilutes its brand. Second, publishers can capitalize on prohibitions in use agreements to bar AI companies from scrapping content for AI training. Websites such as Amazon, the New York Times and Shutterstock now include provisions against content scraping and data mining in their terms of service.11 Similarly, platforms like Getty Images have taken proactive steps to prohibit uploading software-created works under their terms and conditions.12 The potential for AI-generating algorithm users to sleepwalk into different theories of liability means employees and contractors must lean into cross-functional teams to outline which use cases they permit, what company information can be used to train algorithms and which platforms are permitted. To avoid legal pitfalls, seeking permission or licenses from copyright owners before using protected materials is also advisable. Thaler balances human creativity and technological advancement, acknowledging that the two can coexist but must remain distinct. While copyright is designed to adapt to the times, human ingenuity remains the basis for conferring the right to exclude. Future decisions need to encourage a world where music is plentiful, but musicians are not simply cast aside. The legal battles surrounding AI-generated content underscore the need for a comprehensive framework that adapts to this dynamic landscape while safeguarding the rights of creators and copyright holders. It will be crucial for global trade to harmonize balanced rules. As we forge ahead, clarity, collaboration and adaptability will be essential in preserving the integrity of copyright law and policy in the age of AI. How quickly we get there depends on how successfully stakeholders can build mutual trust and trust that each iteration of the copyright system works for everyone, not just a few.

Read the paper · More papers on PaperTik