AI COPYRIGHT LITIGATION: SCRAPING, TRAINING AND OUTPUTS
Generative artificial intelligence (genAI) has transformed from a promising research discipline into one of the most significant commercial technologies of the modern era.
Large language models (LLMs) now power products used by hundreds of millions of individuals and organisations around the world, enabling everything from legal research and software development to content creation and customer service. While investment and competition have accelerated, so too has a wave of copyright litigation that may ultimately shape the future of the AI industry.
For businesses, the emergence of these disputes raises important questions that extend beyond innovation. Courts in Canada, the US and the UK are increasingly being asked to consider whether various aspects of the development and operation of genAI systems engage copyright law. Although the factual circumstances vary from case to case, many of the disputes can be grouped around three distinct AI activities: scraping, training and model outputs.
Scraping refers to the automated collection of online information, training refers to the use of that information to develop AI models, and outputs are the text, images, code or other content generated in response to user prompts.
Each has been alleged to infringe copyright in various proceedings in Canada, the US and the UK, creating a rapidly developing body of litigation in which rights holders, technology companies and courts are grappling with novel applications of longstanding copyright principles. The legal risks associated with these activities are increasingly relevant to organisations developing, deploying or using AI systems.
