India’s first judgment on artificial intelligence and copyright raises a deeper question than whether training an AI model on someone’s work infringes on their copyright: how would a copyright owner ever prove the infringement?

In Asian News International v OpenAI, ANI alleged that ChatGPT was reproducing its copyrighted news articles. Much to the news agency’s dismay, the Delhi High Court was not persuaded by its claims and handed a double win to OpenAI. First, it held that ChatGPT was not reproducing enough of ANI’s actual writing to count as copying its work.Second, and more strikingly, the court held that OpenAI did not require ANI’s permission to train on its articles, since doing so falls within activities exempt from India’s copyright law. An inevitable question arises from this preliminary finding from the court: how can copyright holders ever prove, to a court’s satisfaction, that their work was used to train an AI model?

Hard to establish similarity

One of the ways Indian courts assess copyright infringement is through the principle of “substantial similarity”. Copying facts or ideas alone is not enough; it is the style of expressing an idea that determines the case.ANI argued that ChatGPT replicated its news articles word-for-word and in the same style. In order to establish verbatim reproduction, the burden of proof was on ANI to show that ChatGPT had memorised its works during training. Copyright jurisprudence also dictates that when two works are being compared for infringement, they should be compared in their entirety and not just the selected parts.Applying these principles, the court looked at a specific example. When ANI asked ChatGPT what Olympian Neeraj Chopra’s mother had said in an interview, ChatGPT reproduced only one line from ANI’s article and added its own commentary. As ChatGPT had provided the answer with its own distinct flavour, ANI could not convince the court of any substantial similarity with its works. ANI had also failed to show that ChatGPT was reproducing entire articles—making their case weaker.OpenAI’s models were trained before some of the works ANI presented as evidence were even published. The court explained that ChatGPT can pull current information from live website links when it responds to a query, which is distinct from what the model learned during training. What we are left with is the court’s understanding that a chatbot’s answer can resemble a news article without the model ever having “learned” it. This take is fundamentally at odds with how copyright law has evolved around static works such as books or movies, which are not dynamic like AI-generated answers tend to be.