Abstract
This blog argues that Indian copyright law may not treat AI training as infringing conduct in the first place. While policy discussions proceed on the assumption that training AI on copyrighted works requires authorisation, this assumption sits uneasily with the structure of Section 14 of the Copyright Act, 1957. The blog examines how AI models process works through statistical patterns rather than expressive reproduction and situates this argument within established Indian jurisprudence. Drawing from key judicial precedents, it contends that uses falling outside the reproduction of protected expression may not trigger copyright liability at all. The blog further argues that the real copyright concern arises at the output stage, where substantially similar AI-generated outputs may attract infringement liability.
Introduction: The AI Training Puzzle and a Misplaced Assumption
In 2025, when a San Francisco Court began hearing in Bartz v. Anthropic, a deceptively simple question was raised: does training AI on copyrighted works amount to infringement? In the U.S., that issue has spawned rounds of litigation and dense academic scholarship around the idea of “fair-use”.
India’s institutional response has taken a different route. The Department for Promotion of Industry and Internal Trade (DPIIT) released a working paper proposing a licensing mechanism for AI training, implicitly assuming the use of copyrighted material as infringing. The paper sidesteps a prior basic issue: Does Indian copyright law, as presently structured, treat AI training itself as infringing conduct.
Both, the American litigation and the Indian policy paper proceed on the assumption that AI training on copyrighted works is inherently infringing. That assumption, however, fits uneasily within Indian copyright doctrine, which does not treat every interaction with a protected work as infringing. If a use neither reproduces nor communicates protected expression in a legally meaningful sense, copyright may not be engaged at all. Seen from that angle, the real question should be whether AI training even crosses the S-14 threshold.
What Section 14 Actually Protects
S-14 of the Copyright Act, 1957 defines copyright as bundle of exclusive rights including the right to sell, reproduce adapt and communicate it to the public. However, these rights are confined to acts done “in respect of a work or any substantial part thereof.” That qualifier is not a drafting flourish. It indicates where copyright control runs out. Only uses that reproduce the protected expressions or its communication implicate the S-14 interests. This structural point is crucial when we move from human readers to machine learning systems.
Unlike humans, AI models do not engage with expression while training. They process works as tokens and numerical data to identify statistical patterns across massive datasets, rather than reading or perceiving expressive content. The model does not retain readable copies of books or scripts that humans can access. It internalises statistical patterns through numerical parameters. Even tokenisation stores data only as abstract numerical representations, not as human-readable reproductions of the original work.
Copyright law has distinguished between the expression of a work and the underlying ideas, structures, or patterns from which that expression is built. Protection attaches to the expression and not the latter. Training a model operates overwhelmingly on the latter. This is not merely a point that might later be raised as a defence; it suggests that the right may not be triggered at all u/s-14.
The judicial principles that point the way
This position is anchored in principles laid down by the Indian Courts decades before Generative AI became a reality.
In E.M. Forster v. A.N. Parasuraman, the Madras High Court held that a student guidebook based on “A passage to India” did not infringe the original novel. The guide clearly drew on the plot, characters, and themes of the book, but it presented them in its own language and for an educational purpose. The Court concluded that there was no reproduction of the novel’s protected expression, and therefore, treated the question of fair dealing as irrelevant.
A similar conceptual move is witnessed in Syndicate of the Press of the University of Cambridge v. B.D. Bhandari.The Delhi High Court recognised that a work used for a different purpose and character might not amount to infringement at all, even if the same underlying material was implicated. The Court did not treat the use as presumptively infringing and then saved by Section 52. Instead, it treated it as lying outside the scope of S-14 altogether.
In Barbara Taylor Bradford v. Sahara Media Entertainment Ltd., the Court further explains the point in context of adaptations. The Calcutta High Court interpreted “alteration” in Section 2(a)(v) as incapable of being stretched to cover fundamental transformations of a work. A work that is substantially transformed cannot, merely by virtue of that transformation, be treated as a copy of the original for infringement purposes.
This approach stands in contrast to American law. Unlike American law, which broadly protects transformed or adapted works through derivative rights, Indian law lacks such expansive language. Indian courts have therefore been more cautious in treating heavily transformed or distant uses as reproductions of the original work.
Together, these cases establish that where a use neither reproduces protected expression nor serves the same character or purpose, it falls outside Section 14 itself. In such cases, infringement does not arise in the first place.
The Real Pressure Point: Outputs, Not Training
AI and copyright are not free from tension, but the real issue arises at the output stage and not during training. The reproduction right may be infringed if an AI system generates content substantially similar to a copyrighted work such that an ordinary observer would regard it as a copy. This aligns with the test articulated in R.G. Anand v. Delux Films, which focuses on whether the later work creates the impression of copying the earlier one. Courts can assess AI-generated outputs in the same manner by comparing the output with the claimant’s work and determining whether protected expression, rather than mere ideas or stock elements, has been reproduced.
What remains important is maintaining a clear distinction between training and output. A model that learns linguistic or visual patterns without storing legally recognisable copies is not inherently infringing. Infringement must therefore, be assessed case by case and output by output. Where liability arises at the output stage, the unresolved question is attribution: whether liability should rest with the developer, the user, or both.
Conclusion
The debate on AI and copyright in India cannot be resolved merely by importing foreign ideas of fair use and derivative rights. Indian copyright doctrine already draws meaningful limits around what copyright protects u/s 14. AI training, which operates through pattern recognition rather than expressive reproduction, does not comfortably fit within traditional notions of infringement. The more legally significant issue arises at the output stage, where AI-generated content may reproduce protected expression in a manner that crosses the substantial similarity threshold. In such cases, infringement must be assessed on a case-by-case basis and where liability arises at the output stage, the question shifts to the attribution of liability. Should responsibility lie with the developer who built the model, or the user who prompted the result? The task ahead for the Courts is not doctrinal expansion, but doctrinal clarity.
References
E.M. Forster v. A.N. Parasuraman, A.I.R. 1964 Mad. 331
Syndicate of the Press of the Univ. of Cambridge v. B.D. Bhandari, 2011 (185) D.L.T. 346
Barbara Taylor Bradford v. Sahara Media Entm’t Ltd., 2004 (28) P.T.C. 474 (Cal)
R.G. Anand v. Delux Films, (1978) 4 S.C.C. 118
Working Paper on Generative AI and Copyright, DPIIT, Ministry of Commerce & Industry, Govt. of India (dated December 8, 2025)




