Author: Ikshika
College: Bharat College of Law, Kurukshetra University
LinkedIn : https://www.linkedin.com/in/ikshika-2a052440b
To The Point
Can AI-generated works be copyrighted in India, and does training generative models on copyrighted data violate Indian law?
When we cut through the academic rhetoric and look directly at the statutory text of the Copyright Act, 1957, two core realities emerge: “Indian law does not recognize autonomous AI as an author”, nor does the statute contain an explicit, text-and-data-mining exemption for commercial AI training.
This legal friction presents a two-fold problem for modern legal practice:
1. The Ownership Vacuum (AI Outputs) :
If an AI platform generates a piece of artwork, software code, or literature with only generic prompt engineering, the work exists in a legal void. Under Section 2(d)(vi) of the Act, a computer-generated work is protected only if we can attribute creation to “the person who causes the work to be created.” Where human input is limited to basic text prompts, the human user lacks sufficient creative control over the specific expression produced by the machine. Consequently, purely machine-generated outputs risk being denied copyright registration, placing them directly into the public domain where anyone can freely duplicate or monetize them.
2. The Ingestion & Infringement Risk (AI Inputs) :
Developers building Large Language Models (LLMs) and generative systems rely on massive scraping of copyrighted text, images, and proprietary media. Ingesting this data onto local or cloud servers triggers the copyright holder’s “Exclusive Right of Reproduction” under Section 14 of the Act. While technology companies in the United States rely on flexible “Fair Use” defenses(17 U.S.C. § 107), India operates under a strict, closed-list “Fair Dealing” regime under Section 52.
However, Indian judicial interpretation is actively evolving to meet this challenge. In the landmark interim decision in ‘ANI Media (P) Ltd. v. OpenAIInc.’ (2026), the High Court of Delhi evaluated whether storing retrieved online data to train LLMs or utilizing Retrieval-Augmented Generation (RAG) constitutes copyright infringement or falls under protected fair dealing under Section 52(1)(a). Applying the “doctrine of updating construction,” the court held on a prima facie basis that scraping and storing publicly available literary works to train AI models can qualify as fair dealing for research and private use, provided there is no verbatim memorization, regurgitation, or market substitution in the generated outputs.
While this interim order provides temporary relief for AI developers operating in India, the substantive questions remain open for trial, leaving the long-term interaction between statutory copyright and generative AI in a state of dynamic legal transformation.
Use Of Legal Jargon
To analyze this intersection of technology and law like an Indian legal practitioner, we must examine several fundamental statutory provisions and judicial doctrines:
• Modicum of Creativity: The legal standard established by Indian courts to measure originality. Moving away from the old English “sweat of the brow” doctrine—which rewarded mere labor and financial expenditure—Indian law demands independent human skill, labor, and a minimal degree of creative judgment.
• Statutory Authorship under Section 2(d)(vi): The specific provision defining the author of a computer-generated work as “the person who causes the work to be created.” In generative AI, identifying who “caused” the work—the prompter, the algorithm designer, or the platform owner—is the central point of statutory interpretation.
• Closed-List Fair Dealing vs. Open-Ended Fair Use: Unlike the open-ended judicial discretion permitted under US copyright law, Section 52 of the Indian Copyright Act provides an exhaustive, closed list of statutory exceptions (such as private research, criticism, review, and news reporting). Adapting this closed list to commercial data mining requires careful judicial construction.
• Exclusive Right of Reproduction (Section 14): The economic foundation granting copyright owners sole authority to store, copy, reproduce, or adapt their works in any material form, including digital storage on cloud and machine-learning servers.
• Doctrine of Updating Construction: An interpretive principle utilized by constitutional courts allowing old statutory provisions to be interpreted in light of current technological realities that lawmakers could not have explicitly foreseen when the statute was last amended.
• Three-Step Test under TRIPS: Bound by international treaty obligations like the Berne Convention and TRIPS Agreement, domestic exceptions to copyright must be confined to certain special cases, must not conflict with a normal exploitation of the work, and must not unreasonably prejudice the legitimate interests of the rights holder.
THE PROOF
The necessity for statutory reform and judicial clarity is demonstrated by real-world market conflicts, administrative inconsistencies, and ongoing litigation:
1. Administrative Ambiguity in the Copyright Office :
The practical difficulty of applying Section 2(d)(vi) under existing laws was made obvious when the Indian Copyright Office initially issued a registration co-naming an AI software tool (“RAGHAV Painting App”) as a co-author alongside a human artist. Recognizing the fundamental contradiction with the statutory requirement of human authorship, the Copyright Office subsequently issued a withdrawal notice. This administrative back-and-forth highlights how difficult it is for registration authorities to evaluate software involvement without clear legislative guidelines.
2. High-Stakes Litigation by News Agencies and Creative Industries :
The conflict between content creators and AI developers has moved directly into Indian courtrooms. Media organizations, publishers, and trade bodies—including news agencies like ANI, the Federation of Indian Publishers, the Digital News Publishers Association, and the Indian Music Industry—have raised serious legal challenges against foreign AI firms. Their core grievance centers on the unauthorized commercial ingestion of journalistic archives, literature, and recorded music without express licensing, financial compensation, or mandatory opt-out mechanisms.
3. The Judicial Test Case: ‘ANI Media v. OpenAI'(2026) :
The interim proceedings before the Delhi High Court provided the first major test of how Indian copyright law handles Large Language Models. While the court recognized that news organizations hold valid copyrights in their original reported works under Section 13 and Section 17, it refused to grant an interim injunction against OpenAI. The court reasoned that training AI models on stored digital data constitutes machine learning and pattern extraction—a form of computational research that falls prima facie under Section 52(1)(a)(i). The court emphasized that because the underlying training dataset is not exposed directly to the public, and because ChatGPT generated outputs based on statistical probabilities rather than verbatim regurgitation, no market substitution or irreparable financial injury was proven at the interim stage.
4. Expert Policy Panels and Legislative Studies :
Recognizing these legal gaps, the Ministry of Commerce and Industry and policy bodies across India have undertaken comprehensive studies to evaluate updates to the Copyright Act, 1957. The primary goal is to draft explicit Text and Data Mining (TDM) rules that can protect human creators while preventing Indian tech startups from falling behind in international AI development.
ABSTRACT
The rapid deployment of Artificial Intelligence (AI) in India has created significant tension between technological innovation and traditional statutory intellectual property principles. Drafted in an era when creative output required direct human physical and mental effort, the Copyright Act, 1957 assumes a human author at its core. Today, autonomous generative models produce fine art, software code, journalism, and musical compositions within seconds.
This technological paradigm shift presents two primary legal challenges: evaluating whether machine-generated output satisfies statutory standards of originality and authorship under Section 2(d)(vi), and determining whether training AI models on copyrighted data without explicit consent constitutes civil infringement under Section 51.
This article examines these legal questions under Indian statutory jurisprudence. It analyzes recent judicial developments regarding fair dealing under Section 52, explores international legal comparisons, and outlines balanced legislative solutions to ensure that technological growth does not undermine the economic and moral rights of human creators.
CASE LAWS
1. Eastern Book Company v. D.B. Modak (2008) :
The Supreme Court of India officially rejected the English “sweat of the brow” doctrine and established the “modicum of creativity” standard. The Court ruled that copyright protection requires independent human skill, labor, and creative judgment. This decision forms the foundation for arguing that purely prompt-driven AI outputs lack the necessary human creative judgment to qualify for statutory protection.
2. R.G. Anand v. Delux Films (1978) :
The Supreme Court affirmed the foundational principle that copyright protects only the specific ‘expression’ of an idea, not the underlying idea, plot, theme, or general artistic style itself. This distinction is critical for generative AI platforms, as generating new content in the general style or genre of a famous artist does not constitute infringement unless specific, protected expressions are copied.
3. ANI Media (P) Ltd. v. OpenAI Inc. (2026 Delhi High Court) :
A historic interim ruling in Indian AI copyright jurisprudence. ANI sued OpenAI for scraping and storing its copyrighted news articles to train ChatGPT. Justice Amit Bansal held on a prima facie basis that storing publicly available literary works to train LLMs falls under “fair dealing” for research and private use under Section 52(1)(a) of the Copyright Act. The court noted that LLMs analyze statistical patterns rather than memorizing text for direct public distribution, and that without proof of verbatim regurgitation or direct market loss, an interim injunction would cause irreparable harm to the public interest in technological progress.
4. University of London Press Ltd. v. University Tutorial Press Ltd. [1916]:
A landmark common law decision establishing that an “original” work does not require absolute novelty, but rather that the expression must originate independently from the author. This rule informs ongoing legal debates on whether an AI output, derived through statistical processing of thousands of existing training works, can be said to “originate” from the prompter or software user.
5. Feist Publications Inc. v. Rural Telephone Service Co. 499 U.S. 340 (1991) :
The US Supreme Court ruled that raw factual data lacks copyright protection and that originality requires at least a minimal spark of creative selection or arrangement. This principle supports the defense that scraping unprotectable facts and raw public data to train AI systems does not infringe statutory copyright.
6. Getty Images v. Stability AI & Andersen v. Stability AI (Pending Global Proceedings) :
High-profile international class-action lawsuits closely observed by Indian IP scholars. These disputes challenge whether downloading millions of copyrighted photographs and visual artworks to train generative diffusion models constitutes mass reproduction infringement or protected transformative use.
CONCLUSION
The relationship between generative AI and copyright law represents a pivotal moment in intellectual property jurisprudence. Leaving the Copyright Act, 1957 unaddressed created temporary ambiguity, forcing courts to use broad judicial interpretation to resolve complex technological disputes. While recent judicial decisions like ‘ANI Media v. OpenAI’ have provided interim clarity regarding model training, long-term stability requires comprehensive legislative updates from Parliament.
To achieve a balanced legal ecosystem that rewards human talent while encouraging technological growth, Parliament should consider four key statutory amendments:
1. Clarifying Statutory Authorship under Section 2(d)(vi): Amending the Act to explicitly state that computer-generated works qualify for copyright only when there is demonstrable, substantial human creative control over the final expression.
2. Enacting an Explicit Text and Data Mining (TDM) Exception: Updating Section 52 to establish a clear, non-commercial research exemption for data scraping, paired with a statutory “opt-out” mechanism that allows commercial creators to prevent unauthorized scraping of their digital portfolios.
3. Mandatory Dataset Transparency: Requiring commercial AI developers operating within India to maintain accessible registers detailing the copyrighted sources included in their model training pipelines.
4. Collective Management and Statutory Royalty Licensing: Creating statutory licensing frameworks through registered Copyright Societies, ensuring that human authors receive fair financial compensation when their creative works are utilized in commercial AI training datasets.
As AI tools become an everyday part of modern society, the legal system must ensure that technological development coexists with clear statutory protections and equitable compensation for human intellectual effort.
FAQs
Q1. Why doesn’t the Copyright Act, 1957 explicitly protect AI-generated works?
The statute was enacted in an era when creative works were produced exclusively through direct human effort. Section 2(d)(vi) accounts for computer-generated works, but courts interpret this to require a human author who uses the computer merely as an assisting tool, rather than an autonomous software model acting as the sole creator.
Q2. Can I register copyright for an image or essay created using an AI text prompt?
Under current Indian legal standards, entering simple text prompts generally does not provide the direct creative control required to meet the “modicum of creativity” test. Without significant subsequent human alteration, selection, editing, or creative arrangement, purely prompt-generated works may be denied copyright protection.
Q3. Is scraping internet data to train AI models considered illegal under Indian copyright law?
In light of recent interim judicial interpretations likeANI Media v. OpenAI’ (2026), storing publicly available online data to train AI models does not prima facie constitute copyright infringement if it falls under fair dealing for research under Section 52(1)(a). However, if an AI chatbot consistently outputs verbatim, copyrighted text or replaces the market for the original work, it can still face liability for copyright infringement.
Q4. What is the difference between US “Fair Use” and Indian “Fair Dealing” regarding AI training?
The US uses an open-ended, four-factor test under 17 U.S.C. § 107 (“Fair Use”) that allows courts broad flexibility to excuse commercial transformations. India uses an exhaustive, closed list under Section 52 (“Fair Dealing”), which historically required courts to fit modern digital activities into specific categories like private research or criticism.
Q5. How can human artists and writers protect their work from being scraped by AI developers in India?
Creators can implement technical restrictions like robots.txt opt-out headers on their websites, utilize digital watermarking technologies, or license their content through rights organizations. In addition, upcoming statutory reforms seek to establish a formal right for creators to opt out of commercial training datasets.



