LIBRARY INSIDE MACHINE: COPYRIGHT PROTECTION AND INNOVATION

LIBRARY INSIDE MACHINE: COPYRIGHT PROTECTION AND INNOVATION

Name of author:  Khushi Jain (II Year B.A. L.L.B (Hons.) student at Dr. Ram Manohar Lohiya National Law University, Lucknow)

Artificial Intelligence (“AI”) has transcended the walls of laboratory curiosity and permeated classrooms, lectures, and students’ laptops. While it offers several advantages, it also raises a pertinent question. Do AI-generated educational materials qualify as ‘original works’ deserving copyright protection, or are they unlicensed reproductions that infringe the rights of authors and publishers? There are more than 40 pending AI training cases in many jurisdictions, including U.S., Japan, Europe, China, UK, etc.

Recently, OpenAI is facing a landmark copyright battle in India as news publishers, including ANI and members of the Digital News Publishers Association, accuse it of training ChatGPT on their content without authorization, challenging the boundaries of fair use. Although UNESCO is committed to harnessing the potential of AI technologies for achieving the Education 2030 Agenda, it ensures that its application in educational contexts is guided by inclusion and equity.

Centered on the precise conundrum, the blog explores how copyright law must evolve to balance three competing imperatives: fostering technological innovation, ensuring affordable access to education, and protecting the legitimate rights of authors. It examines whether the act of converting literary works into machine-readable patterns amounts to legitimate learning or unauthorized reproduction. By situating this debate within the ongoing OpenAI case, the blog sheds light on how courts and policymakers might reconcile the twin goals of protecting authors’ rights and promoting technological innovation.

BOOKS AS AI FUEL

The US Copyright Office and investigative reporting revealed that one major AI model was trained using over 2 billion documents, including 4 million books. Pirated repositories like LibGen contain over 7.5 million books and 81 million research papers, many of which have been used for model training without proper authorization. Model developers (like OpenAI CEO Sam Altman) have also confirmed that datasets contain large amounts of copyrighted work.

The legal debate centers on whether such use falls within the scope of fair use or permissible exceptions. Advocates argue that AI training is a transformative use because it does not reproduce the text verbatim but instead extracts patterns of language and meaning. Critics, on the other hand, maintain that incorporating copyrighted books into training datasets without authorization undermines authors’ rights.

A distinction must be drawn between reproduction and inspiration as drawn in the Warhol case. If an AI reproduces a passage from a book verbatim, it may constitute direct infringement. However, if the AI merely generates text that is stylistically similar or conceptually inspired by a book, the legal question is more complex.

THE PUZZLE OF ORIGINALITY

The Copyright Act, 1957 under Section 13 protects original literary works, but the term original has been judicially interpreted rather than statutorily defined. Section 2(d) defines the author as the person who creates the work, which presupposes a natural person exercising intellectual skill and judgment.

It has been recognised that compilations can be protected if they involve sufficient skill, labour, and judgment in selection or arrangement. The Supreme Court refined this by adopting the modicum of creativity test. It rejected the mere sweat of the brow doctrine. It included that the work must reflect some creative spark attributable to human agency. Hence, originality does not mean novelty but requires a minimal degree of creativity and application of the mind.

The Copyright Act, 1957 was enacted in a pre-digital era and does not expressly address authorship or ownership of AI-generated works. The current framework presupposes human authorship and does not contemplate non-human agents.  AI reliance on copyrighted inputs without authorization raises clear infringement under Section 51, while the moral rights claim under Section 57 highlights the risk of distortion and misattribution by AI. At the same time, emphasis lies on the need for access to affordable resources under Article 21A of the Constitution.

Thereby, copyright law must strike a balance between incentivizing creativity and promoting access to knowledge. Similarly, in Eastern Book Company v. D.B. Modak, the Supreme Court emphasized that copyright is not a reward for labour but a protection for creative expression, reinforcing the importance of maintaining originality while preventing monopolies over knowledge.

CROSS JURISDICTIONAL ANALYSIS

A comparison is essential to situate India’s copyright challenges within global practices, offering tested models that can guide balanced reforms between protection and innovation. Several jurisdictions have enacted exceptions allowing for text and data mining (“TDM”) potentially applicable to AI training.

The European Union, through the 2019 Directive on Copyright in the Digital Single Market, directs member states to provide exceptions for “reproductions and extractions” of copyrighted material for use in TDM, in certain circumstances. It applies to “research organisations and cultural heritage institutions in order to carry out, for the purposes of scientific research, text and data mining of works or other subject matter to which they have lawful access.” They have also adopted the EU AI Act.

Singapore’s framework allows the use of copyrighted works for computational data analysis, but only where the user has lawful access to the material. Section 244(1) and (2) of the Copyright Act 2021 states that the copies created during this process must be confined strictly to data analysis and may be shared with others solely for verifying results or conducting collaborative research. Similarly, India could require that AI training be limited to work to which the developer already has lawful access, ensuring that piracy or unauthorized scraping is clearly excluded.

Japan permits the use of copyrighted works for AI development or other data analysis, provided the use is not aimed at “personally enjoying…the thoughts or sentiments expressed in that work.” The exception is excluded where such use would “unreasonably prejudice the interests of the copyright owner.” In its 2024 AI guidelines, Japan’s Copyright Office clarified that “enjoyment” refers to gaining intellectual or emotional satisfaction from the work, such as reading literature, appreciating music, or running computer programs. Creating material that mimics original works may also count as “enjoyment,” and if enjoyment is even a partial purpose, the exception does not apply. Likewise, reproducing a copyrighted database for AI training is not covered when licenses for data analysis are already available in the market.

Drawing lessons, in India, Section 52 of the Copyright Act could be amended to provide a safe harbor for TDM when performed by entities with lawful access, for bona fide research, innovation, or development of AI systems. It should be explicitly clarified that private individuals, educational institutions, and licensed developers can mine data for model training, as long as outputs do not reproduce original works verbatim.

SUGGESTIONS FOR BALANCED COPYRIGHT REFORM

Policy evolution could take several directions. First, India may consider adopting a provision akin to Section 9(3) of the UK Copyright, Designs and Patents Act, 1988, which recognizes computer-generated works and attributes authorship to the person who made the necessary arrangements. It would provide clarity on ownership of AI-assisted works without equating machines to authors.

Second, Section 52 on fair dealing could be expanded to explicitly include educational uses involving AI tools, ensuring that private study, research, and classroom dissemination are clearly protected.  Delhi High Court has upheld educational photocopying as fair dealing, recognising the constitutional importance of access to education. Extending this principle to AI-based educational platforms would align with Article 21A while maintaining limits against commercial exploitation.

Third, a structured licensing mechanism, such as compulsory or statutory licensing under Section 31, could be developed for educational AI platforms. This would allow startups to access works legally while ensuring publishers receive fair remuneration.

Finally, moral rights under Section 57 must be strengthened to prevent misattribution and distortion of academic works in AI outputs. In Amarnath Sehgal v. Union of India, the sanctity of moral rights was underscored, holding that the destruction or distortion of an artist’s work violated his right of integrity.  India must also remain consistent with the Berne Convention and TRIPS obligations, which require protection against unauthorized reproduction and global dissemination.

Therefore, Indian copyright law must recognise AI-assisted works without undermining human creativity, expand educational fair dealing, and create licensing frameworks to balance access and compensation. This balanced approach will ensure that India fosters technological innovation while upholding both the rights of authors and the public’s right to knowledge.

CONCLUSION

The future of AI and copyright will depend on reconciling competing stakeholder interests. Authors and publishers fear loss of revenue and misattribution, pushing for licensing frameworks and collective bargaining. Students and educators highlight the constitutional mandate of affordable access and judicial recognition of educational fair dealing. Meanwhile, Tech innovators seek legal clarity and safe harbours, cautioning that excessive restrictions could stifle experimentation and growth.

A balanced approach that harmonizes these concerns will be crucial to shaping a copyright regime that both protects creativity and nurtures innovation. Book publishers and authors could license works collectively for AI training like music royalties, ensuring fair compensation without individual negotiations. It should be complemented by the author’s autonomy, requiring them to register works that they do not want included in AI training datasets. These reforms ensure that copyright law evolves not as a barrier to progress, but as a bridge between protecting human creativity and enabling the transformative potential of AI.

Leave a Reply

Your email address will not be published. Required fields are marked *