The Challenges of AI and Copyright.

Author
Elzaburu
Date
January 9, 2023

The recent class-action lawsuit filed againstGitHub,Microsoft,OpenAI, and OpenAI Codex, seeking $9 billion, is evidence of a problem that was bound to arise for#artificialintelligence: its development may infringe on#copyright, and great care must be taken.

The lawsuit in question challenges the legality of using GitHub repositories to train GitHub Copilot, a service that auto-completes programming code using artificial intelligence. The lawsuit, filed by Matthew Butterick, alleges that 11 open-source licenses and copyrights have been infringed.

How do artificial intelligence systems work?

The fact is that training AI systems requires feeding them enormous databases (such as those found on GitHub) to develop the large language models (LLMs) that power this technology.

In the case at hand, we are dealing with a large database containing#opensource code. This code, used to train AI, may be copyleft—with viral licenses—or under permissive licenses—which are less open. In any case, they require respect for copyright.

This may require anyone who uses open-source code to disclose its use, attribute it to the author, and comply with the terms of the license, which, among other things, may require keeping the code open source for extended or modified versions of it or for any code into which it is integrated as a component.

Content That Infringes Copyright

Well, this is not the case with the GitHub Copilot service, which would not only violate those rights and terms but would also encourage copyright infringement among its users, since they are unaware that the code snippets used to autocomplete their own code belong to someone else. Thus, they may even be creating commercial code without having true freedom to use the code provided by GitHub Copilot for this purpose.

AI systems from other companies, such as Google and Facebook, are being developed in the same way. And they are not only using programming code to power this technology, but also other types of copyrighted texts, such as literary works, journalistic texts, music, etc.

For this reason, many experts are questioning whether the use of such works to fuel the development of this technology is valid and what measures need to be taken to ensure that it is. Of course, human inspiration draws from sources and not from nothing, and it is legitimate for AI to do the same; but what measures need to be taken to ensure that AI does not generate content that infringes copyright after reading those sources?

At the very least, this will force companies that use GitHub Copilot and other similar tools to conduct thorough code audits. Otherwise, they risk having all their work rendered commercially unusable, among other things.

Alberto López Cazalilla, attorney at ELZABURU