Table of Contents
Artificial Intelligence Translates the Language of Animals
Kim Kyung-jin, Attorney at Law
The Story of AI Learning to Listen to Whales, Dolphins, Birds, and Bees
Twelve chapters on how AI listens to dolphins, sperm whales, humpback whales, birds, and bees to find rules in their sounds, and what this technology means for its risks and for animal rights. Written in simple sentences a child can read, with verified sources in every section.
Table of Contents
AI Deciphers Ancient Scripts
Kim Kyung-jin, Attorney at Law
Ancient Records Revived by AI
In twelve chapters, this book explains how AI revives records once unreadable, from burned scrolls and wooden slips buried in mud to broken clay tablets. It covers virtual unrolling at Herculaneum, virtual collation of oracle-bone texts, reading Silla wooden tablets, computational analysis of undeciphered scripts, and multispectral archives, with verified references for each chapter.
Table of Contents
Artificial Intelligence for New Materials Design and Rocket Propulsion Engineering
Kim Kyung-jin, Attorney at Law
AI Potentials, Self-Driving Laboratories, and Physics-Informed Machine Learning (PIML)
Ten chapters on how artificial intelligence is changing new materials and rocket propulsion: atomic simulation, generative models, self-driving labs, high-temperature alloys, metal 3D printing, combustion, cooling design, and engine diagnosis and control. Written without equations, with verified sources in every chapter.
AI Library
The Age of Autonomous Scientific Discovery
Kim Kyung-jin, Attorney at Law
AI Scientists and Self-Driving Labs
This book follows how AI scientists and self-driving labs are changing the way science generates and verifies claims. It covers literature-based discovery, natural-language protocols translated into robot commands, multi-agent research systems, closed-loop laboratories, materials search, the verification gap, chains of evidence, research harnesses, journal ethics, and legal responsibility.
AI Library
A New Era of Life Sciences Opened by Artificial Intelligence
Structural Proteomics, Genomic Foundation Models, Autonomous Laboratories, and Global Governance
Kim Kyung-jin, Attorney at Law
This book is a research volume compiled with artificial intelligence. A human selected the materials and structured the work, while AI models drafted the sentences and cross-checked the facts.
AI Library
The Double Structure of Digital Sovereignty
Europe’s Departure from Palantir and the Chains of American Big Tech
Kim Kyung-jin, Attorney at Law
This is a record of 2026, when European intelligence agencies and defense ministries began removing analytics tools from America’s Palantir. It covers the replacement decisions made by France’s General Directorate for Internal Security (DGSI), Germany’s Federal Office for the Protection of the Constitution (BfV), and the Netherlands Ministry of Defense; the incident in which US export controls severed an ally’s ac…
New English Edition
Artificial Intelligence in Horticulture
Kim Kyung-jin, Attorney at Law
Across five chapters and ten sections, this book examines computer vision for crop diagnosis, harvesting robots and autonomous field systems, smart greenhouses and digital twins, precision irrigation and supply-chain quality control, high-throughput phenotyping, and predictive breeding.
New English Edition
Artificial Intelligence in Food Crop Agriculture
Kim Kyung-jin, Attorney at Law
Across six chapters and eighteen sections, the book examines digital agricultural infrastructure, remote sensing, crop diagnosis, yield forecasting, precision irrigation, genomics, molecular breeding, agricultural robotics, climate-smart agriculture, and global food security.
New English Edition
The Future of Forestry and Agroforestry
Kim Kyung-jin, Attorney at Law
Driven by Artificial Intelligence and Digital Innovation
Across five chapters and fifteen sections, the book follows satellites, drones, LiDAR, digital twins, forest-specific language models, wildfire and pest forecasting, forestry robotics, agroforestry, timber traceability, and forest carbon markets.
New English Edition
Smart Livestock Farming: AI Enters the Barn
Kim Kyung-jin, Attorney at Law
Sensors listen, cameras watch, and artificial intelligence helps farmers decide.
Across five chapters and fifteen sections, the book follows precision livestock farming from animal health and reproduction to robotic milking, virtual fencing, digital twins, methane reduction, welfare, and data ownership.
Table of Contents
Han Dong-hoon, Busan Buk-gu Gap: A Record of the 100 Days Before and After the Election (Mar. 26-Jul. 3, 2026)
Kim Kyung-jin
Table of Contents and 13 sections
From March 26 to July 3, 2026, this record follows the spring after expulsion, the Busan Buk-gu Gap by-election, victory as an independent, and the first bill submitted in the National Assembly.

Table of Contents
Artificial Intelligence and Medicine
Kim Kyung-jin, Attorney at Law
AI in clinical care, hospitals, education, and research
AI in medical imaging, risk prediction, treatment planning, hospital operations, education, and research, with patient safety, privacy, and accountability.
[AI Library] Chapter 1. Generative AI Training Data and Copyright Disputes
Artificial Intelligence on Trial
Part 1. AI and the Collision with Intellectual Property
Chapter 1. Generative AI Training Data and Copyright Disputes
Attorney Kyungjin Kim
A. The All-Out War Between News Organizations and AI Companies
On the evening of December 27, 2023, lawyers from the New York Times legal team filed a complaint they had been preparing for months with the Southern District of New York federal court in Manhattan. It was a 69-page document. It contained unusual evidence. On the left side sat a Pulitzer Prize-winning New York Times investigative report; on the right, text generated by ChatGPT. The two were strikingly similar. The rhythm of the sentences, the choice of words, even the placement of commas.
(1) New York Times v. OpenAI/Microsoft: Unauthorized Training on News Content and Market Substitution
The Times' lawyers called this phenomenon "regurgitation." It means the AI failed to digest its training data and spat it back out verbatim. The evidence in the complaint showed cases where ChatGPT provided near-perfect excerpts from articles that required a paid subscription. The Times' argument was straightforward. You copied our articles. Millions of them.
There is a concept called Fair Use. Put simply, it works like this. Reading a book at the library is fine, but photocopying the entire book and selling it is not. A student quoting a book in a paper is permitted, but starting a publishing house to print identical copies is a crime. Copyright law judges this boundary using four factors. Whether the use is commercial in purpose. The nature of the original work. How much was used. Whether the use substitutes for the original work's market.
The Times emphasized the fourth factor. Market substitution.
People ask ChatGPT instead of visiting the New York Times website. ChatGPT answers based on Times articles. The person who receives that answer no longer needs a Times subscription. It is like an identical shop opening in front of yours and taking all the customers.
OpenAI pushed back. We did not copy; we "learned." A human author reads thousands of books and develops their own style, and AI learning is no different. The regurgitation evidence was abnormal output obtained by deliberately manipulating the model. In April 2025, Judge Sidney Stein of the Southern District of New York rejected most of OpenAI's arguments. He decided to proceed to trial on the core copyright infringement claims. But the real war erupted during discovery. The Times initially demanded 1.4 billion ChatGPT conversation logs. OpenAI objected fiercely, arguing it would violate user privacy.
In November 2025, Judge Ona T. Wang ordered OpenAI to produce 20 million anonymized conversation logs. OpenAI requested reconsideration in December but was denied. In January 2026, the judge finalized the order. The court's reasoning was clear.
This is not illegal wiretapping. Privacy concerns are adequately protected through anonymization and a protective order.
This litigation is currently proceeding as a multidistrict litigation (MDL) consolidating 16 copyright lawsuits. Major news organizations including the New York Times, the Chicago Tribune, and the Daily News are among the plaintiffs.
For reference, multidistrict litigation (MDL) is, in plain terms, a "consolidated management system for lawsuits."
Problems arise when similar lawsuits against the same defendant pour in from courts across the country. If ten courts each hold separate trials, they might reach different conclusions, wasting time and money.
In 1968, the U.S. Congress created a solution. The Judicial Panel on Multidistrict Litigation (JPML) gathers similar cases into one court. Discovery and common issues are handled all at once. The OpenAI copyright litigation is one such case. Sixteen lawsuits, including those by the New York Times and the Chicago Tribune, were consolidated in the Southern District of New York. A single judge will decide the central question: "Is AI training fair use?"
The practical effect of an MDL is bargaining power. When plaintiffs band together, they become stronger. For defendants, settling everything at once beats fighting in courts across the country. That is why most MDL cases end in settlement before trial. Roughly 60% of all federal civil litigation in the United States currently proceeds through MDL. Asbestos, opioid, and data breach lawsuits have gone through this path. Now AI copyright disputes have joined those ranks.
No verdict has been reached yet, but the market is already moving. Some news organizations have begun signing licensing agreements with OpenAI instead of suing. They chose immediate cash over an uncertain courtroom victory.
(2) Thomson Reuters v. ROSS Intelligence: Unauthorized Copying of a Legal Database and Rejection of the Fair Use Defense
The first page of a Delaware federal court ruling sent a warning signal through the AI industry.
On February 11, 2025, Judge Stephanos Bibas ruled that AI startup ROSS Intelligence's unauthorized copying of headnotes (case summaries) from Thomson Reuters' paid legal database Westlaw for AI training did not qualify as fair use.
Headnotes are not the court opinions themselves but something closer to "summary cards" where editors distill the key points of a decision. Thomson Reuters had accumulated these summaries over decades. ROSS Intelligence took them and tried to build a "robot lawyer."
Judge Bibas based his ruling on two grounds.
First, ROSS's purpose in training its AI was to create a commercial product that directly competed with Thomson Reuters.
Second, ROSS's AI model could not be said to have transformed the original works into something with a new purpose or meaning. Taking a competitor's data without authorization to build a competing product with similar functionality cannot receive fair use protection.
This ruling matters for a specific reason. It shattered Silicon Valley's assumption that "AI training is automatically fair use." The court weighed whether the resulting product was a competitive offering directly targeting the plaintiff's market more heavily than the mere fact that "a machine read it." It was a chilling warning for AI companies.
(3) The Perplexity AI Cases: RAG Technology and Real-Time Content Infringement
In late 2024, a new type of artificial intelligence appeared. Perplexity AI.
Built by former Google engineers, this service summarizes search results and delivers answers directly instead of showing links. It billed itself as an "Answer Engine."
For users, it is convenient. No need to visit news sites plastered with ads.
There is a technology called RAG (Retrieval-Augmented Generation).
It resembles a chef who does not memorize recipes but pulls one from the shelf beside him each time a customer orders and summarizes it on the spot. The problem is that the shelf might belong to someone else's paid collection.
In October 2024, Dow Jones, the publisher of the Wall Street Journal and the New York Post, filed suit against Perplexity. In December 2025, the New York Times joined as well.
The plaintiffs made three claims.
Perplexity bypassed paid subscription paywalls.
It ignored websites' anti-crawling protocols (robots.txt).
It reproduced article content nearly verbatim, eliminating any need for users to click through to the original article.
The Chicago Tribune's lawsuit added one more point.
Perplexity, through "hallucination," fabricated content that the news organization never wrote and attributed it to them as if it were their reporting. The suit raised trademark dilution and defamation as well. This case is testing the boundary between training and real-time display.
If earlier AI systems trained on historical data, today's AI reads and summarizes other people's content in real time. Whether to call this "reference" or "real-time content theft" is the question. The answer will determine the entire business model of search-based AI.
The all-out war between news organizations and AI companies is ultimately an attempt to reshape the flow of money. It is a tug-of-war between the share that belongs to those who produce information and the share claimed by technology companies that process and deliver it. And the next front in this fight is unfolding in the courtroom of authors.
As of January 2026, the lawsuits surrounding Perplexity AI are intensifying.
The most advanced front is Dow Jones.
Filed in October 2024 by the parent company of the Wall Street Journal and the New York Post, this case has now entered full-scale evidentiary combat.
On August 21, 2025, Judge Katherine Polk Failla denied both Perplexity's motion to dismiss and its motion to transfer the case to California. A company that maintains a New York office, employs staff there, and erected a billboard in Times Square could not escape the jurisdiction of New York courts.
The fact discovery deadline has been set for June 4, 2026.
Right now, lawyers on both sides are demanding documents, subpoenaing witnesses, and searching for the other side's weaknesses. Dow Jones is seeking disclosure of Perplexity's source code. Perplexity is refusing. That code contains how the RAG system actually processes content.
On December 5, 2025, two new lawsuits were filed almost simultaneously. The New York Times and the Chicago Tribune. The Times' complaint emphasized one additional point. Not just copyright, but trademark.
A Lanham Act violation. The logic goes like this. Perplexity generated false information and placed the New York Times logo beside it. As if the Times had reported it that way. This constitutes damage to brand value. The claim is that a 170-year-old newspaper's credibility is being harmed by AI hallucinations.
There was an interesting procedural decision. The judge handling the Dow Jones case refused to consolidate the New York Times case as a "related case." The Times case was assigned to Judge Vernon Broderick. For Perplexity, this is a nightmare. It must fight similar issues in two different courts, before two different judges, twice.
The Chicago Tribune's lawsuit also focused heavily on the hallucination problem. Content the Tribune never wrote was displayed as though it were Tribune reporting. The suit raised trademark dilution and defamation as well.
In October 2025, Reddit launched an attack from an entirely different angle.
Not copyright, but DMCA Section 1201, the prohibition on circumventing access controls. Reddit sued not only Perplexity but also three data-scraping intermediaries (SerpApi, Oxylabs, AWMProxy). They used the term "data laundering."
When Perplexity could not scrape Reddit directly, it allegedly went through Google search results as a workaround. Reddit set a trap. They posted test content visible only to Google, and within hours it appeared in Perplexity's answers.
The list of plaintiffs keeps growing. Encyclopaedia Britannica, U.S. News & World Report, Japanese and Italian media outlets. At least six lawsuits against Perplexity are currently pending.
Perplexity's position has been consistent. In the words of communications director Jesse Dwyer: "Media companies have been suing new technology companies for a hundred years. Radio, TV, the internet, social media, now AI. Fortunately, they have never succeeded. If they had, we would still be communicating by telegraph."
But outside the courtroom, other moves are underway. Perplexity has signed revenue-sharing agreements with Time, Fortune, and Der Spiegel. It also formed a partnership with Getty Images.
Litigation and negotiation are running in parallel. Lose in court, and you pay more at the negotiating table. Settle at the table, and the court fight ends. Both sides are doing this math.
The discovery deadline in the Dow Jones case is June 2026. Whether Perplexity's source code will be disclosed by then, or whether a settlement comes first. The answer to this question will determine the future of search-based AI.
B. Authors' Class Action Lawsuits and the Definition of Creativity
George R.R. Martin spent decades writing the original novels behind Game of Thrones. His style is distinctive, his world vast, his characters complex. One day, fans sent him a strange tip. They had typed "Write Game of Thrones Book 6 in the style of George R.R. Martin" into ChatGPT, and it produced a fairly convincing novel. Martin was stunned.
(1) Authors Guild and Writer Coalition Lawsuit: Style Imitation and Requirements for Derivative Works
In September 2023, the Authors Guild, along with prominent writers including George R.R. Martin, John Grisham, and Jodi Picoult, sued OpenAI. Their argument was clear.
"No Consent, No Credit, No Compensation."
There is a concept called a derivative work. You don't play the original song as-is, but you create a remix or a film based on it.
The authors argued that AI ingested their books in their entirety, enabling it to generate text mimicking their voice and style, and that such output constitutes a derivative work of the originals.
The court, however, was cautious. Copyright law protects specific "expression," not an author's "manner" or "style" itself. Style resembles handwriting. Handwriting evokes a person, but it is not itself subject to copyright.
On April 3, 2025, multiple authors' class actions against OpenAI were consolidated into multidistrict litigation (MDL). On October 27 of the same year, Judge Stein denied OpenAI's motion to dismiss. The court found that the plaintiffs had sufficiently stated a prima facie claim for direct copyright infringement. The allegations of actual copying and substantial similarity between ChatGPT's outputs and the authors' works were deemed sufficient.
The issues crystallize into two branches.
One is unauthorized copying during the training phase.
The other is substantial similarity in the outputs.
The authors' side argues that the explanation "the model generates sentences probabilistically" may serve as a smokescreen to obscure an economic structure built on large-scale copying. The defense emphasizes that "individual outputs result from user input and stochastic processes," placing the burden of proving causation and substantial similarity on the plaintiffs.
On April 3, 2025, the Judicial Panel on Multidistrict Litigation (JPML) consolidated the multiple authors' class actions against OpenAI as MDL No. 3143. It was assigned to Judge Sidney Stein in the Southern District of New York. Twelve lawsuits were bundled into one: the Authors Guild suit, the New York Times suit, the Sarah Silverman suit, the Michael Chabon suit, and others.
On October 27 of the same year, Judge Stein denied OpenAI's motion to dismiss. He found that the plaintiffs had sufficiently stated a prima facie claim for direct copyright infringement. The allegation that "ChatGPT's summaries parroted the plot, characters, and themes of the original works" was accepted.
The legal arguments are settled. Now it's a battle over evidence.
As of January 2026, both sides have entered an all-out war in discovery.
The court convened consecutive discovery status conferences on January 15 and February 11. The central issue is the identity of the datasets OpenAI used for training. "Books1" and "Books2." These two datasets are the key to everything. The story goes back to 2018.
An OpenAI employee downloaded millions of books from Library Genesis (LibGen), an illegal piracy site. These became Books1 and Books2. In May 2020, OpenAI publicly disclosed in a research paper that it had used these datasets to train GPT-3. Then, just before ChatGPT's launch in 2022, the company deleted them.
Why did they delete them? This question will determine billions of dollars in damages.
In March 2024, OpenAI's outside counsel Joseph Gratz sent a letter to the authors' attorneys. "Books1 and Books2 were removed from training in late 2021 and deleted in mid-2022 for 'non-use.'"
But when the authors' side pressed on what "non-use" meant, OpenAI changed its story. On June 13, 2025, OpenAI attempted to withdraw that portion of the Gratz letter. It argued that the deletion reason originated from conversations with counsel and was therefore protected by attorney-client privilege.
Judge Ona Wang rejected this. On November 24, 2025, she issued an order: "OpenAI stated the 'reason' (implying it was not privileged), and cannot later claim that 'reason' is privileged. OpenAI has treated its privilege claims like a 'moving target.'"
The court ordered the disclosure of messages from internal OpenAI Slack channels named "project-clear" and "excise-libgen," where employees discussed deleting the datasets. All written communications with in-house counsel from 2022 were also subject to disclosure. Every internal reference to LibGen as well. The deadline was December 8, 2025. Depositions of OpenAI in-house counsel were to be completed by December 19.
OpenAI announced it would appeal. But on December 3, 2025, Judge Wang also denied its motion for reconsideration. On December 5, Judge Stein ordered OpenAI to submit an additional brief. Why does this evidence matter? If "willful infringement" is proven, the game changes. Under copyright law, willful infringement allows statutory damages of up to $150,000 per work. Millions of books are involved.
Theoretical liability could reach tens of billions of dollars. If OpenAI knowingly used pirated material and then deleted it to conceal that fact, damages would skyrocket astronomically.
The authors' attorney Justin Nelson has already won on other fronts. He is investigating whether copyrighted data is being used in models OpenAI currently has in development, and whether the deleted datasets are still in use under different names.
The situation as of January 2026 comes down to this: how transparently OpenAI discloses its training data, and how much evidence of pirated material the authors can uncover within it, will decide the outcome.
Pressure on OpenAI is intensifying outside the courtroom as well, because of a precedent set by competitor Anthropic. In September 2025, Anthropic agreed to pay $1.5 billion. It is the largest copyright settlement in U.S. history. Anthropic's $1.5 billion settlement has now become the benchmark.
If OpenAI loses in court, it will have to pay far more than that. Even if it wins, it has already absorbed years of legal costs and reputational damage. And the next front is already open. Internal documents revealed that Meta used the same LibGen dataset. Evidence shows Mark Zuckerberg approved its use despite knowing it carried "medium-to-high legal risk."
The authors' courtroom battle has only just begun.
(2) Anthropic Class Action and Settlement Developments
On September 5, 2025, a single number was read aloud in a San Francisco courtroom. $1.5 billion. The gallery went quiet. Anthropic had reached the largest settlement in U.S. copyright law history with the authors.
This case began in 2024. Authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic.
The allegation was that Anthropic used data from pirate book sites like Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi) in the process of training its AI model, Claude.
In June 2025, Judge William Alsup of the U.S. District Court for the Northern District of California issued a decisive ruling. Using lawfully purchased books for training is "one of the most transformative things we will see in our lifetimes" and constitutes fair use. But using pirated works is "inherently, irreparably infringing" and cannot qualify as fair use.
Judge Alsup denied summary judgment on the pirated copies and ordered a trial. Under U.S. copyright law, willful infringement can trigger statutory damages of up to $150,000 per work. Anthropic had downloaded roughly 500,000 books from pirate sources. Do the math: potential liability could exceed $70 billion. Enough to destroy the entire company.
Anthropic came to the negotiating table. The settlement terms were as follows. Pay a minimum of $1.5 billion. Distribute approximately $3,000 per book. Destroy all copies of works obtained from pirate sites. But the settlement grants immunity only for past conduct. It does not cover infringement claims related to future training or AI outputs.
On September 25, 2025, Judge Alsup granted preliminary approval of the settlement. A final approval hearing is scheduled for April 2026. Authors Guild CEO Mary Rasenberger said: "This historic settlement is an important step in recognizing that AI companies cannot simply take the creative works of authors because they need books to develop quality large language models."
This settlement left three lessons. First, the path of data acquisition moved to the center of litigation risk. Second, "deletion and cleanup" became not just an ethics issue but a key variable in determining remedies and damages. Third, the licensing market began functioning as a "defensive perimeter" rather than an optional choice.
(3) Silverman, Kadrey, Chabon v. Meta MDL Consolidated Litigation
On July 7, 2023, comedian Sarah Silverman discovered that her memoir, The Bedwetter, had been ingested by Meta's AI. She filed suit against Meta alongside authors Richard Kadrey and Christopher Golden.
The lawsuit soon expanded to include 13 authors, among them Michael Chabon, Junot Diaz, and Andrew Sean Greer. Two Pulitzer Prize-winning works were among them.
What they targeted was a dataset called Books3. Roughly 190,000 books. Most had been illegally copied from a shadow library called Bibliotik. Meta trained LLaMA on this dataset.
Internal Meta documents disclosed in early 2025 were even more shocking. They revealed that Mark Zuckerberg personally approved the use of the LibGen dataset and was fully aware it consisted of pirated material.
On June 25, 2025, Judge Vince Chhabria of the U.S. District Court for the Northern District of California ruled in Meta's favor. He held that copying copyrighted books without authorization for AI training constitutes fair use. He called it "highly transformative."
Two days earlier, a different court had reached the opposite conclusion. That was Judge William Alsup's ruling in the Anthropic case. Judge Alsup recognized that scanning lawfully purchased books to train AI is fair use, calling it "one of the most transformative uses we will see in our lifetimes." But he drew the line at books downloaded from pirate sites, holding that their use is not fair use. Anthropic settled for $1.5 billion.
Two rulings issued in the same week. They appeared to address the same issue, yet reached different results. Why?
The two judges were looking at different questions. Judge Alsup asked: "How was the data acquired?" He concluded that downloading from illegal piracy sites cannot receive the protection of fair use.
Judge Chhabria asked: "How was the data used?" He found that AI training serves an entirely different, transformative purpose from the original works, and therefore qualifies as fair use.
There was another critical difference.
The plaintiffs' failure of proof. Judge Chhabria wrote in his opinion:
"Meta presented evidence that the copying did not cause market harm. The plaintiffs failed to present any empirical evidence to the contrary." There was no evidence that LLaMA generates text substantially similar to the original works. No harm, no victory.
But Judge Chhabria attached an important caveat. "This ruling applies only to the specific circumstances of this case." He added: "Had the plaintiffs presented evidence that LLaMA is permitted to generate works that directly compete with their own, the result might have been different."
That sentence opened a new path for the plaintiffs.
On October 27, 2025, Meta sent notice to the plaintiffs. It concerned "newly discovered evidence" regarding its past downloading of files via torrent protocol from shadow libraries. On November 5, both sides requested a schedule extension. The summary judgment hearing was pushed from April 2, 2026 to April 30.
A new strategy emerged.
Instead of fighting over fair use at the training stage, attack the torrenting conduct itself. The attempt is to apply in Judge Chhabria's court the legal principle that Judge Alsup established in the Anthropic case. Meta's downloading of millions of books from LibGen via BitTorrent protocol is essentially the same as Anthropic's acquisition of data from the same sites. The illegality of the acquisition is not cured by subsequent transformative use.
As of January 2026, the litigation continues. Judge Chhabria's fair use ruling stands, but it is only half the story. The other half, the torrenting issue, awaits the April 30 hearing. If the plaintiffs prevail on this point, the training-stage fair use ruling is effectively neutralized. No matter how transformative the training, data acquired illegally cannot be protected.
Anthropic resolved the matter for $1.5 billion. Meta chose to fight it out in court to the end. Whether that choice was wise will become clear after April 30.
C. Code-Generation AI and Open-Source Licenses
Matthew Butterick is a lawyer and a programmer. A rare combination. In 2022, while using GitHub's AI tool Copilot, he felt an odd sense of recognition. The code snippets Copilot suggested looked too much like code he had written in the past, or code he had seen in the open-source community.
(1) The GitHub Copilot Lawsuit: The Open-Source License Violation Debate
Open-source software is an enormous tower built on the spirit of sharing. Developers make their code visible to anyone. Others take that code and use it. There is one important rule governing this arrangement.
Licenses. Think of a free recipe handed out on the street, but with conditions printed on the paper. "You may take this, but credit the source." "Redistribute under the same terms." Those are the conditions. This is a sacred compact among developers.
On November 3, 2022, Butterick and anonymous developers filed a class action against Microsoft, GitHub, and OpenAI in the U.S. District Court for the Northern District of California.
Copilot trained on billions of lines of open-source code. When users write code, it auto-completes based on what it learned. The problem is that when Copilot produces code, it strips away the original author's name and license notices entirely.
Butterick called this "the largest-scale copyright laundering in software history." The plaintiffs' core claim was a violation of DMCA (Digital Millennium Copyright Act) Section 1202. This provision prohibits the unauthorized removal or alteration of "copyright management information" (CMI). In plain terms, it forbids "tearing the author's name off a book cover and distributing copies." In July 2024, Judge Jon S. Tigar dealt the plaintiffs a major blow.
He dismissed the DMCA Section 1202(b) claim. The judge's reasoning was this: the code Copilot generates is not "identical" to the original. Therefore, the DMCA does not apply. This is the "identicality requirement."
The plaintiffs did not give up. On September 27, 2024, Judge Tigar granted the plaintiffs' request to certify this issue for interlocutory appeal to the Ninth Circuit Court of Appeals.
The core question is this: Does DMCA Section 1202(b) apply only when the AI output is "identical" to the original, or does it also apply when the output is merely "similar"? 17 U.S.C. Section 1202(b)
(b) REMOVAL OR ALTERATION OF COPYRIGHT MANAGEMENT INFORMATION.,No person shall, without the authority of the copyright owner or the law,
(1) intentionally remove or alter any copyright management information,
(2) distribute or import for distribution copyright management information knowing that the copyright management information has been removed or altered without authority of the copyright owner or the law, or (3) distribute, import for distribution, or publicly perform works, copies of works, or phonorecords, knowing that copyright management information has been removed or altered without authority of the copyright owner or the law, knowing, or, with respect to civil remedies under section 1203, having reasonable grounds to know, that it will induce, enable, facilitate, or conceal an infringement of any right under this title.
Section 1202(c) defines "copyright management information":
(c) DEFINITION.,As used in this section, the term "copyright management information" means any of the following information conveyed in connection with copies or phonorecords of a work or performances or displays of a work, including in digital form:
(1) The title and other information identifying the work, including the information set forth on a notice of copyright.
(2) The name of, and other identifying information about, the author of a work.
(3) The name of, and other identifying information about, the copyright owner of the work, including the information set forth in a notice of copyright. The answer to this question could rewrite the rules for the entire AI industry.
If the appellate court confirms the 'identity requirement,' AI companies can escape DMCA liability by making even minor alterations to code. If the court rules that mere 'similarity' is sufficient, coding AI tools will bear the enormous burden of tracking and complying with the licenses of every piece of open-source code in their training data.
Judge Tigar did not dismiss the open-source license violation and breach of contract claims. He treated open-source licenses as actual binding contracts. These claims are ongoing, and the plaintiffs are building evidence that Copilot 'memorizes' their code and reproduces it in its output.
Current status: awaiting oral argument scheduling or a decision at the Ninth Circuit Court of Appeals. The district court proceedings are stayed pending the appellate ruling. This decision is expected to set precedent across the full range of AI copyright disputes.
(2) The Legality of Training Data for Coding AI
Disputes involving coding AI sit at the exact point where 'fair use' logic and 'breach of contract' logic collide head-on. The legality of training data is a question of where you bought your ingredients.
Microsoft and OpenAI's position is firm. Training on publicly available code from GitHub constitutes fair use. Code generated by AI is a transformation of the original, not a copy. Very short code snippets cannot be protected by copyright. A simple loop like 'for (int i=0; i<10; i++)' cannot belong to anyone.
Developers argue the opposite. Open-source code means 'anyone can see it,' not 'anyone can freely use it for commercial purposes.' The GPL license imposes copyleft obligations requiring disclosure of derivative works. Even the MIT license requires author attribution. With Copilot offered as a paid subscription model, platform companies monopolize the profits while exploiting code built on other people's labor.
Three technical issues recur across these cases.
First, whether copying during the training phase constitutes temporary reproduction or permanent reproduction.
Second, whether the output reproduces a 'substantial portion' of code from a specific repository.
Third, whether the system was capable of tracking source attribution and licensing but excluded this function by design.
On this axis, the shield that companies raise is 'stochastic generation.' The sword that plaintiffs wield is 'statistics on duplicated output and pattern reproduction.' GitHub's own FAQ acknowledges that 'in about 1% of cases, a suggestion may contain a code snippet of 150 characters or more that matches the training set.' Independent analysis reports that 'in files where Copilot is active, Copilot accounts for nearly 40% of code in popular programming languages like Python.'
Plaintiffs estimate that statutory damages for DMCA violations alone could exceed $9 billion. This lawsuit is tied to the future of programming as a profession. Ironically, programmers trained the AI that may replace them by sharing their own code.
Within the open-source community, discussions are underway to develop new licenses that include explicit provisions regarding AI training. Some projects are adding 'AI training exclusion' clauses to their licenses. Code has clearer structure than text or images, and copyright licensing rules are comparatively well established. The outcome of this lawsuit is therefore likely to arrive before rulings in the text or image domains, making it a critical benchmark for the AI copyright wars ahead.
The court must now decide whether to allow the open-source ecosystem, built on the spirit of sharing, to paradoxically become fuel for an AI that erodes that very ecosystem. And this decision will shape the principles applied equally to news articles, authors' books, and developers' code.
Kim Kyung-jin
Attorney · Former Member of the National Assembly · AI Policy Researcher
© 2026 Kim Kyung-jin. All rights reserved.
















