Intellectual Property Fragmentation In Machine-Generated Code
Intellectual Property Fragmentation in Machine-Generated Code
1. Introduction
Intellectual Property Fragmentation in Machine-Generated Code refers to the situation in which software code produced or substantially assisted by AI systems may implicate multiple overlapping or uncertain intellectual-property interests.
The problem arises because machine-generated code can simultaneously involve:
human-written source code;
AI-generated code;
open-source components;
proprietary training material;
third-party libraries;
APIs and SDKs;
automatically generated documentation;
generated tests;
copied or memorized code;
model-generated adaptations of existing code.
Consequently, determining who owns, controls, licenses, or may commercially exploit the resulting code can become considerably more complicated than in conventional software development.
The central legal problem is:
When machine-generated code is created from multiple technological and legal inputs, how should copyright, licensing, trade-secret, patent and contractual interests be allocated among the developer, AI provider, software licensors and third-party rights holders?
2. What Is Machine-Generated Code?
Machine-generated code includes code produced through:
large language models;
code-generation systems;
autonomous coding agents;
AI pair programmers;
automated program synthesis;
reinforcement-learning programming systems;
code completion tools.
There are several levels of machine involvement.
Level 1 — Human-written code with AI assistance
The programmer writes the architecture and receives AI suggestions.
Level 2 — AI-generated functions
The developer gives a prompt and accepts generated functions with limited modification.
Level 3 — AI-generated modules
An AI system creates substantial portions of a software application.
Level 4 — Autonomous software development
An AI agent determines architecture, writes code, tests it and modifies it with minimal human intervention.
The further development moves toward Level 4, the more difficult conventional IP concepts become.
3. Why "Fragmentation" Occurs
Fragmentation occurs because the final program may consist of numerous legal layers.
For example:
Final application
→ human-authored architecture
→ AI-generated functions
→ open-source library
→ proprietary API
→ third-party algorithm
→ automatically generated code
→ developer modifications.
Each component may have a different legal status.
Therefore, the question:
"Who owns the software?"
may have no single answer.
4. Copyright Is the Central Issue
Copyright traditionally protects original expression.
Software code can qualify as copyrightable expression.
However, copyright systems generally require some form of human authorship or intellectual creation.
This creates a fundamental difficulty:
Can code generated autonomously by a machine receive copyright protection?
If the answer is uncertain or negative, the resulting code may be commercially valuable but lack conventional copyright exclusivity.
5. Human Contribution
A strong copyright claim becomes more plausible where humans:
formulate detailed creative instructions;
select among AI outputs;
substantially modify generated code;
determine architecture;
arrange modules creatively;
integrate multiple outputs;
debug and transform generated material.
Thus:
AI output + meaningful human creative contribution
may potentially produce protectable human-authored expression.
But:
AI output with no meaningful human creative contribution
raises substantially greater questions concerning copyright ownership.
6. The "Ownership Gap"
Suppose an AI system independently generates:
a novel 500-line software module
Who owns it?
Possibilities include:
the user;
the AI provider;
the developer's employer;
nobody;
the person exercising sufficient creative control;
an entity determined by applicable legislation or contract.
Different jurisdictions may produce different results.
This creates jurisdictional fragmentation.
7. Contractual Ownership Does Not Necessarily Equal Copyright
AI providers may contractually state that users receive rights in generated outputs.
But contractual allocation and copyright ownership are different concepts.
A contract may determine:
who may commercially use output;
warranties;
indemnities;
restrictions;
confidentiality;
liability.
It cannot necessarily create copyright protection where the applicable copyright statute does not recognize the required authorship.
Thus:
Contractual control ≠ statutory copyright ownership.
8. Training-Data Fragmentation
Another layer concerns the material used to train the AI model.
The model may have been trained on:
open-source code;
copyrighted repositories;
public websites;
proprietary datasets;
developer documentation;
licensed databases.
The resulting generated code may therefore raise questions about whether training material has influenced the output.
This produces two separate legal questions:
Question 1
Was the training process itself lawful?
Question 2
Does the generated output infringe third-party rights?
These should not be automatically conflated.
9. Memorization and Substantial Similarity
A particularly difficult problem arises where an AI system reproduces recognizable portions of existing code.
For example, an AI coding assistant might generate:
a distinctive function;
an unusual algorithm implementation;
comments;
variable structures;
a substantial code sequence.
If the underlying code was copyrighted, the generated output may potentially reproduce protected expression.
This creates a risk of latent copyright infringement.
10. Open-Source Fragmentation
Machine-generated code may unintentionally reproduce or resemble open-source code subject to:
MIT licenses;
Apache licenses;
GPL;
LGPL;
AGPL;
BSD licenses;
other copyleft or permissive terms.
This creates a compliance problem.
A developer may believe:
"The AI generated this code, so it is unrestricted."
That assumption may be incorrect.
If the output incorporates protected open-source expression, the relevant licensing conditions may still matter.
11. Copyleft Risk
Copyleft licenses create particularly important concerns.
Suppose an AI-generated module substantially incorporates GPL-licensed code.
If that module is integrated into a larger proprietary software system, questions may arise regarding:
source-code disclosure;
derivative works;
distribution conditions;
license compatibility.
Therefore, AI-generated code may create license contamination or compliance risk.
12. Patent Fragmentation
Copyright is not the only relevant IP right.
Machine-generated code may implement:
patented methods;
patented technical processes;
patented software-related inventions.
Even if the code itself is not copyright-infringing, its implementation may potentially implicate patent rights.
This creates another distinction:
Copyright protects expression; patents can protect qualifying technical inventions.
AI-generated software may therefore require separate patent clearance.
13. Trade Secrets
AI coding systems also create trade-secret concerns.
A developer might inadvertently provide an AI system with:
proprietary source code;
confidential algorithms;
customer information;
internal architecture;
credentials;
unpublished product designs.
The problem becomes particularly serious if confidential information is processed by an external AI service.
The resulting dispute may concern:
unauthorized disclosure;
loss of secrecy;
contractual confidentiality;
ownership;
use of confidential information.
14. Employee and Employer Ownership
Traditional software law frequently allocates rights through employment and work-for-hire principles.
AI complicates this structure.
Suppose:
Employee + AI system → software.
Potential questions include:
Is the employee the author?
Is the employer the copyright owner?
Did the employment contract cover AI-generated output?
Who owns the prompts?
Who owns modifications?
Who bears infringement liability?
Employment contracts therefore increasingly need explicit AI provisions.
15. Case Law
The following cases provide important principles for analysing IP fragmentation in machine-generated code.
16. Feist Publications, Inc. v Rural Telephone Service Co.
499 U.S. 340 (1991)
The U.S. Supreme Court emphasized the requirement of originality for copyright protection.
Relevance
Machine-generated code must be analysed for originality rather than assuming that every automatically generated output qualifies for copyright.
The case is especially important for distinguishing:
mere information;
ideas;
unoriginal material;
from protectable expression.
17. Computer Associates International, Inc. v Altai, Inc.
982 F.2d 693 (2d Cir. 1992)
The case developed the famous abstraction-filtration-comparison approach for software copyright disputes.
Relevance to AI-generated code
This framework is highly useful where an AI-generated program resembles existing software.
The analysis can separate:
ideas;
functional elements;
unprotectable components;
protectable expression.
Thus, similarity between AI-generated code and an existing program does not automatically establish infringement.
18. Google LLC v Oracle America, Inc.
593 U.S. 1 (2021)
The U.S. Supreme Court considered Google's copying of Java API declarations and concluded that Google's use constituted fair use under the circumstances.
Relevance
The case is extremely important for machine-generated programming because APIs are central to AI-generated code.
AI systems routinely generate code using:
APIs;
libraries;
standard programming structures.
The case demonstrates the importance of distinguishing functional software interfaces from expressive implementation code.
19. SAS Institute Inc. v World Programming Ltd.
C-406/10, Court of Justice of the European Union (2012)
The CJEU held that certain functional aspects of software, including ideas and principles underlying programs, are not protected as such by software copyright.
Relevance
AI-generated code frequently reproduces:
algorithms;
functionality;
programming logic;
interfaces.
The case supports careful separation between:
functionality and protected expression.
This is essential when assessing whether AI output infringes existing software.
20. BSA v Ministry of Culture of the Czech Republic
C-393/09, CJEU (2010)
The Court examined the scope of copyright protection for computer programs.
Relevance
The case reinforces the distinction between a computer program's expressive elements and underlying ideas or functions.
This is particularly important when AI generates alternative implementations of the same functionality.
21. Oracle America, Inc. v Google LLC — Earlier Proceedings
The broader Oracle litigation also produced important lower-court decisions concerning:
APIs;
software structure;
copyrightability;
fair use;
functionality.
Relevance
Machine-generated code often works through standardized interfaces.
Therefore, the Oracle litigation provides a useful analytical framework for determining whether generated code merely implements functional interfaces or reproduces protected expressive material.
22. Authors Guild v Google, Inc.
804 F.3d 202 (2d Cir. 2015)
The case concerned Google's mass digitization and transformation of copyrighted works.
Relevance
Although not a software case, it provides important principles concerning large-scale computational processing of copyrighted material.
It is relevant by analogy to AI systems because machine-learning systems often process massive quantities of copyrighted material.
The case highlights the importance of:
purpose;
transformation;
market effects;
copying;
functionality.
23. Thaler v Perlmutter
D.C. Circuit / U.S. copyright proceedings concerning AI-generated artwork
The litigation addressed whether a work generated by an AI system without human authorship could receive copyright protection.
Relevance to machine-generated code
Although involving visual art rather than software, the fundamental question is directly relevant:
Can a machine be the author for copyright purposes?
The answer in U.S. copyright law has strongly emphasized the necessity of human authorship.
This makes the human contribution to AI-generated software particularly important.
24. Naruto v Slater
888 F.3d 418 (9th Cir. 2018)
The Ninth Circuit rejected the idea that an animal could bring a copyright claim concerning photographs it had taken.
Relevance
Although not an AI case, it reinforces the broader principle that statutory copyright systems are structured around legally recognized authors.
The analogy is useful when considering whether an AI system itself can be the copyright author.
25. The Fragmentation Problem in Practice
Consider the following software:
AI-generated application
40% AI-generated code;
30% developer-written code;
15% open-source libraries;
10% third-party APIs;
5% automatically generated configuration.
The legal status is fragmented.
| Component | Potential legal issue |
|---|---|
| Human-written code | Copyright |
| AI-generated code | Authorship uncertainty |
| Open-source code | License compliance |
| APIs | Functionality/copyright |
| Third-party libraries | Copyright/licensing |
| AI model | Contractual terms |
| Training data | Copyright/data rights |
| Proprietary algorithms | Trade secrets/patents |
Therefore, the "software" is legally a composite object.
26. The Problem of Provenance
One of the biggest challenges is proving where code came from.
A company may need to establish:
whether code was AI-generated;
which model generated it;
which prompt was used;
whether the output was modified;
whether similar code existed elsewhere;
whether an open-source component was incorporated.
Without provenance records, defending a copyright or infringement claim becomes difficult.
27. AI Code Provenance as an IP Control
Organizations should maintain:
Prompt → model → output → developer modification → repository commit → final release
This creates an audit trail.
It can help establish:
human contribution;
authorship;
licensing;
provenance;
compliance;
indemnification.
28. License Compatibility
AI-generated code may also combine incompatible licenses.
For example:
Component A — permissive license
Component B — GPL
Component C — proprietary license
Component D — AI-generated output.
The resulting program may have licensing restrictions that are difficult to identify automatically.
This creates license fragmentation in addition to copyright fragmentation.
29. Patent Clearance
AI-generated code should potentially be subjected to patent analysis where the application implements technologically significant processes.
A developer cannot safely assume:
"The code is original, therefore it is legally unrestricted."
Copyright originality and patent freedom-to-operate are separate issues.
30. Trade-Secret Fragmentation
AI development can also fragment secrecy.
For example:
Company source code → AI service → generated output → developer repository.
If confidential information crosses organizational boundaries, the company must assess whether appropriate confidentiality safeguards exist.
This is particularly important for:
proprietary algorithms;
financial models;
security systems;
customer information;
unreleased products.
31. Competition-Law Dimension
Machine-generated code also has an important competition-law dimension.
Suppose one AI coding platform has access to an enormous proprietary code dataset while competitors do not.
The platform could obtain:
superior coding models;
better debugging;
greater developer adoption;
more training data;
stronger network effects.
This can create a data-driven AI coding monopoly.
Thus:
IP fragmentation + AI concentration + proprietary training data
may become a competition concern.
32. Interoperability Problems
Different AI coding systems may produce incompatible:
formats;
APIs;
development environments;
model-specific tools;
proprietary agents.
If one platform becomes dominant, developers may become locked into its ecosystem.
This can transform IP control into technological dependency.
33. Ownership Fragmentation in Corporate Development
A corporate software project may involve:
employee developers;
contractors;
open-source contributors;
AI providers;
third-party libraries;
autonomous coding agents.
Ownership may therefore be dispersed across several contractual and legal relationships.
This creates significant due-diligence challenges during:
acquisitions;
licensing;
investment;
software sales;
IPOs.
34. M&A Due Diligence
When acquiring an AI-developed software company, the purchaser should determine:
What percentage of code was AI-generated?
Which AI tools were used?
What were the provider's contractual terms?
Was open-source code incorporated?
Were licenses compatible?
Are training datasets documented?
Is human authorship adequately established?
Are there third-party claims?
Are patents implicated?
Was confidential information supplied to external AI systems?
Failure to investigate these questions may create significant post-acquisition liability.
35. The "Copyrightability Gap"
A particularly important risk is the possibility that:
AI-generated code is commercially valuable but insufficiently human-authored to receive copyright protection.
A company may invest heavily in automatically generated software only to discover that its ability to prevent competitors from copying portions of that software is limited.
This creates a copyrightability gap between technological value and legal exclusivity.
36. Human-AI Co-Creation
The most legally stable model is generally:
Human conceptual direction + AI assistance + human selection + human modification.
The more substantial the human contribution, the easier it may be to identify human authorship.
However, merely pressing "generate" and making trivial edits should not automatically be treated as sufficient authorship in every jurisdiction.
37. International Fragmentation
The problem becomes even more complex because different legal systems approach AI-generated works differently.
Potential differences concern:
human authorship;
originality;
employee ownership;
software copyright;
database rights;
moral rights;
patent inventorship;
training-data legality;
contractual rights.
A multinational company may therefore have:
copyright protection in one jurisdiction + uncertain protection in another.
This is genuine international IP fragmentation.
38. A Proposed Legal Framework
A useful analytical framework is:
Stage 1 — Identify the source
Was the code:
human-written;
AI-generated;
copied;
adapted;
open-source?
Stage 2 — Determine human contribution
What creative decisions did humans make?
Stage 3 — Separate functionality
Which elements are:
ideas;
algorithms;
APIs;
functional requirements?
Stage 4 — Conduct provenance analysis
Where did the generated code originate?
Stage 5 — Identify third-party rights
Check:
copyright;
patents;
trade secrets;
licenses.
Stage 6 — Examine contractual allocation
Who has contractual rights over the AI output?
Stage 7 — Assess jurisdiction
Which country's law governs?
39. Recommended Compliance Architecture
Organizations using AI-generated code should implement:
AI coding policies
Specify approved tools and prohibited uses.
Code provenance systems
Record AI-generated commits.
Open-source scanning
Automatically identify potentially incorporated open-source components.
Human review
Require developers to review AI-generated code.
License auditing
Track applicable licenses.
Security review
Check generated code for vulnerabilities.
IP review
Identify potential third-party claims.
Contractual safeguards
Address:
ownership;
confidentiality;
indemnification;
permitted use;
model training;
data retention.
40. Key Case-Law Principles
| Case | Principle relevant to AI-generated code |
|---|---|
| Feist v Rural | Originality |
| Computer Associates v Altai | Separating protectable expression from unprotectable elements |
| Google v Oracle | APIs, functionality and fair use |
| SAS Institute v World Programming | Ideas and functionality versus protected expression |
| BSA v Czech Republic | Scope of software copyright |
| Authors Guild v Google | Computational copying and transformation |
| Thaler v Perlmutter | Human authorship requirement |
| Naruto v Slater | Statutory authorship and non-human creators |
41. Core Legal Tensions
The central tensions can be summarized as follows:
Innovation vs exclusivity
AI dramatically lowers software-development costs, but may weaken conventional authorship assumptions.
Openness vs proprietary control
AI-generated code may depend upon open-source ecosystems while being incorporated into proprietary products.
Data access vs copyright
AI models require large datasets, but those datasets may contain copyrighted software.
Automation vs authorship
The more autonomous the machine, the more difficult traditional copyright ownership becomes.
Globalization vs national law
AI-generated software can be developed and distributed internationally, while IP protection remains substantially territorial.
42. Conclusion
Intellectual Property Fragmentation in Machine-Generated Code is fundamentally a problem of multiple overlapping sources of legal entitlement.
A single AI-generated software product may involve:
human copyright + uncertain machine-generated authorship + open-source licenses + third-party APIs + patent rights + trade secrets + contractual restrictions + training-data issues.
The most important cases—including Feist, Altai, Google v Oracle, SAS Institute, BSA, Authors Guild, Thaler and Naruto—show that existing IP law already contains several principles capable of addressing parts of the problem. However, those doctrines were developed largely for human-created software and conventional computing environments.
The emerging challenge is therefore to determine how those principles operate when human creativity, machine generation and third-party computational inputs are combined in a single software artifact.
The most defensible approach is to treat machine-generated code not as a legally homogeneous object but as a layered IP asset. Each layer should be examined separately for authorship, copyrightability, licensing, patent exposure, confidentiality and contractual rights.
Ultimately, the key question is no longer simply "Who wrote the code?" but:
Who contributed creatively, what pre-existing rights were incorporated, what contractual permissions govern the AI system, and which legal rights attach to each component of the resulting software?

comments