Ai Model Training Data Competition Issues
AI Model Training Data Competition Issues
1. Introduction
AI model training data has become a critical competitive input in the development and deployment of large language models, generative AI systems, recommendation engines, computer-vision systems, autonomous technologies, and other machine-learning applications. Access to large, diverse, high-quality, timely and legally usable datasets can substantially affect model performance, development costs, product quality and the ability of new firms to enter AI markets.
From a competition-law perspective, the central issue is not simply whether a company possesses large quantities of data. The relevant questions include whether the data constitutes a competitively significant input, whether rivals can reasonably obtain substitutes, whether access is being restricted through exclusionary conduct, whether data advantages are reinforced by network effects, and whether mergers or contractual arrangements allow a firm to foreclose competing AI developers.
The principal competition concerns include:
- accumulation of unique or difficult-to-replicate training datasets;
- exclusive data-acquisition agreements;
- refusal to provide access to commercially essential datasets;
- discriminatory data licensing;
- tying data access to cloud, platform or AI services;
- self-preferencing in data collection;
- data scraping restrictions used strategically to exclude rivals;
- exclusive partnerships between AI developers and content owners;
- acquisitions of companies possessing strategically important datasets;
- combining user data from multiple markets;
- algorithmic advantages generated by proprietary feedback data; and
- leveraging data advantages from an adjacent market into an emerging AI market.
2. Why Training Data Can Be a Competition Issue
Training data may function as an important input into AI model development.
A simplified AI value chain is:
Data Sources → Data Collection → Data Cleaning/Annotation → Training → Model Development → Fine-Tuning → Deployment → User Feedback → New Data
A firm controlling several stages can potentially create a data feedback loop:
More Users → More Data → Better Model → Better Product → More Users
This may produce competitive advantages that become progressively harder for rivals to overcome.
However, possession of data does not automatically establish market power. Competition authorities would ordinarily examine:
- the relevant product and geographic markets;
- the substitutability of the dataset;
- the uniqueness and quality of the data;
- the cost and legality of obtaining alternative data;
- the speed at which rivals can reproduce the dataset;
- the importance of the data to model performance;
- the firm's position in related markets;
- contractual restrictions on access;
- network and feedback effects; and
- whether the conduct produces exclusionary effects.
3. Relevant Competition-Law Theories
A. Abuse of Dominance
A dominant AI platform possessing a strategically important dataset could potentially engage in exclusionary conduct by:
- refusing access to an important dataset;
- imposing discriminatory licensing terms;
- charging excessive or discriminatory access prices;
- preventing customers from supplying data to competitors;
- restricting interoperability;
- imposing exclusivity obligations; or
- tying access to data with another AI or cloud service.
The difficult question is whether the data is sufficiently important that its restriction materially reduces competition.
4. Data as an Essential or Difficult-to-Replicate Input
The traditional essential-facilities doctrine may become relevant where a dataset is:
- uniquely valuable;
- controlled by a dominant undertaking;
- practically impossible or excessively costly for competitors to reproduce;
- indispensable for competing in the relevant market; and
- capable of being supplied without eliminating legitimate incentives to innovate.
Courts have traditionally applied essential-facilities principles cautiously. Therefore, merely demonstrating that a dataset is useful will generally not be enough.
The stronger argument arises where:
No reasonably viable alternative dataset exists and denial of access substantially prevents effective competition.
5. Data Exclusivity Agreements
A major concern is an agreement under which an AI developer obtains exclusive access to:
- books;
- news archives;
- scientific publications;
- photographs;
- videos;
- consumer reviews;
- maps;
- transaction data;
- financial information;
- industrial datasets; or
- user-generated content.
For example, if an AI developer enters into exclusive agreements with the largest sources of high-quality specialized data, competing developers may face substantially higher costs.
Competition authorities may examine:
- duration of exclusivity;
- market coverage;
- availability of alternative data;
- switching possibilities;
- importance of the data;
- foreclosure percentage;
- entry barriers; and
- efficiencies generated by the agreement.
6. Data Licensing Discrimination
A dominant data owner may provide:
- inexpensive access to its own AI subsidiary;
- expensive access to competitors;
- high-quality data to affiliated companies;
- delayed access to rivals; or
- different technical access conditions.
This can create vertical foreclosure.
For example:
Data Owner → Own AI Model: Full Dataset
but
Data Owner → Rival AI Models: Limited Dataset
If the difference cannot be justified by legitimate commercial considerations, competition concerns may arise.
7. Data Tying and Bundling
A platform may condition access to valuable data on the purchase of another service.
Examples include:
Training dataset + cloud computing
or
Search data + advertising service
or
User-generated data + platform distribution
or
Enterprise data + proprietary AI infrastructure.
Such arrangements can potentially leverage dominance from one market into another.
The competition analysis would examine whether:
- the firm is dominant in the tying market;
- the products are distinct;
- customers are effectively compelled to obtain the tied product;
- competitors are foreclosed; and
- the arrangement produces sufficient efficiencies to justify the conduct.
8. Data Feedback Loops
AI markets present an unusual competitive phenomenon: data can improve the product that generates additional data.
For example:
AI Assistant → More Users → More Queries → More Feedback → Better Model → More Users
A firm with a large installed user base may therefore have an advantage unavailable to smaller competitors.
Competition authorities may examine whether contractual or technical restrictions prevent competitors from obtaining equivalent feedback.
Relevant forms of feedback include:
- prompts;
- corrections;
- user ratings;
- rejected outputs;
- preference signals;
- conversational interactions;
- error reports;
- human annotations; and
- post-deployment performance information.
9. Data Portability and Switching
Competition concerns may arise when customers cannot transfer their relevant data from one AI platform to another.
For example:
Customer → AI Platform A
The platform accumulates:
- prompts;
- fine-tuning data;
- evaluation datasets;
- customized workflows;
- embeddings;
- feedback;
- model-performance information.
If customers cannot export these assets in usable formats, switching to another AI provider becomes more difficult.
This can strengthen incumbency advantages.
10. Scraping Restrictions and Competition
The relationship between web scraping and competition law is particularly complex.
A platform may restrict automated collection of its content through:
- technical barriers;
- robots.txt;
- API restrictions;
- contractual terms;
- authentication requirements;
- rate limits; or
- anti-bot technologies.
These restrictions may be legitimate intellectual-property, privacy, cybersecurity or commercial measures.
However, competition concerns could arise if a dominant platform uses control over data access strategically to exclude competitors while simultaneously using substantially equivalent information for its own AI system.
The legal analysis must therefore distinguish:
legitimate protection of proprietary resources
from
anticompetitive exclusion of competing AI developers.
11. AI Training Data and Merger Control
Competition authorities may examine acquisitions where the target possesses:
- proprietary datasets;
- large user communities;
- specialized scientific data;
- proprietary annotation capabilities;
- valuable behavioral data;
- data-generation technology; or
- access to otherwise difficult-to-obtain information.
The acquisition of a relatively small company may nevertheless have competitive significance if its principal strategic asset is a unique dataset.
Traditional turnover thresholds may therefore fail to capture some data-driven acquisitions, leading jurisdictions to consider transaction-value thresholds or other theories of intervention.
12. Data Combination Across Markets
Suppose a company operates:
- a search engine;
- social-media services;
- cloud computing;
- an advertising platform; and
- an AI model.
Combining data from all these businesses could create a powerful informational advantage.
Competition authorities may ask whether:
Data accumulated through dominance in one market is being used to strengthen market power in another market.
The issue becomes particularly important where competitors cannot replicate the same cross-market dataset.
13. Privacy and Competition
Privacy and competition can overlap.
A reduction in privacy protections may sometimes function as a non-price dimension of competition.
For example, consumers may prefer:
AI Service A → less extensive data collection
over
AI Service B → extensive data collection
If a dominant platform eliminates privacy-enhancing alternatives, the competition analysis may consider deterioration in quality even where monetary prices remain zero.
This reasoning has appeared prominently in digital-platform competition cases.
14. Important Case Laws
The following cases do not all concern AI training datasets directly. They provide established competition-law principles concerning data, digital platforms, access to inputs, exclusion, interoperability, tying, network effects and leveraging, which can be applied to AI-training-data disputes.
Case 1: Google Search (Shopping) – European Commission
Google Search (Shopping), Case AT.39740
The European Commission found that Google abused its dominant position by giving preferential treatment to its own comparison-shopping service in general search results.
Relevance to AI training data
The case demonstrates the importance of:
- control over a major digital platform;
- preferential treatment of affiliated services;
- leveraging platform advantages into adjacent markets; and
- foreclosure of competing digital services.
For AI, a comparable concern could arise if a dominant platform gives its affiliated AI system preferential access to platform-generated data while restricting equivalent access to competing AI developers.
Principle
A platform controlling a strategically important digital interface may face competition-law scrutiny where it uses that control to advantage its own downstream service.
Case 2: Google Android – European Commission
Google Android, Case AT.40099
The European Commission examined Google's conduct concerning Android, including tying and restrictions affecting competing services.
Relevance
The case is important for AI because AI ecosystems increasingly combine:
- operating systems;
- app stores;
- cloud infrastructure;
- search;
- advertising;
- AI assistants; and
- user data.
A dominant platform could potentially use control over one layer to restrict competing AI systems at another layer.
Principle
Competition law can address strategies through which dominance in one technological layer is leveraged into adjacent markets.
Case 3: Microsoft – European Commission
Microsoft, Case COMP/C-3/37.792
The Microsoft proceedings concerned Microsoft's conduct involving interoperability information and tying.
The European Commission required Microsoft to provide interoperability information to competing work-group server products and addressed the linking of products within Microsoft's software ecosystem.
Relevance to AI
The case provides an important framework for considering whether a dominant technological ecosystem can restrict access to information or interoperability necessary for competing products.
AI analogues may include:
- training-data interoperability;
- model portability;
- access to data APIs;
- compatibility with AI infrastructure; and
- restrictions preventing competing models from accessing relevant platform resources.
Principle
Control over technological interoperability can become a competition issue when it materially restricts effective competition.
Case 4: IMS Health v NDC Health
IMS Health GmbH & Co. OHG v NDC Health GmbH & Co. KG
This European competition-law litigation concerned access to pharmaceutical sales-information structures and the application of the essential-facilities doctrine.
The Court of Justice emphasized stringent conditions before refusal to license intellectual property can constitute abuse.
Relevance to AI training data
The case is particularly important where an AI developer argues:
"The dataset is indispensable, and therefore the dominant data owner must provide access."
The decision demonstrates that indispensability is a demanding requirement.
A rival would need to establish much more than the fact that the dataset is commercially valuable.
Principle
Refusal to license a protected resource is not automatically abusive; strict conditions concerning indispensability, elimination of competition and lack of objective justification are relevant.
Case 5: Bronner v Mediaprint
Oscar Bronner GmbH & Co. KG v Mediaprint Zeitungs und Zeitschriftenverlag GmbH & Co. KG
This case concerned access to a newspaper distribution system.
The Court of Justice established a restrictive approach to compulsory access under the essential-facilities doctrine.
Relevance to AI
The analogy can be drawn to:
- proprietary datasets;
- AI training repositories;
- cloud infrastructure;
- specialized data marketplaces; and
- platform APIs.
A competitor cannot ordinarily demand access merely because constructing an alternative system is difficult or expensive.
Principle
The resource must generally be indispensable, and refusal must be capable of eliminating effective competition rather than merely making competition harder.
Case 6: Slovak Telekom v European Commission
Slovak Telekom a.s. v European Commission
The litigation concerned access to telecommunications infrastructure and exclusionary conduct by a dominant undertaking.
Relevance to AI
The case provides an important framework for situations in which a dominant firm controls an upstream infrastructure layer and competitors depend upon that infrastructure.
In AI markets, analogous infrastructure may include:
- training datasets;
- data-access APIs;
- cloud infrastructure;
- model-serving infrastructure;
- specialized annotation systems.
Principle
Dominance over an upstream bottleneck can create competition concerns when access conditions substantially restrict downstream competition.
Case 7: Deutsche Telekom v Commission
Deutsche Telekom AG v European Commission
The case concerned pricing and access conditions in telecommunications infrastructure.
Relevance
Although it did not concern AI data, the case is useful for analysing situations where a vertically integrated undertaking controls an important upstream input while simultaneously competing downstream.
An AI ecosystem could produce a comparable structure:
Data Platform → Training Data → AI Model → AI Application
If the platform provides an input to independent AI developers while operating its own competing AI service, discriminatory access conditions can become particularly significant.
Principle
Vertical integration can intensify competition concerns where the dominant undertaking controls an input used by downstream competitors.
Case 8: Facebook/Meta Data-Related Competition Proceedings
Competition authorities have examined the relationship between Facebook's market position, data collection and its advertising ecosystem.
The German competition proceedings involving Facebook were particularly significant because they addressed the combination of data obtained from different services and the relationship between market power and data-processing conditions.
Relevance to AI
The case illustrates how:
- data aggregation;
- cross-service data combination;
- platform dominance; and
- consumer data practices
can intersect with competition law.
For AI, similar issues may arise when a company combines data from multiple products to train or improve an AI system that competes with firms lacking equivalent data access.
Principle
Data accumulation and combination can become relevant to competitive assessment where they reinforce an undertaking's market position.
15. Case-Law Principles Applied to AI Training Data
| Competition issue | Relevant case-law principle |
|---|---|
| Refusal to provide indispensable data | IMS Health; Bronner |
| Control of essential infrastructure | Slovak Telekom; Deutsche Telekom |
| Leveraging platform dominance | Google Shopping |
| Tying technological services | Google Android; Microsoft |
| Interoperability restrictions | Microsoft |
| Cross-service data aggregation | Facebook/Meta proceedings |
| Vertical foreclosure | Slovak Telekom; Deutsche Telekom |
| Preferential treatment of own AI | Google Shopping analogy |
| Data access discrimination | Infrastructure-access jurisprudence |
| Proprietary dataset access | Essential-facilities jurisprudence |
16. Competition Concerns in AI Training Data Markets
A. Data Hoarding
A dominant undertaking may accumulate large quantities of data without making it available to competitors.
The key question is whether the accumulation itself results from legitimate competition or from exclusionary strategies.
B. Exclusive Data Partnerships
Exclusive agreements can prevent rivals from obtaining:
- high-quality text;
- specialist scientific information;
- premium audiovisual content;
- proprietary consumer data; or
- domain-specific datasets.
Long-term exclusivity covering a substantial proportion of commercially useful data can increase entry barriers.
C. Data Quality Advantage
Not all datasets are equivalent.
Competitive significance may depend upon:
- accuracy;
- diversity;
- freshness;
- geographic coverage;
- language coverage;
- domain specificity;
- labeling quality;
- temporal depth; and
- feedback quality.
Thus, one billion low-quality records may be less competitively important than a much smaller unique dataset.
17. Data Annotation as a Competitive Bottleneck
Training-data competition also extends beyond raw data.
AI developers require:
- human labeling;
- reinforcement-learning feedback;
- safety classification;
- preference ranking;
- domain-specific annotation; and
- expert validation.
If one AI company obtains exclusive access to scarce expert annotators or annotation providers, competitors may face higher costs.
This may create a secondary data-processing bottleneck.
18. Synthetic Data and Competition
Synthetic data can reduce dependence on scarce human-generated datasets.
However, synthetic-data ecosystems may themselves become concentrated.
For example:
Dominant Model → Generates Synthetic Data → Trains New Models → New Models Depend on Dominant Model
This creates a potential circular dependency.
Competition authorities may therefore need to consider whether synthetic-data generation merely reduces entry barriers or instead creates another layer of dependency on an incumbent model.
19. Data Poisoning and Strategic Data Manipulation
Competitive concerns may also arise where firms deliberately manipulate data environments.
Potential practices include:
- inserting misleading training material;
- deliberately degrading datasets;
- flooding data repositories with low-quality material;
- restricting access to high-quality corrections;
- manipulating feedback systems; or
- preventing competitors from verifying dataset quality.
Such conduct may become relevant where it has the purpose or effect of weakening competing AI systems.
20. Data Scraping and Exclusionary Conduct
A dominant platform may simultaneously:
- prevent rivals from scraping its content;
- restrict API access;
- impose contractual limitations; and
- use the same content internally to train its AI system.
The competition question becomes whether the restriction is:
legitimate protection of proprietary content
or
strategic exclusion of competitors.
Intellectual-property law, contract law, privacy law and competition law may overlap.
21. Consumer Data as a Competitive Asset
Consumer-facing AI services may generate enormous quantities of behavioral information.
Examples include:
- prompts;
- preferences;
- corrections;
- interaction histories;
- purchasing patterns;
- search behavior;
- clicks;
- ratings; and
- usage frequency.
This creates a potential dynamic data advantage.
The competitive advantage may therefore increase over time even if the original dataset was not uniquely valuable.
22. Data Portability as a Competition Remedy
Potential remedies could include:
- data portability;
- API access;
- interoperability requirements;
- non-discrimination obligations;
- licensing commitments;
- restrictions on exclusivity;
- data silos;
- divestiture of datasets in merger cases; and
- prohibition of discriminatory data access.
However, mandatory access remedies must account for:
- privacy;
- cybersecurity;
- intellectual-property rights;
- data provenance;
- confidential information;
- incentives to collect data; and
- technical feasibility.
23. Merger Remedies
Where a merger creates significant data concentration, authorities could potentially consider:
Structural remedies
- divestiture of a dataset;
- divestiture of a data-generating business;
- separation of competing AI operations.
Behavioral remedies
- non-exclusive licensing;
- API access;
- interoperability;
- data portability;
- non-discrimination;
- restrictions on cross-use of data.
The appropriate remedy would depend upon the competitive theory established in the particular case.
24. India-Specific Competition Perspective
In India, the principal framework is the Competition Act, 2002, administered by the Competition Commission of India.
AI training-data practices may potentially engage:
- Section 3 — anti-competitive agreements;
- Section 4 — abuse of dominant position;
- Section 5 — combinations;
- Section 6 — regulation of combinations;
- relevant CCI regulations and merger-control principles.
Potential Section 4 theories could include:
- discriminatory access to data;
- denial of market access;
- tying or bundling;
- leveraging dominance;
- exclusionary contractual arrangements; and
- discriminatory conditions imposed on competing AI developers.
For Section 3, attention may be given to:
- exclusive supply arrangements;
- exclusive distribution arrangements;
- refusal-to-deal arrangements;
- information-sharing arrangements;
- coordinated data acquisition; and
- agreements restricting access to competitively important datasets.
25. Indian Competition-Law Analogies
Indian competition jurisprudence concerning digital platforms, market access, interoperability, vertical restraints and dominance provides useful analytical foundations for AI-data disputes.
The most important questions would include:
- What is the relevant market?
- Does the undertaking possess dominance?
- Is the dataset commercially indispensable?
- Are substitutes realistically available?
- Is access being denied or merely commercially priced?
- Does the restriction foreclose competitors?
- Is there objective justification?
- Are privacy or IP interests involved?
- Does the conduct leverage dominance into another market?
- Are efficiencies sufficient to explain the conduct?
26. Key Legal Test for AI Training Data
A useful analytical framework is:
Step 1 — Define the market
Identify whether the relevant market concerns:
- raw training data;
- specialized datasets;
- data annotation;
- foundation models;
- AI APIs;
- cloud computing; or
- downstream AI applications.
Step 2 — Identify the data advantage
Determine:
- uniqueness;
- quality;
- scale;
- freshness;
- exclusivity;
- provenance; and
- replicability.
Step 3 — Examine market power
Consider:
- market share;
- entry barriers;
- network effects;
- switching costs;
- vertical integration;
- access to users; and
- alternative data sources.
Step 4 — Identify exclusionary conduct
Examples:
Refusal → Discrimination → Exclusivity → Tying → Self-preferencing → Bundling → Data withholding
Step 5 — Examine competitive effects
Ask whether the conduct:
- raises rivals' costs;
- prevents entry;
- reduces innovation;
- restricts model quality;
- increases switching costs;
- eliminates competing business models; or
- reinforces network effects.
Step 6 — Consider objective justification
Possible justifications include:
- privacy;
- cybersecurity;
- intellectual-property protection;
- confidentiality;
- data quality;
- technical limitations;
- contractual obligations; and
- legitimate investment incentives.
27. Major Competition Risks
The principal AI-training-data risks can therefore be summarized as:
1. Dataset concentration
A small number of firms control unique datasets.
2. Exclusive access
Major datasets are contractually reserved for one AI developer.
3. Data discrimination
Affiliated AI systems receive superior access.
4. Data foreclosure
Competitors cannot obtain sufficiently comparable information.
5. Cross-market leveraging
Data from a dominant platform is used to strengthen an AI market position.
6. Feedback-loop advantages
Large user bases continuously generate additional training information.
7. Data portability barriers
Users cannot transfer relevant information to competing AI providers.
8. Annotation concentration
Specialized human labeling capacity becomes concentrated.
9. Merger-driven concentration
Acquisitions combine datasets that previously belonged to competing ecosystems.
10. Ecosystem lock-in
Data, cloud, applications and AI models become integrated into a single ecosystem.
28. Conclusion
AI training data is increasingly capable of functioning as a strategic competitive input. Competition law, however, should not treat every large dataset as an essential facility or every refusal to license data as abusive.
The central legal distinction is between legitimate data-based competition and strategic use of data control to exclude competitors.
The jurisprudence in Google Shopping, Google Android, Microsoft, IMS Health, Bronner, Slovak Telekom, Deutsche Telekom and the Facebook/Meta proceedings provides useful principles concerning platform leveraging, interoperability, refusal to deal, essential facilities, vertical foreclosure, tying and data aggregation.
For AI markets, the most significant future disputes are likely to concern the intersection of training-data exclusivity, data portability, proprietary datasets, cross-platform data combination, feedback loops, AI/cloud vertical integration, merger control and access to indispensable data resources.
The emerging legal question is therefore not simply “Who owns the data?”, but:
Whether control over strategically important training data enables an undertaking to acquire, maintain or extend market power by restricting effective competition in AI markets.

comments