Competition Law And Ai Training Data Concentration .

 

Competition Law and AI Training Data Concentration

1. Introduction

AI training-data concentration describes a situation in which a relatively small number of firms control, possess, or can economically obtain the very large datasets needed to develop competitive artificial-intelligence models. The issue has become particularly important for foundation models and generative AI because competitive performance can depend not only on algorithms and computing power but also on access to sufficiently large, diverse, current, and legally usable datasets.

Competition authorities increasingly treat data as an important input into AI markets. The UK Competition and Markets Authority (CMA), for example, has examined access to data, computing power, expertise, distribution channels, and partnerships when considering competition in foundation-model markets. The U.S. Federal Trade Commission (FTC) has similarly identified data advantages and access to training data as issues deserving competition scrutiny.

There is not yet a large body of final judgments specifically deciding that concentration of AI training data itself violates competition law. Therefore, the most useful case law comes from earlier digital-platform, search, data-access, interoperability, tying, merger, and essential-input cases. These precedents provide the legal principles that authorities could apply to AI training-data concentration.

2. Why Training Data Matters for Competition

Training data can function as an important competitive input. A company operating a search engine, social network, marketplace, operating system, cloud platform, productivity suite, or other widely used digital service may continuously receive enormous amounts of information.

This may create a feedback mechanism:

Large user base → more data → potentially better AI products → more users/customers → still more data.

This cycle does not automatically constitute an antitrust violation. Competition law generally does not punish a company simply because it has accumulated valuable assets or developed a superior dataset through lawful competition.

The concern becomes greater where control of data is combined with conduct that excludes rivals—for example, discriminatory access conditions, exclusivity agreements, tying, self-preferencing, restrictive partnerships, acquisitions of important data sources, or practices that make it exceptionally difficult for customers or suppliers to switch.

The FTC's 2025 study of major cloud/AI partnerships specifically identified questions about competitive advantages arising from access to large datasets, including internally generated user data and licensing agreements. It also found that some AI/cloud partnerships involved sharing technical information, resources, and certain training data.

3. Relevant Market Definition

An AI training-data case would normally begin with defining the relevant market.

Authorities could potentially distinguish markets such as:

  • foundation-model development;
  • generative-AI services;
  • AI search services;
  • cloud AI infrastructure;
  • specialized training datasets;
  • data-licensing services;
  • particular downstream applications; or
  • specialized datasets such as scientific, financial, legal, linguistic, image, mapping, or commercial data.

The important question is whether alternative data sources are realistically substitutable.

For example, billions of publicly available webpages may not necessarily substitute for specialized medical, scientific, commercial, or real-time behavioural information. Consequently, authorities would examine the quality, scale, freshness, uniqueness, legality, cost and reproducibility of the relevant datasets.

4. Data Concentration Is Not Automatically Illegal

Competition law normally distinguishes between market power and abuse of market power.

A company might lawfully possess a uniquely valuable dataset because it created the service that generated the information. Competition problems are more likely to arise when additional conduct protects or extends that advantage artificially.

An authority would therefore investigate questions such as:

Can competitors reproduce the dataset? If equivalent information can easily be purchased or collected independently, concentration may present a smaller problem.

Are network effects involved? More users may produce more information, potentially improving a service and attracting still more users.

Are economies of scale significant? Collecting, cleaning, labeling and processing enormous datasets can require substantial investment.

Are there legal restrictions? Copyright, privacy, contractual rights and database protections may limit alternative sources.

Does the company control several important inputs simultaneously? A firm possessing data, cloud infrastructure, specialized chips, distribution channels and downstream applications could have a stronger strategic position than a firm controlling data alone.

The CMA's foundation-model work has consequently examined competition through multiple interconnected inputs rather than treating data in isolation.

5. Important Case Laws and Competition Precedents

Case 1: Google Search (Shopping) — European Commission / EU Courts

The European Commission's Google Shopping proceedings are highly relevant to AI competition even though they were not about AI training datasets.

The case concerned Google's treatment of its own comparison-shopping service relative to competing comparison-shopping services within Google search.

The broader principle is important: control over a strategically important digital gateway may raise competition concerns where a dominant company uses that position in a manner that disadvantages competing downstream services.

Relevance to AI Training Data

Imagine that a dominant digital platform possesses data generated by millions of users and develops its own AI model.

If it provides its own AI operation with privileged access to commercially important information while imposing materially different conditions on competing AI developers, competition authorities could investigate whether the arrangement represents exclusionary conduct.

The analogy must nevertheless be used carefully: Google Shopping concerned treatment and visibility in search results, not a general legal obligation to share proprietary datasets.

Case 2: Google Android — European Commission / General Court

The Android proceedings concerned contractual practices surrounding Google's Android ecosystem, including requirements associated with Google Search and Chrome and restrictions relating to alternative Android versions.

The General Court substantially confirmed the Commission's findings while modifying part of the decision and fine.

Relevance to AI

The case demonstrates why authorities examine ecosystems rather than isolated products.

An AI company may participate simultaneously in:

operating systems + cloud infrastructure + consumer applications + search + data collection + foundation models.

Training-data concentration can therefore become more significant when combined with distribution power.

For example, control over an operating system might give a company access to substantial user interactions, while integration of its AI assistant could increase usage and generate additional data.

Competition authorities could consequently investigate whether contractual or technical restrictions reinforce these interconnected advantages.

Case 3: Microsoft — European Commission

The historic Microsoft interoperability proceedings provide another important analogy.

Competition concerns included Microsoft's position in PC operating systems and the availability of interoperability information required by competing work-group server operating systems.

The case became an important European precedent concerning dominant digital ecosystems, interoperability and exclusionary practices.

Application to AI Data

Suppose competitors cannot develop effective AI products without access to particular interfaces or information controlled by an incumbent.

Authorities might investigate whether refusing access or providing degraded access prevents effective competition.

However, Microsoft does not establish a general principle that successful technology companies must provide competitors with their valuable data. Any access obligation would require the applicable legal requirements to be established.

Case 4: IMS Health GmbH & Co. OHG v NDC Health GmbH & Co. KG

Case C-418/01

IMS Health concerned access to a copyrighted structure used for pharmaceutical sales information.

The European Court of Justice considered the exceptional circumstances under which refusal by a dominant undertaking to license protected material could constitute abuse.

Important considerations included whether the requested input was indispensable and whether refusal prevented the emergence of a new product for which potential consumer demand existed, lacked objective justification and reserved a market to the rights holder.

Importance for AI Training Data

This case is especially relevant where a dataset or data structure is genuinely difficult or impossible to reproduce.

An AI developer could theoretically argue:

A dominant undertaking controls information that is indispensable for competing in a downstream AI market.

But the threshold is deliberately demanding.

Being merely useful, cheaper or advantageous would generally not establish indispensability. Authorities would need to examine whether realistic alternatives actually existed.

Case 5: Oscar Bronner GmbH v Mediaprint

Case C-7/97

Bronner is one of the major EU cases concerning refusal to provide access to infrastructure.

A newspaper publisher wanted access to a rival publisher's home-delivery system.

The Court adopted a restrictive approach to compulsory access. Among other considerations, it examined whether access was indispensable and whether realistic alternatives existed.

AI Application

Bronner is particularly important because it prevents an overly simple argument that:

"A company has valuable data, therefore it must share that data."

Competition law does not normally work that way.

For AI training data, an authority would have to investigate whether:

  1. alternative datasets exist;
  2. the requesting company could construct an alternative;
  3. synthetic or licensed information provides an adequate substitute;
  4. public data is sufficient;
  5. another commercial supplier could provide comparable information.

Therefore, Bronner provides an important limitation on expansive AI-data access claims.

Case 6: Slovak Telekom v European Commission

Case C-165/19 P

Slovak Telekom concerned conduct affecting competitors' access to telecommunications infrastructure.

One significant legal issue was the relationship between refusal-to-supply principles and other exclusionary conduct concerning access.

The Court clarified that the stringent Bronner indispensability conditions do not automatically govern every type of conduct connected with access to infrastructure.

Importance for AI

This distinction could become important in AI markets.

There is a difference between:

A. an AI competitor demanding completely new access to a company's proprietary dataset;

and

B. an incumbent already providing access but imposing discriminatory or exclusionary contractual conditions.

The first situation may raise strict refusal-to-deal principles.

The second could potentially be assessed under broader abuse-of-dominance principles depending upon the jurisdiction and facts.

Case 7: Google and DoubleClick — FTC Merger Review

The Google/DoubleClick transaction was reviewed by the FTC in the United States.

Although the FTC ultimately allowed the transaction to proceed, the matter became an important early example of debate surrounding combinations of large datasets in digital markets.

AI Training-Data Significance

Modern AI mergers may similarly require authorities to ask whether combining two companies also combines strategically valuable datasets.

For example:

Company A: major foundation-model developer.

Company B: possesses a unique dataset unavailable at comparable scale elsewhere.

A merger could potentially remove an independent supplier of data or provide the combined company with an input advantage.

Therefore, data concentration can be relevant not only under monopolization or abuse-of-dominance rules but also under merger control.

Case 8: Meta/Kustomer — European Commission Merger Review

The European Commission examined Meta's acquisition of Kustomer, a customer-relationship-management software provider.

The Commission ultimately cleared the transaction subject to commitments.

The proceeding is useful because it illustrates how competition authorities may examine digital ecosystems, access to important inputs and relationships between platform services and downstream businesses.

AI Relevance

A comparable AI transaction might involve:

large digital platform + specialized data company + AI model development.

Authorities would investigate whether post-merger control over an important input could disadvantage competing AI developers.

Potential remedies could focus on continued access, interoperability or non-discrimination where legally appropriate.

6. AI-Specific Regulatory Developments

The traditional cases above now operate alongside direct regulatory examination of generative-AI markets.

The FTC studied the relationships involving Microsoft–OpenAI, Amazon–Anthropic and Google–Anthropic. Its January 2025 report identified several potential competition issues, including access to computing resources and engineering talent, switching costs, exclusivity-related provisions and access to sensitive technical and business information.

Importantly for training-data concentration, FTC staff specifically identified as a subject for further investigation the competitive effects of data advantages enjoyed by firms possessing large datasets through sources such as internally generated user content or licensing arrangements.

The CMA has likewise examined foundation models through competition principles involving access, diversity of business models, choice, transparency and flexibility.

There is also an increasingly important distinction between competition law and ex-ante digital regulation. In July 2026, for example, the European Commission adopted measures under the Digital Markets Act concerning Google Search data sharing and Android interoperability with competing AI services. The Commission stated that eligible competing search services, including AI chatbots offering search functionality, can benefit from the search-data-sharing framework established under the DMA.

This is regulatory intervention under the DMA rather than a judgment establishing that AI training-data concentration itself constitutes an antitrust infringement.

7. Major Competition Risks

AI training-data concentration can create several distinct competition concerns.

Input foreclosure occurs where a vertically integrated company controls an important dataset and prevents downstream competitors from obtaining it.

Raising rivals' costs may arise where competitors technically receive access but only at significantly worse prices or conditions.

Exclusive data agreements can matter where an AI developer secures exclusive rights to particularly important datasets and comparable alternatives are scarce.

Data-driven network effects can arise where larger usage produces more useful information, which improves the service and attracts additional users.

Self-preferencing may become relevant where a platform controls both valuable information and competing AI services and systematically advantages its own downstream operation.

Merger-driven concentration may occur where acquisitions consolidate previously independent datasets.

Ecosystem leveraging can arise where advantages in search, operating systems, cloud services or productivity software strengthen a company's AI position.

Partnership concentration is another developing concern. The FTC found that major cloud/AI partnerships can involve substantial exchanges of computing resources, technical information and data and can create contractual or technical switching costs.

8. Possible Defences and Objective Justifications

Not every limitation on data access is anticompetitive.

Companies may have legitimate reasons for restricting access, including:

Privacy: personal information cannot simply be supplied to competitors.

Copyright and intellectual property: datasets may contain protected material or proprietary database structures.

Cybersecurity: unrestricted access could create security risks.

Investment incentives: compulsory sharing could reduce incentives to invest in collecting and improving datasets.

Quality control: companies may need restrictions to protect data integrity.

Contractual restrictions: the company itself may only possess limited licensing rights.

Technical capacity: providing continuous data access may require substantial infrastructure.

Competition authorities therefore generally have to balance exclusionary effects against legitimate commercial and technical explanations.

9. Potential Remedies

Where an infringement is actually established, remedies would depend heavily on the jurisdiction and specific conduct.

Possible approaches include requirements concerning non-discriminatory access, restrictions on exclusionary contractual clauses, interoperability obligations, data portability, limitations on exclusivity agreements, merger commitments, separation of certain datasets, or restrictions preventing information obtained from business partners from being used to disadvantage those partners.

Structural intervention would generally represent a considerably stronger remedy and would require the applicable legal standards to be satisfied.

Importantly, competition law is only one part of the regulatory framework. Privacy law, copyright law, trade-secret protection, database rights and AI-specific regulation can constrain what data-sharing remedies are legally possible.

10. Overall Legal Position

The central competition-law issue is not simply whether one company possesses more training data than another.

The more important questions are:

How did the company obtain the data?

Can competitors realistically reproduce or substitute for it?

Does the company possess market power?

Is it restricting competitors' access through exclusionary conduct?

Are exclusivity agreements locking up important sources?

Does vertical integration allow data advantages to be leveraged into adjacent AI markets?

Would compulsory access undermine legitimate privacy, IP, security or investment interests?

The established cases—Google Shopping, Google Android, Microsoft, IMS Health, Bronner, Slovak Telekom, Google/DoubleClick and Meta/Kustomer—provide useful legal frameworks, but they should not be misdescribed as judgments specifically finding unlawful concentration of generative-AI training data.

Current regulatory work is filling that gap. The FTC has expressly identified large-dataset advantages as an area requiring further investigation, while the CMA has incorporated data and other critical inputs into its examination of foundation-model competition. The EU's 2026 DMA implementation also demonstrates a movement toward specific access and interoperability requirements for certain powerful digital gatekeepers.

LEAVE A COMMENT