Competition Law And Cognitive Automation Market Power
Competition Law and Synthetic Data Market Concentration
1. Introduction
Synthetic data is artificially generated information designed to reproduce important statistical or structural characteristics of real-world data. It can include synthetic images, text, financial records, customer behaviour, medical-style datasets, sensor information, or simulated environments used to train and test artificial-intelligence systems.
From a competition-law perspective, synthetic data is important because access to large quantities of high-quality training data is increasingly an important input in AI development. Competition authorities have already identified control over data, computing resources, technical expertise, and distribution channels as potential sources of market power in AI markets. For example, the UK Competition and Markets Authority (CMA) has identified concentrated control of critical AI inputs as a potential competition risk.
There is not yet a large body of decided antitrust cases specifically concerning a standalone synthetic-data market. Therefore, the most relevant precedents come from competition cases concerning control of datasets, digital platforms, interoperability, essential inputs, technological ecosystems, and data-driven mergers. These cases provide the legal principles likely to be applied if synthetic-data markets become highly concentrated.
2. Meaning of Synthetic Data Market Concentration
Market concentration occurs when a relatively small number of companies control a substantial part of a market.
In synthetic data, concentration could arise where a few firms control important parts of the production chain, such as:
- proprietary real-world datasets used to create synthetic data;
- synthetic-data generation software;
- foundation models capable of generating datasets;
- cloud infrastructure and AI computing capacity;
- simulation platforms;
- data validation and quality-assurance technology;
- specialised technical expertise;
- distribution platforms and marketplaces for datasets.
This is particularly important because the AI value chain can involve the same large firms at several different levels. The CMA has observed major technology firms operating across compute, data, foundation-model development, partnerships and downstream deployment.
Consequently, competition analysis cannot simply ask who sells synthetic datasets. Authorities may have to examine the complete technological ecosystem.
3. Why Synthetic Data Could Become Concentrated
Economies of Scale
Generating sophisticated synthetic datasets can require significant computing resources, technical expertise and access to high-quality source data.
Large companies may therefore produce synthetic data at lower average cost than small competitors.
This creates a potential cycle:
More data → better synthetic generation → better AI products → more customers → more data and revenue → further investment.
If sufficiently strong, these advantages can make entry difficult.
Access to Real-World Data
Synthetic data does not necessarily eliminate the competitive importance of real data.
A company may need representative real-world information to train, calibrate, evaluate or validate a synthetic-data generator. Consequently, firms possessing extensive proprietary datasets may retain important competitive advantages.
Competition authorities have specifically identified access to data at scale as one of the critical inputs in foundation-model competition.
Computing Infrastructure
Large-scale synthetic-data generation can also depend upon cloud computing and specialised AI chips.
If the firms controlling important datasets also have substantial positions in cloud computing or AI models, vertical integration may strengthen their competitive position.
The CMA has warned more generally that firms controlling critical AI inputs may potentially restrict access in ways that protect them from competition.
4. Relevant Market Definition
A major competition-law question would be defining the relevant market.
Authorities could consider relatively broad markets such as:
AI training data services
or narrower markets such as:
synthetic image datasets for autonomous systems
synthetic financial datasets
synthetic-language training datasets
synthetic-data generation platforms
simulation environments for robotics
The correct definition would depend on whether customers consider real data, synthetic data and alternative generation technologies sufficiently interchangeable.
Authorities generally examine both:
Product Market
Which products customers regard as reasonable substitutes.
Geographic Market
Whether competitive conditions are national, regional or worldwide.
Synthetic-data markets may often have an international dimension because digital datasets can be supplied electronically across borders, although regulation, language, privacy requirements and industry-specific standards can sometimes create narrower markets.
5. Horizontal Concentration
Horizontal concentration can arise where competing synthetic-data providers merge or where one repeatedly acquires smaller competitors.
Suppose four companies are significant suppliers of synthetic datasets for autonomous-vehicle development. If two merge, competition authorities could examine whether the transaction would:
- eliminate an important competitive constraint;
- increase prices;
- reduce dataset quality;
- reduce innovation;
- decrease customer choice;
- reduce incentives to develop alternative generation techniques.
In rapidly developing technology markets, competition authorities can also consider innovation competition rather than focusing exclusively on existing revenues.
6. Vertical Concentration
Vertical concentration may be especially significant.
Consider a company operating simultaneously in:
Cloud infrastructure → Foundation models → Synthetic-data generation → AI development tools → AI applications.
Such integration can produce genuine efficiencies. But it can also create opportunities to disadvantage competitors.
Possible conduct could include:
- denying competitors access to synthetic datasets;
- charging discriminatory prices;
- providing inferior API access;
- making synthetic-data tools work better with the firm's own AI services;
- bundling datasets with cloud services;
- imposing restrictive contractual conditions.
The CMA's work on foundation models similarly identifies the possibility that integrated firms or powerful partnerships could restrict rivals' access to critical inputs.
7. Data Foreclosure
A particularly important theory of harm is input foreclosure.
Imagine Company A possesses unique industrial information from millions of machines. It uses that information to develop highly accurate synthetic manufacturing datasets.
If competing AI developers cannot obtain comparable real or synthetic data, Company A could potentially gain substantial market power.
Competition authorities would examine:
- whether the input is competitively important;
- whether alternatives exist;
- whether the company has market power;
- whether it can restrict competitors' access;
- whether it has an economic incentive to do so;
- whether the restriction materially harms downstream competition.
Simply possessing valuable data does not automatically violate competition law. The concern normally arises from market power combined with conduct or transactions capable of materially restricting competition.
8. Network Effects and Feedback Loops
Synthetic-data platforms could develop indirect network effects.
For example:
More customers → more feedback → improved generation system → better datasets → more customers.
Once sufficiently developed, this can increase entry barriers.
The concern becomes stronger where a company also controls complementary assets such as cloud infrastructure, software-development tools or distribution channels.
The CMA's AI principles therefore emphasise continued access to critical inputs and the importance of preventing early advantages, economies of scale and feedback loops from becoming disproportionate barriers to effective competition.
Important Case Laws
Because dedicated synthetic-data antitrust litigation remains limited, the following cases provide the closest established principles for analysing synthetic-data concentration.
1. Google/DoubleClick — European Commission, COMP/M.4731 (2008)
Google's acquisition of DoubleClick raised important questions concerning online advertising and data.
The European Commission examined whether combining Google's position with DoubleClick's technology could create competition problems.
The transaction was ultimately approved.
Relevance to Synthetic Data
The case demonstrates that authorities may examine whether combining datasets, technology and existing market positions creates competitive advantages unavailable to rivals.
In synthetic-data markets, similar questions could arise where an AI company acquires a business possessing unique datasets necessary for producing high-quality synthetic information.
2. Facebook/WhatsApp — European Commission, Case M.7217 (2014)
The European Commission reviewed Facebook's acquisition of WhatsApp.
Among other issues, the Commission considered online communications, social networking and online advertising and examined the significance of user data.
The acquisition was cleared under EU merger rules.
Relevance
The case illustrates that control over large quantities of user information can form part of competition analysis even where data itself is not treated as a separate product market.
Similarly, a synthetic-data merger might require authorities to examine the competitive significance of the underlying information available to the merging companies.
3. Microsoft/LinkedIn — European Commission, Case M.8124 (2016)
Microsoft's acquisition of LinkedIn involved the combination of a major software ecosystem with a large professional social network and substantial professional-user information.
The Commission approved the transaction subject to commitments addressing particular competition concerns.
Relevance
This case is particularly useful for understanding ecosystem effects.
A synthetic-data company could obtain advantages when combined with:
- operating systems;
- cloud infrastructure;
- professional datasets;
- AI development tools;
- distribution networks.
Authorities may therefore examine the entire ecosystem rather than looking at synthetic data independently.
4. Google Search (Shopping) — European Commission, Case AT.39740 (2017)
The European Commission found that Google had abused its dominant position in general internet search by giving favourable treatment to its own comparison-shopping service relative to competing services.
The General Court later largely upheld the Commission's decision.
Relevance
The principle is important where a vertically integrated synthetic-data platform operates both the infrastructure and competing downstream products.
For example, competition concerns could potentially arise if a dominant AI platform systematically advantages its own synthetic datasets while making rival datasets materially harder for customers to discover or use.
5. Google Android — European Commission, Case AT.40099 (2018)
The European Commission concluded that Google imposed certain contractual restrictions associated with Android and Google's applications that unlawfully strengthened its search position.
Relevance
Synthetic-data platforms may similarly develop ecosystems containing:
data + model + API + cloud + application marketplace.
Competition law can examine contractual arrangements that make it difficult for customers or developers to use competing services where the relevant legal requirements for abuse of dominance are satisfied.
6. Microsoft — European Commission, Case COMP/C-3/37.792 (2004)
The European Commission found competition-law violations concerning Microsoft's conduct relating to interoperability information and Windows Media Player.
One important issue concerned competitors' ability to obtain information necessary for interoperability with Microsoft's systems.
Relevance
Interoperability could become crucial in synthetic-data markets.
A dominant synthetic-data platform could potentially disadvantage rivals if proprietary formats or interfaces prevent datasets from working effectively with competing AI systems.
The Microsoft proceedings therefore provide important principles concerning interoperability, technological ecosystems and exclusionary conduct by dominant firms.
7. IMS Health — Case C-418/01, Court of Justice of the European Union (2004)
IMS Health concerned intellectual-property rights relating to a structure used for pharmaceutical sales data.
The Court addressed the exceptional circumstances under which refusal to license intellectual property by a dominant company can constitute abuse.
Relevance
Synthetic datasets may be protected through copyright, database rights, contracts, trade secrets or other intellectual-property mechanisms depending upon jurisdiction and circumstances.
IMS Health is therefore relevant where competitors claim that access to a proprietary data structure or related protected resource is indispensable.
However, competition law does not normally create a general obligation to share valuable proprietary data. Compulsory-access principles apply only under demanding legal conditions.
8. Slovak Telekom — Case C-165/19 P (CJEU, 2021)
This case concerned access to telecommunications infrastructure and alleged exclusionary conduct by a dominant undertaking.
The Court clarified aspects of the relationship between refusal-to-supply principles and other forms of conduct concerning access.
Relevance
This distinction could matter significantly for synthetic data.
Competition law may distinguish between:
- an outright refusal to provide a proprietary dataset; and
- discriminatory or unfair conditions imposed where access is already being supplied.
The applicable legal test can therefore depend upon the precise nature of the allegedly exclusionary behaviour.
9. Merger Control and Synthetic Data
Synthetic-data concentration may increasingly become relevant to merger authorities.
Authorities could investigate transactions involving:
Synthetic-data company + cloud provider
AI developer + specialist dataset provider
Foundation-model developer + simulation company
Technology platform + synthetic-data startup
A merger may be problematic where it eliminates an important competitive constraint or gives the merged company the ability and incentive to foreclose competitors.
The broader AI market already contains extensive investment and partnership relationships. The CMA reported an interconnected network of more than 90 partnerships and strategic investments involving major technology and AI firms in its 2024 analysis.
That does not mean such partnerships are inherently anticompetitive. Authorities recognise that investments and partnerships can supply funding, compute and expertise and can therefore increase competition as well as potentially restrict it.
10. Abuse of Dominance
A company possessing a dominant position in a properly defined synthetic-data market would not violate competition law merely because it is large or successful.
The central question is normally how that position is used.
Potential theories of abuse could include:
Exclusive dealing: customers are prevented from obtaining datasets from competitors.
Tying: access to important synthetic data requires purchasing unrelated cloud or AI services.
Discriminatory access: downstream competitors receive inferior access compared with the dominant company's own business.
Interoperability restrictions: competitors cannot integrate their datasets effectively.
Predatory conduct: pricing is allegedly structured to exclude efficient competitors and later exploit reduced competition, subject to the jurisdiction's legal test.
Self-preferencing: an integrated platform advantages its own synthetic-data services.
Each theory requires proof under the relevant jurisdiction's competition rules; concentration alone is insufficient.
11. Intellectual Property Versus Competition
Synthetic data creates an important tension between intellectual-property protection and competition.
Companies need incentives to invest in expensive data-generation technologies. Strong protection can encourage:
- research;
- dataset creation;
- privacy-enhancing technologies;
- better simulation;
- improved AI models.
However, highly restrictive control over strategically important datasets can sometimes contribute to entry barriers.
Competition law therefore generally seeks to distinguish legitimate exploitation of intellectual property from exclusionary conduct by firms possessing substantial market power.
Cases such as IMS Health and Microsoft demonstrate how difficult this balance can become.
12. Synthetic Data Can Also Increase Competition
Synthetic data should not automatically be treated as a concentration problem.
It can actually reduce barriers to entry.
A startup that cannot obtain millions of real customer records might generate synthetic information capable of supporting model development while reducing dependence on proprietary real-world datasets.
Synthetic data can therefore potentially:
- reduce data-acquisition costs;
- facilitate experimentation;
- reduce some privacy constraints;
- enable smaller AI developers;
- create alternatives to incumbent datasets;
- improve market entry.
This makes the competitive effect highly context-dependent.
13. Competition-Law Assessment Framework
A competition authority examining synthetic-data concentration would generally need to consider:
Market definition → Market shares and concentration → Entry barriers → Control of source data → Access to compute → Intellectual-property protection → Network effects → Vertical integration → Switching costs → Interoperability → Exclusivity arrangements → Merger effects → Innovation effects → Efficiencies → Consumer harm.
The key question is not simply whether a company possesses enormous quantities of synthetic data.
It is whether the company's position and conduct substantially reduce effective competitive constraints.
14. Current Regulatory Direction
Competition authorities are increasingly examining AI as a complete value chain rather than considering models in isolation.
A 2024 joint statement from competition authorities in the United Kingdom, European Union and United States highlighted three particularly relevant risks: concentrated control over critical AI inputs, extension or entrenchment of existing market power, and arrangements among important market participants that could amplify those risks.
These principles are directly relevant to future synthetic-data markets because synthetic data can itself become an AI input and can depend upon other concentrated inputs such as source data, compute and foundation models.
Conclusion
Competition Law and Synthetic Data Market Concentration concerns the possibility that control over synthetic datasets, source data, AI models, compute infrastructure and distribution channels becomes concentrated among a limited number of companies.
The strongest competition concerns would arise where concentration produces durable barriers to entry or enables exclusionary conduct—for example, data foreclosure, restrictive interoperability, tying, discriminatory access or anticompetitive acquisitions. At the same time, synthetic data can be strongly pro-competitive by giving smaller companies alternatives to scarce proprietary real-world information.
The most relevant established precedents include Google/DoubleClick, Facebook/WhatsApp, Microsoft/LinkedIn, Google Shopping, Google Android, Microsoft, IMS Health, and Slovak Telekom. None is a direct synthetic-data antitrust judgment; instead, together they provide legal principles concerning data accumulation, ecosystem power, interoperability, essential inputs, vertical foreclosure and digital-market concentration that are likely to shape future synthetic-data competition cases.

comments