Inference Infrastructure Concentration And Latency-Based Dominance .
Inference Infrastructure Concentration And Latency-Based Dominance
Introduction
Inference infrastructure concentration refers to a situation in which a small number of firms control the computing, networking, accelerator, cloud, model-serving, or API infrastructure required to run AI models at scale. Latency-based dominance arises where a firm's competitive advantage depends not merely on the quality or price of its AI service, but on its ability to deliver model outputs faster, more reliably, and at greater geographic proximity to users.
This creates a distinctive competition-law problem. A firm may become dominant because competitors cannot reproduce the speed, capacity, reliability, and geographical distribution of its inference infrastructure, even where the underlying AI model itself is technically replicable.
The issue can therefore be understood as:
Infrastructure concentration → scarce inference capacity → lower latency → better user experience → increased demand → greater scale → further infrastructure advantage → higher entry barriers.
Traditional competition law can address these concerns through market definition, dominance/monopolisation, refusal to supply, essential facilities, tying, exclusive dealing, discriminatory access, margin squeeze, self-preferencing, and merger control.
1. Meaning of Inference Infrastructure
AI inference is the process through which a trained model produces an output in response to an input.
The infrastructure supporting inference can include:
- GPUs and AI accelerators;
- TPUs and specialised inference chips;
- cloud computing;
- model-serving infrastructure;
- inference APIs;
- data centres;
- high-speed networking;
- memory and storage;
- distributed computing systems;
- edge-computing facilities;
- content-delivery networks;
- model-routing systems;
- orchestration software;
- inference optimisation tools; and
- geographic deployment capacity.
A modern AI provider may therefore depend on a vertically integrated stack:
Model → accelerator → cloud → networking → inference server → API → application.
Competition concerns arise when one firm controls several of these layers simultaneously.
2. What Is Latency-Based Dominance?
Latency is the time between a user's request and the delivery of an AI response.
In conventional markets, small differences in delivery time may have limited competitive significance. In AI markets, latency can become a central competitive variable.
For example:
- a financial-trading AI system may require millisecond-level responses;
- an autonomous vehicle system requires extremely rapid inference;
- a voice assistant must respond almost immediately;
- an industrial-control system may require real-time processing;
- a gaming application may be highly sensitive to delay;
- an AI search engine may lose users if responses are substantially slower.
Consequently, the relevant competitive parameter may become:
Price + accuracy + reliability + latency + throughput.
A provider with substantially lower latency may therefore gain market power even where competitors offer comparable model quality.
3. How Infrastructure Concentration Creates Market Power
A. Scarcity of Accelerators
Advanced AI accelerators are expensive and may have limited availability.
A firm with privileged access to large accelerator clusters can offer:
- greater inference capacity;
- faster response times;
- lower congestion;
- higher throughput;
- better reliability.
Competitors unable to obtain comparable capacity may be unable to match these characteristics.
B. Economies of Scale
Inference infrastructure often exhibits substantial economies of scale.
Large providers can distribute:
- hardware costs;
- energy costs;
- networking costs;
- software optimisation;
- engineering expenditure; and
- data-centre costs
over enormous volumes of inference requests.
This can produce a reinforcing cycle:
More users → greater utilisation → lower average cost → better service → more users.
4. Latency as a Competition Parameter
Latency can operate as a quality parameter rather than merely a technical specification.
Suppose two AI providers have:
| Factor | Provider A | Provider B |
|---|---|---|
| Accuracy | Similar | Similar |
| Price | Similar | Similar |
| Reliability | High | High |
| Latency | 50 ms | 500 ms |
For many applications, Provider A may have a substantial competitive advantage.
The competition authority should therefore avoid defining the market solely by:
"AI models"
and instead examine whether competition occurs in a narrower market for:
high-performance, low-latency inference services.
5. Geographic Concentration
Inference infrastructure is also geographically significant.
A provider with data centres close to major population or commercial centres may offer substantially lower latency.
This may create geographic competitive advantages through:
- proximity to users;
- local data centres;
- edge nodes;
- private networking;
- regional cloud zones;
- dedicated capacity.
A competitor with equivalent computing power located much farther away may nevertheless be unable to replicate the same service quality.
Thus, infrastructure concentration can generate location-based market power.
6. Network Effects and Feedback Loops
Latency advantages can create network effects.
A simplified model is:
Lower latency
↓
Better user experience
↓
More users
↓
More inference traffic
↓
More revenue
↓
More infrastructure investment
↓
Greater geographic coverage
↓
Even lower latency
This can produce a self-reinforcing competitive advantage.
The important antitrust question is whether this advantage represents competition on the merits or whether exclusionary conduct prevents rivals from obtaining comparable infrastructure.
7. Vertical Integration
Inference infrastructure concentration becomes particularly problematic when a company operates across multiple levels.
For example:
AI accelerator
↓
Cloud infrastructure
↓
Model
↓
Inference API
↓
Consumer application
A vertically integrated firm could potentially disadvantage independent AI developers by:
- reserving scarce computing capacity for its own models;
- charging rivals higher infrastructure prices;
- restricting access to APIs;
- degrading interoperability;
- bundling cloud and inference services;
- offering preferential latency to affiliated products.
This creates potential vertical foreclosure.
8. Relevant Competition-Law Theories
A. Abuse of Dominance
A dominant infrastructure provider may infringe competition law if it uses its position to exclude competitors rather than merely competing through superior technology.
Potential conduct includes:
- discriminatory access;
- refusal to supply;
- discriminatory pricing;
- exclusive agreements;
- tying;
- bundling;
- self-preferencing;
- capacity reservation;
- degradation of rivals' performance.
B. Essential-Facilities Theory
Inference infrastructure may become comparable to an essential facility where:
- the infrastructure is indispensable;
- competitors cannot reasonably duplicate it;
- access is technically feasible;
- refusal substantially eliminates competition; and
- access can be provided without disproportionate difficulty.
Not every important infrastructure asset qualifies as an essential facility.
The doctrine generally requires necessity, not simply commercial usefulness.
9. Latency and Refusal to Deal
Suppose a dominant cloud provider operates the only economically viable low-latency inference infrastructure in a particular region.
It refuses access to a competing AI developer while continuing to provide the same infrastructure to its own downstream AI service.
The competition authority might investigate:
- whether the infrastructure is indispensable;
- whether alternatives exist;
- whether access is technically possible;
- whether refusal eliminates effective competition;
- whether the provider has legitimate business justification.
This closely resembles traditional refusal-to-deal analysis but involves computational infrastructure rather than physical infrastructure.
10. Margin Squeeze
A vertically integrated company could potentially engage in a margin squeeze by:
- charging competitors a high price for cloud/inference infrastructure; while
- pricing its own downstream AI service aggressively.
The competitor may technically have access but be unable to compete profitably.
The issue therefore becomes:
Can an equally efficient downstream competitor survive using the infrastructure supplied by the dominant provider?
This is particularly important where the infrastructure provider controls the input necessary for achieving competitive latency.
11. Tying and Bundling
A dominant cloud provider could require customers purchasing inference capacity to purchase:
- its own model;
- its API;
- its monitoring system;
- its storage;
- its networking service; or
- another proprietary AI product.
This could raise tying concerns if the conduct forecloses independent providers.
The competition authority would examine:
- dominance in the tying product;
- separability of products;
- coercion;
- foreclosure;
- consumer harm;
- efficiencies.
12. Exclusive Capacity Agreements
An infrastructure provider might enter long-term agreements reserving scarce accelerator or data-centre capacity exclusively for selected AI firms.
Such agreements can be commercially legitimate because infrastructure investment requires predictable demand.
But competition concerns may arise if:
the agreements deprive rivals of the minimum infrastructure necessary to compete.
The duration, coverage, exclusivity, switching costs and availability of alternative infrastructure would be relevant.
13. Self-Preferencing
A vertically integrated AI infrastructure provider may give its own downstream AI applications:
- lower latency;
- priority inference queues;
- preferential access to GPUs;
- higher throughput;
- better network routing;
- guaranteed capacity.
Competitors may technically receive access but receive inferior performance.
This is particularly important because discrimination can be hidden in technical parameters rather than explicit prices.
14. Latency Discrimination
Latency discrimination deserves separate consideration.
Imagine an infrastructure provider gives:
- affiliated AI service: 40 ms average latency;
- independent competitor: 250 ms average latency.
Even if both pay similar prices, the infrastructure provider may have altered the competitive conditions through technical discrimination.
A competition investigation could therefore examine:
- routing;
- queue priority;
- GPU allocation;
- network peering;
- geographic deployment;
- API throttling;
- congestion management;
- workload scheduling.
The relevant evidence may exist in technical logs rather than conventional commercial documents.
15. Merger-Control Concerns
Infrastructure concentration can also arise through acquisitions.
Potential transactions include:
- cloud provider acquiring an AI accelerator company;
- AI model provider acquiring a specialised inference platform;
- cloud provider acquiring an inference API;
- data-centre operator acquiring an AI infrastructure provider;
- chip company acquiring model-serving software.
Authorities may examine whether the merger creates:
Input foreclosure
Rivals lose access to essential inference infrastructure.
Customer foreclosure
Infrastructure suppliers lose independent AI customers.
Data foreclosure
The merged firm combines infrastructure and proprietary data advantages.
Innovation foreclosure
Competing architectures become commercially unviable.
16. Relevant Case Laws
The following cases do not all concern AI inference specifically. They establish competition-law principles that can be applied to AI infrastructure concentration, access, latency discrimination and vertically integrated digital ecosystems.
1. United Brands v Commission
The Court of Justice recognised that dominance involves a position of economic strength enabling a firm to behave to an appreciable extent independently of competitors, customers and consumers.
Relevance
An inference infrastructure provider controlling scarce computational capacity could acquire dominance where competitors cannot constrain its conduct.
Latency advantages may strengthen that position where customers cannot practically switch to alternatives providing equivalent performance.
Principle: Market power depends on the ability to behave independently, not merely on market share.
2. Commercial Solvents Corp v Commission
This case established important principles concerning refusal to supply by a dominant undertaking controlling an input necessary for downstream competition.
Relevance to AI inference
An infrastructure provider supplying computational resources to independent AI developers could potentially face similar concerns if it:
- controls an indispensable input;
- competes downstream; and
- refuses supply to downstream competitors.
The analogy becomes particularly strong where alternative inference capacity is unavailable.
Principle: A dominant firm controlling an indispensable input cannot necessarily use that control to eliminate downstream competition.
3. Bronner v Mediaprint
The Court imposed a demanding standard for compulsory access to infrastructure.
The facility must generally be indispensable, and there must be no actual or potential substitute capable of providing viable competition.
Relevance
This is crucial for inference infrastructure.
A competition authority should not automatically treat every expensive GPU cluster or cloud platform as an essential facility.
It should ask:
- Can rivals obtain equivalent accelerators?
- Can they build their own data centres?
- Can they use another cloud provider?
- Can they tolerate higher latency?
- Can edge computing provide an alternative?
- Is duplication economically feasible?
Principle: Importance alone does not make infrastructure an essential facility.
4. IMS Health v Commission
The Court dealt with access to an indispensable information infrastructure protected by intellectual-property rights.
The case developed stringent conditions for compulsory licensing/access.
Relevance
AI infrastructure may combine:
- proprietary software;
- model-serving technology;
- data;
- specialised hardware;
- network architecture.
If competitors require access to such infrastructure, IMS Health provides a framework for determining when competition law can require access despite proprietary rights.
Principle: Competition law may intervene where control over indispensable infrastructure prevents effective downstream competition, but the threshold is high.
5. Microsoft Corp v Commission
The EU Microsoft decision is highly relevant to digital infrastructure.
Microsoft was found to have abused dominance through conduct involving interoperability information and tying.
Relevance to AI ecosystems
Modern AI infrastructure similarly involves interoperability between:
- models;
- APIs;
- operating environments;
- cloud platforms;
- developer tools;
- hardware;
- applications.
A dominant infrastructure provider that deliberately restricts interoperability could potentially prevent competing AI systems from achieving comparable functionality or latency.
Principle: Control over technological interfaces can become a source of exclusionary market power.
6. Google Shopping
The Google Shopping decision illustrates how a dominant digital platform can use control over an important ecosystem to favour its own downstream service.
Relevance
Inference infrastructure creates a similar possibility.
A vertically integrated infrastructure provider could potentially favour its own AI application through:
- preferential inference capacity;
- superior routing;
- faster response times;
- better geographic deployment;
- privileged access to accelerator clusters.
The important question is whether the conduct distorts competition beyond competition on the merits.
Principle: Digital self-preferencing can be particularly significant where infrastructure control determines downstream competitive conditions.
7. Slovak Telekom v Commission
The Court examined exclusionary conduct involving a dominant vertically integrated telecommunications operator and access conditions for competitors.
Relevance
The case provides a useful analogy for AI infrastructure because telecommunications networks and inference infrastructure both involve:
- substantial fixed investment;
- network effects;
- capacity constraints;
- access conditions;
- downstream services.
A dominant AI infrastructure provider could potentially use pricing or access conditions to make downstream competition commercially impossible.
Principle: Vertical control over infrastructure can facilitate exclusion of downstream competitors.
8. Bronner + Commercial Solvents Compared
The combination of these cases is especially useful for AI.
Commercial Solvents demonstrates that control over an input can create exclusionary concerns.
Bronner demonstrates that compulsory access requires a much stronger showing of indispensability.
Thus:
"Competitors need my infrastructure" ≠ automatically "I must provide access."
The competition authority must establish why the infrastructure is indispensable and why alternative sources cannot realistically discipline the dominant provider.
17. Latency as a Relevant Market Dimension
Competition authorities may traditionally examine:
- price;
- quality;
- output;
- market shares.
For AI inference, they may also need to examine:
- time-to-first-token;
- tokens per second;
- round-trip latency;
- geographic proximity;
- inference throughput;
- uptime;
- capacity guarantees;
- peak-load performance.
This produces a more sophisticated concept of quality-adjusted market power.
18. The "Latency Gap" Test
A useful analytical framework could compare:
Latency of dominant provider ÷ latency of effective competitor
For example:
- Provider A: 50 ms
- Provider B: 200 ms
The fourfold latency difference may matter greatly in real-time applications.
But latency must be considered together with:
- price;
- accuracy;
- reliability;
- functionality;
- switching costs;
- application requirements.
A latency difference that matters enormously to autonomous vehicles may matter little for a document-generation service.
19. Relevant Counterfactual
The competition authority should ask:
What would competition look like if rivals had access to equivalent inference infrastructure?
If competitors immediately become viable once equivalent infrastructure is available, infrastructure access may be a significant competitive bottleneck.
Conversely, if the incumbent retains its advantage because of superior algorithms, data, model quality or innovation, infrastructure concentration may not itself constitute an antitrust problem.
20. Consumer and Business Harm
Infrastructure concentration can cause:
Higher prices
Cloud and inference prices may remain above competitive levels.
Lower innovation
Start-ups may be unable to experiment at sufficient scale.
Reduced choice
Developers become dependent on one inference provider.
Lower quality
Competitors may be unable to offer comparable response times.
Reduced resilience
System-wide dependence on one infrastructure provider creates systemic outage risks.
Switching costs
Applications may become deeply integrated with proprietary APIs and infrastructure.
21. Entry Barriers
Inference markets may exhibit unusually high barriers to entry because new entrants need:
- capital;
- accelerator supply;
- electricity;
- data centres;
- networking;
- software expertise;
- specialised engineers;
- model optimisation;
- customer relationships.
Consequently, an incumbent's infrastructure advantage may persist even where the underlying AI technology is theoretically replicable.
22. Infrastructure as a Bottleneck Asset
The most important conceptual development is to recognise inference infrastructure as a potential bottleneck asset.
A bottleneck asset is infrastructure that competitors cannot readily bypass.
Possible bottlenecks include:
GPU clusters → cloud capacity → high-speed networking → inference serving → low-latency API.
Control of one or more bottlenecks can produce substantial downstream market power.
23. Competition Between Infrastructure Architectures
Authorities should not assume that concentration in one infrastructure architecture necessarily means market power.
Potential substitutes could include:
- CPUs;
- GPUs;
- TPUs;
- ASICs;
- edge inference;
- local inference;
- specialised accelerators;
- distributed computing;
- smaller models;
- model compression;
- quantisation.
The crucial question is economic substitutability, not merely technical substitutability.
24. Remedies
Possible competition-law remedies include:
Access remedies
Require non-discriminatory access to scarce infrastructure.
Interoperability
Require APIs and interfaces to remain interoperable.
Non-discrimination
Prevent preferential latency for affiliated products.
Capacity commitments
Require allocation of sufficient capacity to independent customers.
Transparency
Require disclosure of material technical access conditions.
Structural remedies
In exceptional cases, separation of infrastructure and downstream AI operations may be considered.
Merger remedies
Authorities could impose:
- access commitments;
- licensing;
- interoperability obligations;
- capacity guarantees;
- restrictions on exclusive arrangements.
25. Key Legal Test
A competition authority examining inference infrastructure concentration could proceed through the following sequence:
1. Define the relevant market
↓
2. Identify inference infrastructure suppliers
↓
3. Measure capacity and geographic coverage
↓
4. Examine latency and throughput
↓
5. Assess switching possibilities
↓
6. Determine dominance
↓
7. Identify exclusionary conduct
↓
8. Examine legitimate efficiencies
↓
9. Assess actual/potential foreclosure
↓
10. Design proportionate remedies
26. Important Distinction: Superior Performance vs Exclusion
Competition law should not punish a firm merely because it has faster inference infrastructure.
If the firm achieves low latency through:
- superior engineering;
- better hardware;
- efficient data centres;
- innovative scheduling;
- legitimate investment;
- better algorithms;
then the resulting advantage is ordinarily a form of competition on the merits.
The legal problem arises when the dominant firm uses infrastructure control to artificially prevent rivals from obtaining comparable competitive opportunities.
This distinction is fundamental.
27. Emerging AI-Specific Concern
The most significant future problem may be the combination of:
compute concentration + model concentration + cloud concentration + latency advantage.
A company controlling all four can potentially create a vertically integrated AI ecosystem in which competitors technically have access to AI models but cannot achieve commercially viable performance.
This produces a new form of market power:
performance-based infrastructural dominance.
The competitor is not necessarily denied access altogether; instead, it receives access on conditions that make its product too slow, too expensive, or too unreliable to compete effectively.
Conclusion
Inference infrastructure concentration and latency-based dominance represent an important emerging competition-law problem in AI markets.
The central issue is not simply who owns the largest number of GPUs or data centres. It is whether control over scarce computational and networking infrastructure enables a firm to:
- achieve persistent latency advantages;
- control downstream AI access;
- discriminate against competitors;
- foreclose rival models or applications;
- impose excessive infrastructure costs;
- restrict interoperability;
- reserve scarce capacity; or
- leverage infrastructure dominance into downstream AI markets.
The traditional cases of Commercial Solvents, Bronner, IMS Health, Microsoft, Google Shopping, Slovak Telekom and United Brands provide the principal doctrinal building blocks.
The key legal distinction is:
A low-latency advantage obtained through superior innovation is legitimate competition; a low-latency advantage maintained by exclusionary control over indispensable infrastructure may constitute an abuse of dominance.
Thus, future competition enforcement may increasingly have to treat latency, inference capacity, geographic deployment and infrastructure access as competitive variables, rather than treating AI infrastructure merely as a technical backend.

comments