Inference Infrastructure Concentration And Latency-Based Dominance .

Inference Infrastructure Concentration And Latency-Based Dominance

Introduction

Inference infrastructure concentration refers to a situation in which a small number of firms control the computing, networking, accelerator, cloud, model-serving, or API infrastructure required to run AI models at scale. Latency-based dominance arises where a firm's competitive advantage depends not merely on the quality or price of its AI service, but on its ability to deliver model outputs faster, more reliably, and at greater geographic proximity to users.

This creates a distinctive competition-law problem. A firm may become dominant because competitors cannot reproduce the speed, capacity, reliability, and geographical distribution of its inference infrastructure, even where the underlying AI model itself is technically replicable.

The issue can therefore be understood as:

Infrastructure concentration → scarce inference capacity → lower latency → better user experience → increased demand → greater scale → further infrastructure advantage → higher entry barriers.

Traditional competition law can address these concerns through market definition, dominance/monopolisation, refusal to supply, essential facilities, tying, exclusive dealing, discriminatory access, margin squeeze, self-preferencing, and merger control.

1. Meaning of Inference Infrastructure

AI inference is the process through which a trained model produces an output in response to an input.

The infrastructure supporting inference can include:

  • GPUs and AI accelerators;
  • TPUs and specialised inference chips;
  • cloud computing;
  • model-serving infrastructure;
  • inference APIs;
  • data centres;
  • high-speed networking;
  • memory and storage;
  • distributed computing systems;
  • edge-computing facilities;
  • content-delivery networks;
  • model-routing systems;
  • orchestration software;
  • inference optimisation tools; and
  • geographic deployment capacity.

A modern AI provider may therefore depend on a vertically integrated stack:

Model → accelerator → cloud → networking → inference server → API → application.

Competition concerns arise when one firm controls several of these layers simultaneously.

2. What Is Latency-Based Dominance?

Latency is the time between a user's request and the delivery of an AI response.

In conventional markets, small differences in delivery time may have limited competitive significance. In AI markets, latency can become a central competitive variable.

For example:

  • a financial-trading AI system may require millisecond-level responses;
  • an autonomous vehicle system requires extremely rapid inference;
  • a voice assistant must respond almost immediately;
  • an industrial-control system may require real-time processing;
  • a gaming application may be highly sensitive to delay;
  • an AI search engine may lose users if responses are substantially slower.

Consequently, the relevant competitive parameter may become:

Price + accuracy + reliability + latency + throughput.

A provider with substantially lower latency may therefore gain market power even where competitors offer comparable model quality.

3. How Infrastructure Concentration Creates Market Power

A. Scarcity of Accelerators

Advanced AI accelerators are expensive and may have limited availability.

A firm with privileged access to large accelerator clusters can offer:

  • greater inference capacity;
  • faster response times;
  • lower congestion;
  • higher throughput;
  • better reliability.

Competitors unable to obtain comparable capacity may be unable to match these characteristics.

B. Economies of Scale

Inference infrastructure often exhibits substantial economies of scale.

Large providers can distribute:

  • hardware costs;
  • energy costs;
  • networking costs;
  • software optimisation;
  • engineering expenditure; and
  • data-centre costs

over enormous volumes of inference requests.

This can produce a reinforcing cycle:

More users → greater utilisation → lower average cost → better service → more users.

4. Latency as a Competition Parameter

Latency can operate as a quality parameter rather than merely a technical specification.

Suppose two AI providers have:

FactorProvider AProvider B
AccuracySimilarSimilar
PriceSimilarSimilar
ReliabilityHighHigh
Latency50 ms500 ms

For many applications, Provider A may have a substantial competitive advantage.

The competition authority should therefore avoid defining the market solely by:

"AI models"

and instead examine whether competition occurs in a narrower market for:

high-performance, low-latency inference services.

5. Geographic Concentration

Inference infrastructure is also geographically significant.

A provider with data centres close to major population or commercial centres may offer substantially lower latency.

This may create geographic competitive advantages through:

  • proximity to users;
  • local data centres;
  • edge nodes;
  • private networking;
  • regional cloud zones;
  • dedicated capacity.

A competitor with equivalent computing power located much farther away may nevertheless be unable to replicate the same service quality.

Thus, infrastructure concentration can generate location-based market power.

6. Network Effects and Feedback Loops

Latency advantages can create network effects.

A simplified model is:

Lower latency

↓

Better user experience

↓

More users

↓

More inference traffic

↓

More revenue

↓

More infrastructure investment

↓

Greater geographic coverage

↓

Even lower latency

This can produce a self-reinforcing competitive advantage.

The important antitrust question is whether this advantage represents competition on the merits or whether exclusionary conduct prevents rivals from obtaining comparable infrastructure.

7. Vertical Integration

Inference infrastructure concentration becomes particularly problematic when a company operates across multiple levels.

For example:

AI accelerator

↓

Cloud infrastructure

↓

Model

↓

Inference API

↓

Consumer application

A vertically integrated firm could potentially disadvantage independent AI developers by:

  • reserving scarce computing capacity for its own models;
  • charging rivals higher infrastructure prices;
  • restricting access to APIs;
  • degrading interoperability;
  • bundling cloud and inference services;
  • offering preferential latency to affiliated products.

This creates potential vertical foreclosure.

8. Relevant Competition-Law Theories

A. Abuse of Dominance

A dominant infrastructure provider may infringe competition law if it uses its position to exclude competitors rather than merely competing through superior technology.

Potential conduct includes:

  • discriminatory access;
  • refusal to supply;
  • discriminatory pricing;
  • exclusive agreements;
  • tying;
  • bundling;
  • self-preferencing;
  • capacity reservation;
  • degradation of rivals' performance.

B. Essential-Facilities Theory

Inference infrastructure may become comparable to an essential facility where:

  1. the infrastructure is indispensable;
  2. competitors cannot reasonably duplicate it;
  3. access is technically feasible;
  4. refusal substantially eliminates competition; and
  5. access can be provided without disproportionate difficulty.

Not every important infrastructure asset qualifies as an essential facility.

The doctrine generally requires necessity, not simply commercial usefulness.

9. Latency and Refusal to Deal

Suppose a dominant cloud provider operates the only economically viable low-latency inference infrastructure in a particular region.

It refuses access to a competing AI developer while continuing to provide the same infrastructure to its own downstream AI service.

The competition authority might investigate:

  • whether the infrastructure is indispensable;
  • whether alternatives exist;
  • whether access is technically possible;
  • whether refusal eliminates effective competition;
  • whether the provider has legitimate business justification.

This closely resembles traditional refusal-to-deal analysis but involves computational infrastructure rather than physical infrastructure.

10. Margin Squeeze

A vertically integrated company could potentially engage in a margin squeeze by:

  • charging competitors a high price for cloud/inference infrastructure; while
  • pricing its own downstream AI service aggressively.

The competitor may technically have access but be unable to compete profitably.

The issue therefore becomes:

Can an equally efficient downstream competitor survive using the infrastructure supplied by the dominant provider?

This is particularly important where the infrastructure provider controls the input necessary for achieving competitive latency.

11. Tying and Bundling

A dominant cloud provider could require customers purchasing inference capacity to purchase:

  • its own model;
  • its API;
  • its monitoring system;
  • its storage;
  • its networking service; or
  • another proprietary AI product.

This could raise tying concerns if the conduct forecloses independent providers.

The competition authority would examine:

  • dominance in the tying product;
  • separability of products;
  • coercion;
  • foreclosure;
  • consumer harm;
  • efficiencies.

12. Exclusive Capacity Agreements

An infrastructure provider might enter long-term agreements reserving scarce accelerator or data-centre capacity exclusively for selected AI firms.

Such agreements can be commercially legitimate because infrastructure investment requires predictable demand.

But competition concerns may arise if:

the agreements deprive rivals of the minimum infrastructure necessary to compete.

The duration, coverage, exclusivity, switching costs and availability of alternative infrastructure would be relevant.

13. Self-Preferencing

A vertically integrated AI infrastructure provider may give its own downstream AI applications:

  • lower latency;
  • priority inference queues;
  • preferential access to GPUs;
  • higher throughput;
  • better network routing;
  • guaranteed capacity.

Competitors may technically receive access but receive inferior performance.

This is particularly important because discrimination can be hidden in technical parameters rather than explicit prices.

14. Latency Discrimination

Latency discrimination deserves separate consideration.

Imagine an infrastructure provider gives:

  • affiliated AI service: 40 ms average latency;
  • independent competitor: 250 ms average latency.

Even if both pay similar prices, the infrastructure provider may have altered the competitive conditions through technical discrimination.

A competition investigation could therefore examine:

  • routing;
  • queue priority;
  • GPU allocation;
  • network peering;
  • geographic deployment;
  • API throttling;
  • congestion management;
  • workload scheduling.

The relevant evidence may exist in technical logs rather than conventional commercial documents.

15. Merger-Control Concerns

Infrastructure concentration can also arise through acquisitions.

Potential transactions include:

  • cloud provider acquiring an AI accelerator company;
  • AI model provider acquiring a specialised inference platform;
  • cloud provider acquiring an inference API;
  • data-centre operator acquiring an AI infrastructure provider;
  • chip company acquiring model-serving software.

Authorities may examine whether the merger creates:

Input foreclosure

Rivals lose access to essential inference infrastructure.

Customer foreclosure

Infrastructure suppliers lose independent AI customers.

Data foreclosure

The merged firm combines infrastructure and proprietary data advantages.

Innovation foreclosure

Competing architectures become commercially unviable.

16. Relevant Case Laws

The following cases do not all concern AI inference specifically. They establish competition-law principles that can be applied to AI infrastructure concentration, access, latency discrimination and vertically integrated digital ecosystems.

1. United Brands v Commission

The Court of Justice recognised that dominance involves a position of economic strength enabling a firm to behave to an appreciable extent independently of competitors, customers and consumers.

Relevance

An inference infrastructure provider controlling scarce computational capacity could acquire dominance where competitors cannot constrain its conduct.

Latency advantages may strengthen that position where customers cannot practically switch to alternatives providing equivalent performance.

Principle: Market power depends on the ability to behave independently, not merely on market share.

2. Commercial Solvents Corp v Commission

This case established important principles concerning refusal to supply by a dominant undertaking controlling an input necessary for downstream competition.

Relevance to AI inference

An infrastructure provider supplying computational resources to independent AI developers could potentially face similar concerns if it:

  • controls an indispensable input;
  • competes downstream; and
  • refuses supply to downstream competitors.

The analogy becomes particularly strong where alternative inference capacity is unavailable.

Principle: A dominant firm controlling an indispensable input cannot necessarily use that control to eliminate downstream competition.

3. Bronner v Mediaprint

The Court imposed a demanding standard for compulsory access to infrastructure.

The facility must generally be indispensable, and there must be no actual or potential substitute capable of providing viable competition.

Relevance

This is crucial for inference infrastructure.

A competition authority should not automatically treat every expensive GPU cluster or cloud platform as an essential facility.

It should ask:

  • Can rivals obtain equivalent accelerators?
  • Can they build their own data centres?
  • Can they use another cloud provider?
  • Can they tolerate higher latency?
  • Can edge computing provide an alternative?
  • Is duplication economically feasible?

Principle: Importance alone does not make infrastructure an essential facility.

4. IMS Health v Commission

The Court dealt with access to an indispensable information infrastructure protected by intellectual-property rights.

The case developed stringent conditions for compulsory licensing/access.

Relevance

AI infrastructure may combine:

  • proprietary software;
  • model-serving technology;
  • data;
  • specialised hardware;
  • network architecture.

If competitors require access to such infrastructure, IMS Health provides a framework for determining when competition law can require access despite proprietary rights.

Principle: Competition law may intervene where control over indispensable infrastructure prevents effective downstream competition, but the threshold is high.

5. Microsoft Corp v Commission

The EU Microsoft decision is highly relevant to digital infrastructure.

Microsoft was found to have abused dominance through conduct involving interoperability information and tying.

Relevance to AI ecosystems

Modern AI infrastructure similarly involves interoperability between:

  • models;
  • APIs;
  • operating environments;
  • cloud platforms;
  • developer tools;
  • hardware;
  • applications.

A dominant infrastructure provider that deliberately restricts interoperability could potentially prevent competing AI systems from achieving comparable functionality or latency.

Principle: Control over technological interfaces can become a source of exclusionary market power.

6. Google Shopping

The Google Shopping decision illustrates how a dominant digital platform can use control over an important ecosystem to favour its own downstream service.

Relevance

Inference infrastructure creates a similar possibility.

A vertically integrated infrastructure provider could potentially favour its own AI application through:

  • preferential inference capacity;
  • superior routing;
  • faster response times;
  • better geographic deployment;
  • privileged access to accelerator clusters.

The important question is whether the conduct distorts competition beyond competition on the merits.

Principle: Digital self-preferencing can be particularly significant where infrastructure control determines downstream competitive conditions.

7. Slovak Telekom v Commission

The Court examined exclusionary conduct involving a dominant vertically integrated telecommunications operator and access conditions for competitors.

Relevance

The case provides a useful analogy for AI infrastructure because telecommunications networks and inference infrastructure both involve:

  • substantial fixed investment;
  • network effects;
  • capacity constraints;
  • access conditions;
  • downstream services.

A dominant AI infrastructure provider could potentially use pricing or access conditions to make downstream competition commercially impossible.

Principle: Vertical control over infrastructure can facilitate exclusion of downstream competitors.

8. Bronner + Commercial Solvents Compared

The combination of these cases is especially useful for AI.

Commercial Solvents demonstrates that control over an input can create exclusionary concerns.

Bronner demonstrates that compulsory access requires a much stronger showing of indispensability.

Thus:

"Competitors need my infrastructure" ≠ automatically "I must provide access."

The competition authority must establish why the infrastructure is indispensable and why alternative sources cannot realistically discipline the dominant provider.

17. Latency as a Relevant Market Dimension

Competition authorities may traditionally examine:

  • price;
  • quality;
  • output;
  • market shares.

For AI inference, they may also need to examine:

  • time-to-first-token;
  • tokens per second;
  • round-trip latency;
  • geographic proximity;
  • inference throughput;
  • uptime;
  • capacity guarantees;
  • peak-load performance.

This produces a more sophisticated concept of quality-adjusted market power.

18. The "Latency Gap" Test

A useful analytical framework could compare:

Latency of dominant provider ÷ latency of effective competitor

For example:

  • Provider A: 50 ms
  • Provider B: 200 ms

The fourfold latency difference may matter greatly in real-time applications.

But latency must be considered together with:

  • price;
  • accuracy;
  • reliability;
  • functionality;
  • switching costs;
  • application requirements.

A latency difference that matters enormously to autonomous vehicles may matter little for a document-generation service.

19. Relevant Counterfactual

The competition authority should ask:

What would competition look like if rivals had access to equivalent inference infrastructure?

If competitors immediately become viable once equivalent infrastructure is available, infrastructure access may be a significant competitive bottleneck.

Conversely, if the incumbent retains its advantage because of superior algorithms, data, model quality or innovation, infrastructure concentration may not itself constitute an antitrust problem.

20. Consumer and Business Harm

Infrastructure concentration can cause:

Higher prices

Cloud and inference prices may remain above competitive levels.

Lower innovation

Start-ups may be unable to experiment at sufficient scale.

Reduced choice

Developers become dependent on one inference provider.

Lower quality

Competitors may be unable to offer comparable response times.

Reduced resilience

System-wide dependence on one infrastructure provider creates systemic outage risks.

Switching costs

Applications may become deeply integrated with proprietary APIs and infrastructure.

21. Entry Barriers

Inference markets may exhibit unusually high barriers to entry because new entrants need:

  • capital;
  • accelerator supply;
  • electricity;
  • data centres;
  • networking;
  • software expertise;
  • specialised engineers;
  • model optimisation;
  • customer relationships.

Consequently, an incumbent's infrastructure advantage may persist even where the underlying AI technology is theoretically replicable.

22. Infrastructure as a Bottleneck Asset

The most important conceptual development is to recognise inference infrastructure as a potential bottleneck asset.

A bottleneck asset is infrastructure that competitors cannot readily bypass.

Possible bottlenecks include:

GPU clusters → cloud capacity → high-speed networking → inference serving → low-latency API.

Control of one or more bottlenecks can produce substantial downstream market power.

23. Competition Between Infrastructure Architectures

Authorities should not assume that concentration in one infrastructure architecture necessarily means market power.

Potential substitutes could include:

  • CPUs;
  • GPUs;
  • TPUs;
  • ASICs;
  • edge inference;
  • local inference;
  • specialised accelerators;
  • distributed computing;
  • smaller models;
  • model compression;
  • quantisation.

The crucial question is economic substitutability, not merely technical substitutability.

24. Remedies

Possible competition-law remedies include:

Access remedies

Require non-discriminatory access to scarce infrastructure.

Interoperability

Require APIs and interfaces to remain interoperable.

Non-discrimination

Prevent preferential latency for affiliated products.

Capacity commitments

Require allocation of sufficient capacity to independent customers.

Transparency

Require disclosure of material technical access conditions.

Structural remedies

In exceptional cases, separation of infrastructure and downstream AI operations may be considered.

Merger remedies

Authorities could impose:

  • access commitments;
  • licensing;
  • interoperability obligations;
  • capacity guarantees;
  • restrictions on exclusive arrangements.

25. Key Legal Test

A competition authority examining inference infrastructure concentration could proceed through the following sequence:

1. Define the relevant market

↓

2. Identify inference infrastructure suppliers

↓

3. Measure capacity and geographic coverage

↓

4. Examine latency and throughput

↓

5. Assess switching possibilities

↓

6. Determine dominance

↓

7. Identify exclusionary conduct

↓

8. Examine legitimate efficiencies

↓

9. Assess actual/potential foreclosure

↓

10. Design proportionate remedies

26. Important Distinction: Superior Performance vs Exclusion

Competition law should not punish a firm merely because it has faster inference infrastructure.

If the firm achieves low latency through:

  • superior engineering;
  • better hardware;
  • efficient data centres;
  • innovative scheduling;
  • legitimate investment;
  • better algorithms;

then the resulting advantage is ordinarily a form of competition on the merits.

The legal problem arises when the dominant firm uses infrastructure control to artificially prevent rivals from obtaining comparable competitive opportunities.

This distinction is fundamental.

27. Emerging AI-Specific Concern

The most significant future problem may be the combination of:

compute concentration + model concentration + cloud concentration + latency advantage.

A company controlling all four can potentially create a vertically integrated AI ecosystem in which competitors technically have access to AI models but cannot achieve commercially viable performance.

This produces a new form of market power:

performance-based infrastructural dominance.

The competitor is not necessarily denied access altogether; instead, it receives access on conditions that make its product too slow, too expensive, or too unreliable to compete effectively.

Conclusion

Inference infrastructure concentration and latency-based dominance represent an important emerging competition-law problem in AI markets.

The central issue is not simply who owns the largest number of GPUs or data centres. It is whether control over scarce computational and networking infrastructure enables a firm to:

  • achieve persistent latency advantages;
  • control downstream AI access;
  • discriminate against competitors;
  • foreclose rival models or applications;
  • impose excessive infrastructure costs;
  • restrict interoperability;
  • reserve scarce capacity; or
  • leverage infrastructure dominance into downstream AI markets.

The traditional cases of Commercial Solvents, Bronner, IMS Health, Microsoft, Google Shopping, Slovak Telekom and United Brands provide the principal doctrinal building blocks.

The key legal distinction is:

A low-latency advantage obtained through superior innovation is legitimate competition; a low-latency advantage maintained by exclusionary control over indispensable infrastructure may constitute an abuse of dominance.

Thus, future competition enforcement may increasingly have to treat latency, inference capacity, geographic deployment and infrastructure access as competitive variables, rather than treating AI infrastructure merely as a technical backend.

 

 

LEAVE A COMMENT