Ai Infrastructure Stack Concentration (Chips, Cloud, Models, Apps) .

AI Inference Latency as a Competitive Parameter

1. Introduction

AI inference latency is the time between an AI system receiving a request and producing a usable response. In generative AI, it may include time-to-first-token (TTFT), inter-token latency, total response time, and end-to-end application latency.

Competition law traditionally focuses on price, output, quality, innovation and consumer choice. In AI markets, however, latency can itself be a dimension of competition. Two AI models may provide broadly comparable answers, yet the provider offering substantially faster inference can obtain an important competitive advantage.

Latency becomes particularly significant where AI is integrated into:

  • search engines;
  • coding assistants;
  • financial trading systems;
  • healthcare applications;
  • autonomous systems;
  • customer-service platforms;
  • advertising technology;
  • cloud APIs;
  • enterprise software;
  • real-time translation;
  • robotics and industrial systems.

A dominant AI or cloud provider could therefore potentially harm competition not merely by raising prices, but by degrading rivals' access to low-latency infrastructure, prioritising its own inference traffic, tying model access to proprietary hardware, imposing discriminatory API conditions, or using technical architecture to foreclose competing AI services.

2. Meaning of AI Inference Latency

The AI inference process can broadly be represented as:

User request → API/network → model processing → first token → subsequent tokens → completed response

Latency can consequently be divided into several competitive variables.

A. Time-to-First-Token

TTFT measures how quickly the system begins responding.

This is particularly important for:

  • conversational AI;
  • search;
  • customer-service chatbots;
  • coding assistants.

B. Inter-Token Latency

This measures the interval between successive generated tokens.

A system may therefore have:

  • fast initial response but slow generation; or
  • slower initial response but rapid subsequent generation.

C. End-to-End Latency

This measures the complete time required to produce the requested output.

D. Tail Latency

The P95/P99 latency experienced by users can be more commercially important than average latency.

For example:

Provider A: average response = 300 ms
Provider B: average response = 280 ms

The difference may appear insignificant.

But if Provider A has reliable P99 latency while Provider B experiences substantial delays during peak demand, enterprise customers may prefer Provider A.

3. Why Latency Is a Competition Variable

AI competition is increasingly multi-dimensional.

Instead of:

Price → Quality → Output

the competitive framework becomes:

Price + Accuracy + Reliability + Latency + Privacy + Interoperability + Innovation

Latency can therefore constitute a form of non-price competition.

A customer may select an AI provider because:

  1. it responds faster;
  2. its service remains fast during peak demand;
  3. its API provides predictable response times;
  4. its geographic infrastructure is closer to users;
  5. its accelerator architecture reduces inference time;
  6. its model is optimised for real-time applications.

4. Relationship Between Compute and Latency

Inference latency is closely connected to infrastructure.

Important inputs include:

  • GPUs;
  • AI accelerators;
  • TPUs;
  • memory bandwidth;
  • networking;
  • data-centre location;
  • model architecture;
  • quantisation;
  • batching;
  • caching;
  • inference software;
  • compiler optimisation;
  • model size.

This creates an important competition-law question:

Can control over an upstream AI-infrastructure input allow a dominant undertaking to control downstream AI competition through latency?

For example:

AI accelerator dominance

↓

Limited access to high-performance compute

↓

Higher inference latency for rivals

↓

Inferior user experience

↓

Customer migration

↓

Reduced downstream competition

5. Latency as a Quality Parameter

Competition law frequently recognises that quality can be competitively significant even where price does not change.

AI makes this particularly important because a slower AI product may effectively be a lower-quality product.

For example, a coding assistant that generates correct code in 1 second and another that generates comparable code in 8 seconds are not necessarily equivalent products.

Latency may influence:

  • developer productivity;
  • user retention;
  • advertising conversion;
  • search engagement;
  • customer-service completion rates;
  • trading execution;
  • medical workflow efficiency.

Consequently, an AI provider may compete aggressively through latency optimisation rather than price reduction.

6. Relevant Competition-Law Theories

A. Abuse of Dominance

Where an undertaking possesses substantial market power, deliberate manipulation of latency could potentially become relevant to an abuse-of-dominance analysis.

Examples include:

  • slowing access to competing AI models;
  • imposing artificial throttling;
  • discriminatory API rate limits;
  • denying access to high-performance inference infrastructure;
  • reserving scarce accelerator capacity for the dominant firm's own services.

The relevant question would not simply be:

"Is the competitor slower?"

It would be:

"Did the dominant undertaking engage in conduct that materially disadvantaged competitors and harmed competition?"

7. Refusal or Discrimination in Access to Compute

Suppose a cloud provider controls a scarce supply of advanced AI accelerators.

It provides:

  • its own AI models → premium accelerator access;
  • affiliated applications → priority scheduling;
  • independent AI developers → restricted capacity.

The resulting latency difference could become competitively significant.

This could raise issues concerning:

  • discriminatory access;
  • refusal to supply;
  • essential-facility-type theories;
  • self-preferencing;
  • vertical foreclosure;
  • exclusionary conduct.

However, mere superiority in infrastructure does not automatically constitute an antitrust violation. Legitimate technical, security, capacity and efficiency explanations would also have to be considered.

8. Self-Preferencing Through Latency

One particularly important AI scenario is latency-based self-preferencing.

Suppose an AI platform operates:

  1. its own foundation model;
  2. an AI marketplace;
  3. an inference API;
  4. competing third-party models.

The platform could theoretically give its own model:

  • priority GPU scheduling;
  • lower network latency;
  • preferential caching;
  • higher rate limits;
  • reserved inference capacity.

Third-party models might technically remain available but become commercially inferior because they respond more slowly.

This is potentially more difficult to detect than an outright exclusion.

9. AI APIs and Latency Discrimination

API markets introduce another issue.

A platform may offer:

Internal AI service:
50 ms latency

Third-party AI service:
250 ms latency

Even if both models have similar accuracy, developers may integrate the internal service because latency affects the performance of their own applications.

Potential competitive concerns include:

  • discriminatory routing;
  • unequal API infrastructure;
  • throttling;
  • differential rate limits;
  • preferential caching;
  • geographic routing discrimination;
  • differential access to inference accelerators.

10. Latency and Network Effects

Latency can reinforce network effects.

A simplified cycle is:

More users

↓

More inference demand

↓

More infrastructure investment

↓

Lower latency

↓

Better user experience

↓

More users

↓

More data and revenue

↓

Further infrastructure investment

This can create a latency-based competitive feedback loop.

A well-funded incumbent may therefore achieve advantages that smaller competitors cannot easily reproduce.

11. Latency and Economies of Scale

Inference infrastructure has substantial fixed and variable costs.

Large providers can potentially:

  • purchase accelerators in bulk;
  • build specialised data centres;
  • optimise inference software;
  • deploy geographically distributed infrastructure;
  • negotiate favourable network arrangements;
  • operate large inference clusters.

As scale increases, the provider may reduce latency and cost simultaneously.

This may create economies of scale that constitute a barrier to entry.

12. Latency and Switching Costs

Latency can also contribute to customer lock-in.

An enterprise may build its software around:

  • a particular inference API;
  • a particular model;
  • a particular accelerator architecture;
  • proprietary optimisation tools.

After substantial integration, switching to another provider may increase latency.

The customer therefore faces:

migration costs + engineering costs + latency degradation.

This can strengthen incumbent market power.

13. Latency and Vertical Integration

Vertical integration creates particularly interesting competition questions.

Consider:

AI accelerator manufacturer

↓

Cloud infrastructure

↓

Inference platform

↓

Foundation model

↓

AI application

A vertically integrated undertaking may control several layers simultaneously.

If it can optimise latency throughout the stack, independent competitors may have difficulty matching the performance.

The competition issue is therefore not simply:

"Does the company have the fastest model?"

It is:

"Does control over multiple complementary markets allow the undertaking to exclude equally efficient competitors?"

14. Six Important Case Laws

Because there are relatively few reported decisions specifically concerning AI inference latency, the most useful authorities are established competition cases involving quality, technical performance, access to infrastructure, interoperability, self-preferencing, exclusion and innovation. Their application to AI latency is therefore analogical rather than a claim that the courts decided AI-latency disputes.

Case 1: United Brands v Commission

Case: United Brands Company and United Brands Continentaal BV v Commission, Case 27/76, Court of Justice of the European Communities (1978).

Principle

The Court examined market power and the competitive conditions surrounding the relevant product.

Relevance to AI latency

The case demonstrates that competition analysis cannot be reduced mechanically to price.

For AI services, relevant competitive parameters can include:

  • performance;
  • responsiveness;
  • reliability;
  • technical characteristics;
  • customer preferences.

Therefore, when defining an AI inference market, latency may help establish whether customers regard two AI services as sufficiently substitutable.

AI application

If users systematically switch between models based upon response speed, latency can form part of the relevant product-market analysis.

Case 2: Microsoft Corp. v Commission

Case: Microsoft Corp. v Commission, Case T-201/04, General Court (2007).

Principle

The case involved Microsoft's control over important software technologies and interoperability issues.

The Court accepted the significance of access to interoperability information in enabling competing products to operate effectively.

AI latency relevance

AI ecosystems similarly depend upon interoperability among:

  • models;
  • APIs;
  • cloud infrastructure;
  • applications;
  • data;
  • accelerators.

If a dominant platform provides substantially better technical performance to its own downstream product while restricting competitors' technical access, latency could become part of a broader exclusionary strategy.

Key lesson

Technical interoperability can be a competition parameter, not merely a technical matter.

Case 3: Intel v Commission

Case: Intel Corporation v European Commission, Case C-413/14 P, Court of Justice (2017).

Principle

The case concerned exclusionary conduct involving rebates and the assessment of whether conduct was capable of producing anticompetitive foreclosure.

The Court emphasised the importance of analysing the actual competitive effects where appropriate.

AI latency relevance

Suppose an AI infrastructure provider provides favourable inference capacity to certain customers while making competing AI providers operate under materially worse latency conditions.

Competition analysis should examine:

  • market coverage;
  • duration;
  • actual or potential foreclosure;
  • competitive alternatives;
  • customer dependence;
  • efficiency explanations.

AI lesson

Latency advantages should not be treated as automatically anticompetitive. The mechanism and effects matter.

Case 4: Google Shopping

Case: Google and Alphabet v Commission, Case T-612/17, General Court (2021).

Principle

The case concerned Google's treatment of its own comparison-shopping service within its general search results.

The central competition issue included the possibility that a dominant platform could use its position in one market to advantage its own service in another.

AI latency relevance

The analogous AI scenario is:

Dominant AI platform

↓

Controls infrastructure/interface

↓

Own model receives preferential technical treatment

↓

Third-party models receive inferior latency

↓

Users disproportionately choose the integrated model

This illustrates how technical or algorithmic preference can become a competitive parameter.

Important distinction

Google Shopping did not concern AI inference latency. Its relevance is the broader principle concerning conduct by a dominant platform that advantages its own downstream service.

Case 5: Slovak Telekom

Case: European Commission v Slovak Telekom a.s. and Deutsche Telekom AG, Joined Cases C-152/19 P and C-165/19 P (2021).

Principle

The case involved exclusionary conduct concerning access to telecommunications infrastructure.

The Court examined how a dominant undertaking's control over an upstream network could affect downstream competition.

AI latency relevance

AI inference infrastructure increasingly resembles a layered network:

Compute → networking → cloud → inference → application

Where a dominant infrastructure provider controls an important upstream input, discriminatory access can potentially affect downstream AI competitors.

A latency differential may therefore be evidence of:

  • discriminatory access;
  • degradation;
  • foreclosure;
  • strategic prioritisation.

Case 6: Bronner v Mediaprint

Case: Oscar Bronner GmbH & Co. KG v Mediaprint Zeitungs und Zeitschriftenverlag GmbH & Co. KG, Case C-7/97, Court of Justice (1998).

Principle

The Court established a stringent framework for refusal-to-supply claims involving infrastructure.

A facility must generally be indispensable, and the refusal must be capable of eliminating effective competition, among other requirements.

AI latency relevance

An AI competitor might argue:

"Access to this particular inference infrastructure is necessary to compete at commercially viable latency."

That does not automatically establish an antitrust violation.

The Bronner framework highlights the importance of determining:

  • whether the infrastructure is genuinely indispensable;
  • whether alternatives exist;
  • whether replication is realistically possible;
  • whether refusal eliminates effective competition.

AI significance

This is especially relevant to:

  • scarce GPUs;
  • specialised AI accelerators;
  • hyperscale inference infrastructure;
  • proprietary inference networks.

Case 7: IMS Health

Case: IMS Health GmbH & Co. OHG v NDC Health GmbH & Co. KG, Joined Cases C-418/01 and C-7/01, Court of Justice.

Principle

The case developed the exceptional circumstances surrounding refusal to license intellectual-property-related infrastructure.

AI latency relevance

AI infrastructure increasingly involves proprietary:

  • inference software;
  • optimisation systems;
  • model-serving technology;
  • accelerator interfaces;
  • specialised datasets.

Where a dominant undertaking controls technology that competitors require to achieve commercially viable performance, questions may arise concerning whether refusal or discriminatory access crosses the competition-law threshold.

Again, technical indispensability must be established rather than assumed.

15. Comparative Case-Law Matrix

CaseCore principleAI latency relevance
United BrandsMarket power and product characteristicsLatency can be a product-quality parameter
MicrosoftInteroperability and technical accessAPI/infrastructure interoperability
IntelExclusionary conduct and foreclosure analysisLatency discrimination may require effects analysis
Google ShoppingPreferential treatment by dominant platformSelf-preferencing through technical performance
Slovak TelekomAccess to upstream infrastructureAI compute/inference infrastructure
BronnerRefusal to supply / indispensabilityAccess to low-latency infrastructure
IMS HealthExceptional access/licensing circumstancesProprietary AI technology and infrastructure

16. Latency-Based Foreclosure Theory

A particularly important theory can be expressed as follows:

Step 1 — Dominant infrastructure

A company controls scarce AI compute.

Step 2 — Differential allocation

Its own services receive priority access.

Step 3 — Latency divergence

Competitors experience:

  • higher TTFT;
  • greater P95 latency;
  • throttling;
  • lower throughput.

Step 4 — Customer migration

Developers choose the incumbent's model because applications need fast responses.

Step 5 — Network effects

More users generate more revenue and investment.

Step 6 — Entrenchment

Rivals lose scale and cannot economically reproduce the incumbent's infrastructure.

This is a classic vertical foreclosure hypothesis adapted to AI.

17. Latency as an Essential Competitive Parameter

Latency becomes particularly important where the application itself is time-sensitive.

High-latency sensitivity

  • algorithmic trading;
  • autonomous vehicles;
  • robotics;
  • gaming;
  • live translation;
  • fraud detection;
  • cybersecurity;
  • industrial control;
  • customer-service voice systems.

Lower-latency sensitivity

  • long-form document analysis;
  • batch processing;
  • offline summarisation;
  • archival research.

Consequently, the relevant market may need to be segmented according to use case and latency sensitivity.

18. Latency and Market Definition

Competition authorities may consider whether:

AI Model A with 100 ms latency

and

AI Model B with 1,000 ms latency

are actually substitutes for a real-time application.

The answer may differ by customer.

For example:

ApplicationImportance of latency
Real-time voice assistantVery high
SearchHigh
Coding assistantHigh
Customer supportHigh
Financial analyticsPotentially very high
Medical documentationModerate
Long-form researchLower
Batch data processingRelatively lower

Thus, latency can affect both product-market definition and competitive-effects analysis.

19. P95 and P99 Latency as Competition Metrics

Average latency may conceal important competitive differences.

For example:

ProviderAverageP95P99
A300 ms450 ms600 ms
B250 ms1,500 ms4,000 ms

Provider B may appear superior using average latency, but enterprise customers may prefer Provider A because its performance is more predictable.

Competition authorities should therefore potentially examine:

  • median latency;
  • P95;
  • P99;
  • peak-period performance;
  • geographic variation;
  • capacity throttling.

20. Latency Degradation as a Possible Exclusionary Strategy

A particularly subtle form of conduct would be strategic degradation.

Rather than banning competitors, a dominant platform could theoretically:

  • increase API processing delays;
  • reduce priority;
  • restrict accelerator availability;
  • impose lower throughput limits;
  • increase queue times;
  • deprioritise third-party inference workloads.

The service technically remains available.

But its competitive quality deteriorates.

This is analogous to competition concerns involving technical degradation in other network and platform markets.

21. Efficiency Defences

Latency advantages are not inherently unlawful.

A provider may legitimately have lower latency because of:

  • better hardware;
  • better algorithms;
  • superior model architecture;
  • efficient quantisation;
  • better caching;
  • geographically distributed data centres;
  • legitimate capacity planning;
  • investment in infrastructure.

Competition law should therefore distinguish:

Legitimate innovation

"Our model is faster because we developed superior inference technology."

from:

Potential exclusion

"We deliberately slow competitors using infrastructure they depend upon."

The former can represent vigorous competition; the latter may raise competition concerns depending upon market power, conduct and effects.

22. Consumer Welfare Effects

Latency can produce both positive and negative welfare effects.

Positive effects

Competition on latency can generate:

  • faster AI services;
  • better productivity;
  • lower waiting time;
  • improved accessibility;
  • more responsive applications;
  • technological innovation.

Potential negative effects from exclusion

If competition is reduced, consumers may ultimately face:

  • higher prices;
  • fewer AI providers;
  • slower innovation;
  • reduced choice;
  • greater dependence on one ecosystem;
  • weaker privacy or quality competition.

23. Evidence Required in an AI Latency Investigation

A competition authority examining latency should potentially collect:

Technical evidence

  • TTFT;
  • token-generation speed;
  • P95/P99 latency;
  • throughput;
  • GPU utilisation;
  • accelerator allocation;
  • API logs.

Commercial evidence

  • customer switching data;
  • contracts;
  • pricing;
  • customer complaints;
  • procurement documents;
  • internal strategy documents.

Infrastructure evidence

  • GPU availability;
  • cloud capacity;
  • networking arrangements;
  • geographical deployment;
  • data-centre architecture.

Algorithmic evidence

  • scheduling algorithms;
  • routing systems;
  • caching;
  • throttling;
  • queue prioritisation.

This allows the authority to determine whether latency differences arise from legitimate technical superiority or exclusionary conduct.

24. Possible Remedies

If competition harm were established, potential remedies could include:

A. Non-discrimination

Require equivalent infrastructure access for competing services under comparable conditions.

B. API transparency

Require disclosure of relevant:

  • rate limits;
  • throttling;
  • service-level conditions.

C. Interoperability

Allow competitors to connect to relevant infrastructure on reasonable terms.

D. Structural separation

In exceptional cases, competition authorities could consider separation between infrastructure and downstream AI services.

E. Monitoring

Require periodic reporting of:

  • latency;
  • throughput;
  • capacity allocation;
  • outage rates.

F. Behavioural commitments

Prohibit preferential scheduling or discriminatory technical treatment.

25. China Competition-Law Perspective

Under China's Anti-Monopoly Law, AI latency could become relevant primarily through established categories such as:

  • abuse of dominant market position;
  • refusal to deal;
  • discriminatory treatment;
  • tying;
  • unreasonable trading conditions;
  • exclusionary conduct;
  • concentration-related concerns.

The important analytical question would be whether latency manipulation constitutes a mechanism through which market power is exercised or competitors are foreclosed.

Chinese digital-platform enforcement also makes algorithmic conduct, platform rules, data and technological infrastructure increasingly relevant to competition analysis.

26. India Competition-Law Perspective

Under the Competition Act 2002, latency can potentially be considered within broader questions concerning:

  • relevant market;
  • market power;
  • denial of market access;
  • discriminatory conditions;
  • leveraging;
  • refusal to deal;
  • tying/bundling;
  • abuse of dominant position.

For example, if a dominant cloud platform provides its own AI service with materially better inference infrastructure while imposing discriminatory technical restrictions on competing AI providers, latency could form part of the evidence concerning denial of market access or leveraging.

27. United States Perspective

Under U.S. antitrust principles, latency can potentially matter in:

  • monopolisation analysis;
  • exclusionary conduct;
  • vertical foreclosure;
  • essential-input disputes;
  • platform competition;
  • merger review.

The key distinction remains between:

competition on the merits

and

artificial conduct designed to exclude rivals.

A superior AI architecture that legitimately reduces latency is generally evidence of innovation rather than exclusion.

28. Merger-Control Implications

AI mergers may raise latency concerns even where traditional market-share analysis appears modest.

Consider:

Cloud Provider + Foundation Model Developer

The transaction may combine:

  • compute;
  • cloud;
  • model;
  • API;
  • application distribution.

The combined undertaking could potentially improve its own model's latency while disadvantaging rival models.

Authorities may therefore examine:

  • access to compute;
  • accelerator capacity;
  • API neutrality;
  • model routing;
  • interoperability;
  • vertical foreclosure;
  • innovation effects.

29. A Useful Legal Test

A structured AI-latency competition analysis can use the following framework:

1. Market Power

Does the undertaking possess substantial power in compute, cloud, inference or AI applications?

↓

2. Competitive Parameter

Is latency materially important to customers?

↓

3. Conduct

Has the undertaking manipulated, degraded or discriminated in latency?

↓

4. Rival Dependence

Do competitors depend upon the relevant infrastructure?

↓

5. Foreclosure

Does the conduct materially impair rival access or competitiveness?

↓

6. Effects

Are price, quality, innovation, choice or market entry harmed?

↓

7. Legitimate Justification

Can the latency difference be explained by genuine technical or efficiency considerations?

↓

8. Remedy

Would access, non-discrimination, interoperability or other measures restore competition?

30. Conclusion

AI inference latency is emerging as an important non-price dimension of competition. It can affect product quality, customer choice, market definition, entry barriers, vertical foreclosure and innovation.

The most important competition-law distinction is between:

superior latency achieved through innovation

and

artificial latency differences created through exclusionary control of infrastructure or platforms.

The existing case law—particularly Microsoft, Intel, Google Shopping, Slovak Telekom, Bronner and IMS Health—provides useful legal frameworks even though those decisions did not directly concern AI inference latency.

Accordingly, future AI competition investigations may increasingly treat TTFT, P95/P99 latency, throughput, accelerator access, API throttling and technical prioritisation as evidence of how competition actually occurs in AI markets.

 

 

LEAVE A COMMENT