Ai Infrastructure Stack Concentration (Chips, Cloud, Models, Apps) .
AI Inference Latency as a Competitive Parameter
1. Introduction
AI inference latency is the time between an AI system receiving a request and producing a usable response. In generative AI, it may include time-to-first-token (TTFT), inter-token latency, total response time, and end-to-end application latency.
Competition law traditionally focuses on price, output, quality, innovation and consumer choice. In AI markets, however, latency can itself be a dimension of competition. Two AI models may provide broadly comparable answers, yet the provider offering substantially faster inference can obtain an important competitive advantage.
Latency becomes particularly significant where AI is integrated into:
- search engines;
- coding assistants;
- financial trading systems;
- healthcare applications;
- autonomous systems;
- customer-service platforms;
- advertising technology;
- cloud APIs;
- enterprise software;
- real-time translation;
- robotics and industrial systems.
A dominant AI or cloud provider could therefore potentially harm competition not merely by raising prices, but by degrading rivals' access to low-latency infrastructure, prioritising its own inference traffic, tying model access to proprietary hardware, imposing discriminatory API conditions, or using technical architecture to foreclose competing AI services.
2. Meaning of AI Inference Latency
The AI inference process can broadly be represented as:
User request → API/network → model processing → first token → subsequent tokens → completed response
Latency can consequently be divided into several competitive variables.
A. Time-to-First-Token
TTFT measures how quickly the system begins responding.
This is particularly important for:
- conversational AI;
- search;
- customer-service chatbots;
- coding assistants.
B. Inter-Token Latency
This measures the interval between successive generated tokens.
A system may therefore have:
- fast initial response but slow generation; or
- slower initial response but rapid subsequent generation.
C. End-to-End Latency
This measures the complete time required to produce the requested output.
D. Tail Latency
The P95/P99 latency experienced by users can be more commercially important than average latency.
For example:
Provider A: average response = 300 ms
Provider B: average response = 280 ms
The difference may appear insignificant.
But if Provider A has reliable P99 latency while Provider B experiences substantial delays during peak demand, enterprise customers may prefer Provider A.
3. Why Latency Is a Competition Variable
AI competition is increasingly multi-dimensional.
Instead of:
Price → Quality → Output
the competitive framework becomes:
Price + Accuracy + Reliability + Latency + Privacy + Interoperability + Innovation
Latency can therefore constitute a form of non-price competition.
A customer may select an AI provider because:
- it responds faster;
- its service remains fast during peak demand;
- its API provides predictable response times;
- its geographic infrastructure is closer to users;
- its accelerator architecture reduces inference time;
- its model is optimised for real-time applications.
4. Relationship Between Compute and Latency
Inference latency is closely connected to infrastructure.
Important inputs include:
- GPUs;
- AI accelerators;
- TPUs;
- memory bandwidth;
- networking;
- data-centre location;
- model architecture;
- quantisation;
- batching;
- caching;
- inference software;
- compiler optimisation;
- model size.
This creates an important competition-law question:
Can control over an upstream AI-infrastructure input allow a dominant undertaking to control downstream AI competition through latency?
For example:
AI accelerator dominance
↓
Limited access to high-performance compute
↓
Higher inference latency for rivals
↓
Inferior user experience
↓
Customer migration
↓
Reduced downstream competition
5. Latency as a Quality Parameter
Competition law frequently recognises that quality can be competitively significant even where price does not change.
AI makes this particularly important because a slower AI product may effectively be a lower-quality product.
For example, a coding assistant that generates correct code in 1 second and another that generates comparable code in 8 seconds are not necessarily equivalent products.
Latency may influence:
- developer productivity;
- user retention;
- advertising conversion;
- search engagement;
- customer-service completion rates;
- trading execution;
- medical workflow efficiency.
Consequently, an AI provider may compete aggressively through latency optimisation rather than price reduction.
6. Relevant Competition-Law Theories
A. Abuse of Dominance
Where an undertaking possesses substantial market power, deliberate manipulation of latency could potentially become relevant to an abuse-of-dominance analysis.
Examples include:
- slowing access to competing AI models;
- imposing artificial throttling;
- discriminatory API rate limits;
- denying access to high-performance inference infrastructure;
- reserving scarce accelerator capacity for the dominant firm's own services.
The relevant question would not simply be:
"Is the competitor slower?"
It would be:
"Did the dominant undertaking engage in conduct that materially disadvantaged competitors and harmed competition?"
7. Refusal or Discrimination in Access to Compute
Suppose a cloud provider controls a scarce supply of advanced AI accelerators.
It provides:
- its own AI models → premium accelerator access;
- affiliated applications → priority scheduling;
- independent AI developers → restricted capacity.
The resulting latency difference could become competitively significant.
This could raise issues concerning:
- discriminatory access;
- refusal to supply;
- essential-facility-type theories;
- self-preferencing;
- vertical foreclosure;
- exclusionary conduct.
However, mere superiority in infrastructure does not automatically constitute an antitrust violation. Legitimate technical, security, capacity and efficiency explanations would also have to be considered.
8. Self-Preferencing Through Latency
One particularly important AI scenario is latency-based self-preferencing.
Suppose an AI platform operates:
- its own foundation model;
- an AI marketplace;
- an inference API;
- competing third-party models.
The platform could theoretically give its own model:
- priority GPU scheduling;
- lower network latency;
- preferential caching;
- higher rate limits;
- reserved inference capacity.
Third-party models might technically remain available but become commercially inferior because they respond more slowly.
This is potentially more difficult to detect than an outright exclusion.
9. AI APIs and Latency Discrimination
API markets introduce another issue.
A platform may offer:
Internal AI service:
50 ms latency
Third-party AI service:
250 ms latency
Even if both models have similar accuracy, developers may integrate the internal service because latency affects the performance of their own applications.
Potential competitive concerns include:
- discriminatory routing;
- unequal API infrastructure;
- throttling;
- differential rate limits;
- preferential caching;
- geographic routing discrimination;
- differential access to inference accelerators.
10. Latency and Network Effects
Latency can reinforce network effects.
A simplified cycle is:
More users
↓
More inference demand
↓
More infrastructure investment
↓
Lower latency
↓
Better user experience
↓
More users
↓
More data and revenue
↓
Further infrastructure investment
This can create a latency-based competitive feedback loop.
A well-funded incumbent may therefore achieve advantages that smaller competitors cannot easily reproduce.
11. Latency and Economies of Scale
Inference infrastructure has substantial fixed and variable costs.
Large providers can potentially:
- purchase accelerators in bulk;
- build specialised data centres;
- optimise inference software;
- deploy geographically distributed infrastructure;
- negotiate favourable network arrangements;
- operate large inference clusters.
As scale increases, the provider may reduce latency and cost simultaneously.
This may create economies of scale that constitute a barrier to entry.
12. Latency and Switching Costs
Latency can also contribute to customer lock-in.
An enterprise may build its software around:
- a particular inference API;
- a particular model;
- a particular accelerator architecture;
- proprietary optimisation tools.
After substantial integration, switching to another provider may increase latency.
The customer therefore faces:
migration costs + engineering costs + latency degradation.
This can strengthen incumbent market power.
13. Latency and Vertical Integration
Vertical integration creates particularly interesting competition questions.
Consider:
AI accelerator manufacturer
↓
Cloud infrastructure
↓
Inference platform
↓
Foundation model
↓
AI application
A vertically integrated undertaking may control several layers simultaneously.
If it can optimise latency throughout the stack, independent competitors may have difficulty matching the performance.
The competition issue is therefore not simply:
"Does the company have the fastest model?"
It is:
"Does control over multiple complementary markets allow the undertaking to exclude equally efficient competitors?"
14. Six Important Case Laws
Because there are relatively few reported decisions specifically concerning AI inference latency, the most useful authorities are established competition cases involving quality, technical performance, access to infrastructure, interoperability, self-preferencing, exclusion and innovation. Their application to AI latency is therefore analogical rather than a claim that the courts decided AI-latency disputes.
Case 1: United Brands v Commission
Case: United Brands Company and United Brands Continentaal BV v Commission, Case 27/76, Court of Justice of the European Communities (1978).
Principle
The Court examined market power and the competitive conditions surrounding the relevant product.
Relevance to AI latency
The case demonstrates that competition analysis cannot be reduced mechanically to price.
For AI services, relevant competitive parameters can include:
- performance;
- responsiveness;
- reliability;
- technical characteristics;
- customer preferences.
Therefore, when defining an AI inference market, latency may help establish whether customers regard two AI services as sufficiently substitutable.
AI application
If users systematically switch between models based upon response speed, latency can form part of the relevant product-market analysis.
Case 2: Microsoft Corp. v Commission
Case: Microsoft Corp. v Commission, Case T-201/04, General Court (2007).
Principle
The case involved Microsoft's control over important software technologies and interoperability issues.
The Court accepted the significance of access to interoperability information in enabling competing products to operate effectively.
AI latency relevance
AI ecosystems similarly depend upon interoperability among:
- models;
- APIs;
- cloud infrastructure;
- applications;
- data;
- accelerators.
If a dominant platform provides substantially better technical performance to its own downstream product while restricting competitors' technical access, latency could become part of a broader exclusionary strategy.
Key lesson
Technical interoperability can be a competition parameter, not merely a technical matter.
Case 3: Intel v Commission
Case: Intel Corporation v European Commission, Case C-413/14 P, Court of Justice (2017).
Principle
The case concerned exclusionary conduct involving rebates and the assessment of whether conduct was capable of producing anticompetitive foreclosure.
The Court emphasised the importance of analysing the actual competitive effects where appropriate.
AI latency relevance
Suppose an AI infrastructure provider provides favourable inference capacity to certain customers while making competing AI providers operate under materially worse latency conditions.
Competition analysis should examine:
- market coverage;
- duration;
- actual or potential foreclosure;
- competitive alternatives;
- customer dependence;
- efficiency explanations.
AI lesson
Latency advantages should not be treated as automatically anticompetitive. The mechanism and effects matter.
Case 4: Google Shopping
Case: Google and Alphabet v Commission, Case T-612/17, General Court (2021).
Principle
The case concerned Google's treatment of its own comparison-shopping service within its general search results.
The central competition issue included the possibility that a dominant platform could use its position in one market to advantage its own service in another.
AI latency relevance
The analogous AI scenario is:
Dominant AI platform
↓
Controls infrastructure/interface
↓
Own model receives preferential technical treatment
↓
Third-party models receive inferior latency
↓
Users disproportionately choose the integrated model
This illustrates how technical or algorithmic preference can become a competitive parameter.
Important distinction
Google Shopping did not concern AI inference latency. Its relevance is the broader principle concerning conduct by a dominant platform that advantages its own downstream service.
Case 5: Slovak Telekom
Case: European Commission v Slovak Telekom a.s. and Deutsche Telekom AG, Joined Cases C-152/19 P and C-165/19 P (2021).
Principle
The case involved exclusionary conduct concerning access to telecommunications infrastructure.
The Court examined how a dominant undertaking's control over an upstream network could affect downstream competition.
AI latency relevance
AI inference infrastructure increasingly resembles a layered network:
Compute → networking → cloud → inference → application
Where a dominant infrastructure provider controls an important upstream input, discriminatory access can potentially affect downstream AI competitors.
A latency differential may therefore be evidence of:
- discriminatory access;
- degradation;
- foreclosure;
- strategic prioritisation.
Case 6: Bronner v Mediaprint
Case: Oscar Bronner GmbH & Co. KG v Mediaprint Zeitungs und Zeitschriftenverlag GmbH & Co. KG, Case C-7/97, Court of Justice (1998).
Principle
The Court established a stringent framework for refusal-to-supply claims involving infrastructure.
A facility must generally be indispensable, and the refusal must be capable of eliminating effective competition, among other requirements.
AI latency relevance
An AI competitor might argue:
"Access to this particular inference infrastructure is necessary to compete at commercially viable latency."
That does not automatically establish an antitrust violation.
The Bronner framework highlights the importance of determining:
- whether the infrastructure is genuinely indispensable;
- whether alternatives exist;
- whether replication is realistically possible;
- whether refusal eliminates effective competition.
AI significance
This is especially relevant to:
- scarce GPUs;
- specialised AI accelerators;
- hyperscale inference infrastructure;
- proprietary inference networks.
Case 7: IMS Health
Case: IMS Health GmbH & Co. OHG v NDC Health GmbH & Co. KG, Joined Cases C-418/01 and C-7/01, Court of Justice.
Principle
The case developed the exceptional circumstances surrounding refusal to license intellectual-property-related infrastructure.
AI latency relevance
AI infrastructure increasingly involves proprietary:
- inference software;
- optimisation systems;
- model-serving technology;
- accelerator interfaces;
- specialised datasets.
Where a dominant undertaking controls technology that competitors require to achieve commercially viable performance, questions may arise concerning whether refusal or discriminatory access crosses the competition-law threshold.
Again, technical indispensability must be established rather than assumed.
15. Comparative Case-Law Matrix
| Case | Core principle | AI latency relevance |
|---|---|---|
| United Brands | Market power and product characteristics | Latency can be a product-quality parameter |
| Microsoft | Interoperability and technical access | API/infrastructure interoperability |
| Intel | Exclusionary conduct and foreclosure analysis | Latency discrimination may require effects analysis |
| Google Shopping | Preferential treatment by dominant platform | Self-preferencing through technical performance |
| Slovak Telekom | Access to upstream infrastructure | AI compute/inference infrastructure |
| Bronner | Refusal to supply / indispensability | Access to low-latency infrastructure |
| IMS Health | Exceptional access/licensing circumstances | Proprietary AI technology and infrastructure |
16. Latency-Based Foreclosure Theory
A particularly important theory can be expressed as follows:
Step 1 — Dominant infrastructure
A company controls scarce AI compute.
Step 2 — Differential allocation
Its own services receive priority access.
Step 3 — Latency divergence
Competitors experience:
- higher TTFT;
- greater P95 latency;
- throttling;
- lower throughput.
Step 4 — Customer migration
Developers choose the incumbent's model because applications need fast responses.
Step 5 — Network effects
More users generate more revenue and investment.
Step 6 — Entrenchment
Rivals lose scale and cannot economically reproduce the incumbent's infrastructure.
This is a classic vertical foreclosure hypothesis adapted to AI.
17. Latency as an Essential Competitive Parameter
Latency becomes particularly important where the application itself is time-sensitive.
High-latency sensitivity
- algorithmic trading;
- autonomous vehicles;
- robotics;
- gaming;
- live translation;
- fraud detection;
- cybersecurity;
- industrial control;
- customer-service voice systems.
Lower-latency sensitivity
- long-form document analysis;
- batch processing;
- offline summarisation;
- archival research.
Consequently, the relevant market may need to be segmented according to use case and latency sensitivity.
18. Latency and Market Definition
Competition authorities may consider whether:
AI Model A with 100 ms latency
and
AI Model B with 1,000 ms latency
are actually substitutes for a real-time application.
The answer may differ by customer.
For example:
| Application | Importance of latency |
|---|---|
| Real-time voice assistant | Very high |
| Search | High |
| Coding assistant | High |
| Customer support | High |
| Financial analytics | Potentially very high |
| Medical documentation | Moderate |
| Long-form research | Lower |
| Batch data processing | Relatively lower |
Thus, latency can affect both product-market definition and competitive-effects analysis.
19. P95 and P99 Latency as Competition Metrics
Average latency may conceal important competitive differences.
For example:
| Provider | Average | P95 | P99 |
|---|---|---|---|
| A | 300 ms | 450 ms | 600 ms |
| B | 250 ms | 1,500 ms | 4,000 ms |
Provider B may appear superior using average latency, but enterprise customers may prefer Provider A because its performance is more predictable.
Competition authorities should therefore potentially examine:
- median latency;
- P95;
- P99;
- peak-period performance;
- geographic variation;
- capacity throttling.
20. Latency Degradation as a Possible Exclusionary Strategy
A particularly subtle form of conduct would be strategic degradation.
Rather than banning competitors, a dominant platform could theoretically:
- increase API processing delays;
- reduce priority;
- restrict accelerator availability;
- impose lower throughput limits;
- increase queue times;
- deprioritise third-party inference workloads.
The service technically remains available.
But its competitive quality deteriorates.
This is analogous to competition concerns involving technical degradation in other network and platform markets.
21. Efficiency Defences
Latency advantages are not inherently unlawful.
A provider may legitimately have lower latency because of:
- better hardware;
- better algorithms;
- superior model architecture;
- efficient quantisation;
- better caching;
- geographically distributed data centres;
- legitimate capacity planning;
- investment in infrastructure.
Competition law should therefore distinguish:
Legitimate innovation
"Our model is faster because we developed superior inference technology."
from:
Potential exclusion
"We deliberately slow competitors using infrastructure they depend upon."
The former can represent vigorous competition; the latter may raise competition concerns depending upon market power, conduct and effects.
22. Consumer Welfare Effects
Latency can produce both positive and negative welfare effects.
Positive effects
Competition on latency can generate:
- faster AI services;
- better productivity;
- lower waiting time;
- improved accessibility;
- more responsive applications;
- technological innovation.
Potential negative effects from exclusion
If competition is reduced, consumers may ultimately face:
- higher prices;
- fewer AI providers;
- slower innovation;
- reduced choice;
- greater dependence on one ecosystem;
- weaker privacy or quality competition.
23. Evidence Required in an AI Latency Investigation
A competition authority examining latency should potentially collect:
Technical evidence
- TTFT;
- token-generation speed;
- P95/P99 latency;
- throughput;
- GPU utilisation;
- accelerator allocation;
- API logs.
Commercial evidence
- customer switching data;
- contracts;
- pricing;
- customer complaints;
- procurement documents;
- internal strategy documents.
Infrastructure evidence
- GPU availability;
- cloud capacity;
- networking arrangements;
- geographical deployment;
- data-centre architecture.
Algorithmic evidence
- scheduling algorithms;
- routing systems;
- caching;
- throttling;
- queue prioritisation.
This allows the authority to determine whether latency differences arise from legitimate technical superiority or exclusionary conduct.
24. Possible Remedies
If competition harm were established, potential remedies could include:
A. Non-discrimination
Require equivalent infrastructure access for competing services under comparable conditions.
B. API transparency
Require disclosure of relevant:
- rate limits;
- throttling;
- service-level conditions.
C. Interoperability
Allow competitors to connect to relevant infrastructure on reasonable terms.
D. Structural separation
In exceptional cases, competition authorities could consider separation between infrastructure and downstream AI services.
E. Monitoring
Require periodic reporting of:
- latency;
- throughput;
- capacity allocation;
- outage rates.
F. Behavioural commitments
Prohibit preferential scheduling or discriminatory technical treatment.
25. China Competition-Law Perspective
Under China's Anti-Monopoly Law, AI latency could become relevant primarily through established categories such as:
- abuse of dominant market position;
- refusal to deal;
- discriminatory treatment;
- tying;
- unreasonable trading conditions;
- exclusionary conduct;
- concentration-related concerns.
The important analytical question would be whether latency manipulation constitutes a mechanism through which market power is exercised or competitors are foreclosed.
Chinese digital-platform enforcement also makes algorithmic conduct, platform rules, data and technological infrastructure increasingly relevant to competition analysis.
26. India Competition-Law Perspective
Under the Competition Act 2002, latency can potentially be considered within broader questions concerning:
- relevant market;
- market power;
- denial of market access;
- discriminatory conditions;
- leveraging;
- refusal to deal;
- tying/bundling;
- abuse of dominant position.
For example, if a dominant cloud platform provides its own AI service with materially better inference infrastructure while imposing discriminatory technical restrictions on competing AI providers, latency could form part of the evidence concerning denial of market access or leveraging.
27. United States Perspective
Under U.S. antitrust principles, latency can potentially matter in:
- monopolisation analysis;
- exclusionary conduct;
- vertical foreclosure;
- essential-input disputes;
- platform competition;
- merger review.
The key distinction remains between:
competition on the merits
and
artificial conduct designed to exclude rivals.
A superior AI architecture that legitimately reduces latency is generally evidence of innovation rather than exclusion.
28. Merger-Control Implications
AI mergers may raise latency concerns even where traditional market-share analysis appears modest.
Consider:
Cloud Provider + Foundation Model Developer
The transaction may combine:
- compute;
- cloud;
- model;
- API;
- application distribution.
The combined undertaking could potentially improve its own model's latency while disadvantaging rival models.
Authorities may therefore examine:
- access to compute;
- accelerator capacity;
- API neutrality;
- model routing;
- interoperability;
- vertical foreclosure;
- innovation effects.
29. A Useful Legal Test
A structured AI-latency competition analysis can use the following framework:
1. Market Power
Does the undertaking possess substantial power in compute, cloud, inference or AI applications?
↓
2. Competitive Parameter
Is latency materially important to customers?
↓
3. Conduct
Has the undertaking manipulated, degraded or discriminated in latency?
↓
4. Rival Dependence
Do competitors depend upon the relevant infrastructure?
↓
5. Foreclosure
Does the conduct materially impair rival access or competitiveness?
↓
6. Effects
Are price, quality, innovation, choice or market entry harmed?
↓
7. Legitimate Justification
Can the latency difference be explained by genuine technical or efficiency considerations?
↓
8. Remedy
Would access, non-discrimination, interoperability or other measures restore competition?
30. Conclusion
AI inference latency is emerging as an important non-price dimension of competition. It can affect product quality, customer choice, market definition, entry barriers, vertical foreclosure and innovation.
The most important competition-law distinction is between:
superior latency achieved through innovation
and
artificial latency differences created through exclusionary control of infrastructure or platforms.
The existing case law—particularly Microsoft, Intel, Google Shopping, Slovak Telekom, Bronner and IMS Health—provides useful legal frameworks even though those decisions did not directly concern AI inference latency.
Accordingly, future AI competition investigations may increasingly treat TTFT, P95/P99 latency, throughput, accelerator access, API throttling and technical prioritisation as evidence of how competition actually occurs in AI markets.

comments