Ai Infrastructure Dependency Mapping In Markets .
AI Inference Latency as a Competitive Parameter
Detailed Explanation with At Least 6 Case Laws
1. Introduction
AI inference latency is the time between an AI system receiving a request and producing a response. In generative AI, it may be measured through metrics such as:
- Time to First Token (TTFT) – time before the first output appears;
- Time Per Output Token (TPOT) – time required to generate subsequent tokens;
- end-to-end response latency;
- tail latency – performance at the slowest percentiles, such as P95 or P99;
- queueing delay;
- network/API latency; and
- latency caused by model routing, GPU availability, memory bandwidth or cloud infrastructure.
Latency can therefore become a non-price competitive parameter, alongside price, accuracy, reliability, privacy and functionality.
This is particularly significant where AI services are used for real-time search, coding assistants, customer service, autonomous systems, fraud detection, financial trading, healthcare applications, gaming or industrial control. A model that is marginally less accurate but responds substantially faster may be commercially preferable for certain applications.
Importantly, there is not yet a mature body of reported antitrust case law specifically holding that AI inference latency itself constitutes an independent competition parameter. The existing cases provide analogies through the treatment of quality, innovation, interoperability, access, exclusionary conduct and non-price competition.
2. Why Latency Can Be a Competition Parameter
Traditional competition analysis often focuses on:
Price + output + market share
Digital and AI markets increasingly require:
Price + quality + speed + innovation + interoperability + access
Latency matters because it can directly influence the commercial usefulness of an AI model.
For example:
| AI application | Importance of latency |
|---|---|
| Conversational AI | High |
| Search assistants | Very high |
| Coding assistants | High |
| Autonomous vehicles | Extremely high |
| Fraud detection | Extremely high |
| Financial trading | Extremely high |
| Healthcare decision support | High |
| Customer-service bots | High |
| Batch document analysis | Relatively lower |
| AI model training | Latency less directly important than throughput |
Thus, competition authorities may need to examine whether a firm is using control over compute, GPUs, cloud infrastructure, APIs or model routing to manipulate latency experienced by competitors.
3. Latency as a Non-Price Dimension of Competition
Suppose an AI platform provides its own model with:
- 100 ms API response time,
while an independent competing model receives:
- 1,500 ms response time,
despite technically comparable infrastructure.
The platform could potentially obtain a competitive advantage without increasing its price.
This raises an important antitrust question:
Can deliberately degraded latency constitute discriminatory or exclusionary treatment?
Potential theories include:
- refusal to supply;
- discriminatory access;
- self-preferencing;
- tying or leveraging;
- raising rivals' costs;
- foreclosure;
- margin squeeze;
- interoperability discrimination;
- exclusionary technical design; and
- exploitation of an essential infrastructure bottleneck.
4. AI Inference Stack and Latency Control
Latency can be controlled at several layers.
A. Hardware layer
- GPUs
- TPUs
- AI accelerators
- memory bandwidth
- interconnects
B. Cloud layer
- compute allocation
- geographic location
- capacity reservation
- network architecture
- scheduling
C. Model layer
- model size
- quantisation
- speculative decoding
- batching
- caching
- mixture-of-experts routing
D. API layer
- rate limits
- priority queues
- throttling
- geographic routing
- token limits
E. Distribution layer
- default model selection
- search integration
- operating-system integration
- application-store integration
- assistant routing
Consequently, a company controlling several layers could potentially influence rivals' latency without explicitly refusing access.
The FTC has already identified computing resources, switching costs and vertical relationships between cloud providers and AI developers as competition issues requiring scrutiny.
5. Case Law
Case 1 — United Brands v Commission
Case 27/76, United Brands v Commission (1978)
Principle
The European Court of Justice recognised that competition analysis cannot be reduced exclusively to price. Product characteristics, quality and consumer preferences can be relevant to the competitive assessment.
Relevance to AI latency
Latency can be treated as a quality characteristic of an AI service.
Consider two AI APIs:
- Model A: 95% accuracy, 100 ms response;
- Model B: 96% accuracy, 2 seconds response.
For a real-time application, customers may value Model A more highly.
Therefore, competition authorities examining AI markets may need to consider:
speed-adjusted quality rather than accuracy alone.
Competition significance
A dominant firm could theoretically reduce a rival's competitiveness by:
- degrading API response speed;
- allocating inferior infrastructure;
- imposing artificial queueing;
- limiting geographic deployment;
- restricting access to high-performance accelerators.
Latency can consequently form part of the quality dimension of competition.
6. Case 2 — Intel v Commission
Case C-413/14 P, Intel Corp. v Commission (2017)
Principle
The Intel litigation concerned exclusionary rebates and the competitive effects of conduct by a dominant undertaking.
The case is particularly relevant because the Court emphasised the importance of examining whether conduct is capable of restricting competition rather than relying exclusively upon formal classifications.
Application to AI latency
Suppose an AI/cloud provider offers:
- extremely low latency to its own AI model;
- premium GPU scheduling to its own downstream applications;
- slower inference access to competing AI developers.
Even if competitors technically remain able to access the infrastructure, the practical competitive conditions may be unequal.
The question becomes:
Does the latency differential materially impair the rival's ability to compete?
This resembles the broader logic of examining the actual competitive capability of rivals rather than merely asking whether access formally exists.
AI example
A cloud provider could theoretically give its affiliated AI assistant:
P99 latency = 300 ms
while independent AI providers receive:
P99 latency = 2,500 ms.
The independent provider may technically have access, but its product could become commercially unattractive for real-time applications.
7. Case 3 — Google Shopping
Case T-612/17, Google and Alphabet v Commission
Principle
The Google Shopping litigation concerned Google's treatment of competing comparison-shopping services within its search ecosystem.
The case demonstrates the importance of technical positioning and platform-controlled distribution in digital competition.
Relevance to AI inference latency
AI platforms increasingly perform model selection and routing.
A platform might:
- route its own model through premium infrastructure;
- route rival models through slower infrastructure;
- place its own model as the default;
- give its own model preferential compute allocation.
This creates a potential form of technical self-preferencing.
Hypothetical example
An AI operating system has access to five models:
| Model | Average latency |
|---|---|
| Platform's model | 250 ms |
| Rival A | 900 ms |
| Rival B | 1,200 ms |
| Rival C | 1,500 ms |
If the platform controls routing and infrastructure, competition analysis could ask whether the latency differences arise from legitimate technical factors or preferential treatment.
The lesson from Google Shopping is that control over the interface through which consumers access competing services can itself have competitive significance.
8. Case 4 — Bronner v Mediaprint
Case C-7/97, Oscar Bronner GmbH v Mediaprint (1998)
Principle
The case established important limits concerning refusal to supply and essential facilities under Article 102 TFEU.
Access to infrastructure controlled by a dominant firm does not automatically create an obligation to provide access. The stringent conditions associated with the essential-facilities doctrine remain important.
Relevance to AI inference
AI inference may increasingly depend upon infrastructure such as:
- high-end GPUs;
- AI accelerators;
- hyperscale cloud;
- specialised inference clusters;
- high-bandwidth interconnects.
If a particular infrastructure becomes indispensable for competing AI applications, the question could arise whether denying or materially degrading access amounts to exclusionary conduct.
Latency-specific issue
The relevant conduct need not necessarily be:
"You cannot access our infrastructure."
It could instead be:
"You can access it, but your requests are consistently placed in a substantially slower service tier."
That makes quality-of-access potentially as important as access itself.
9. Case 5 — Slovak Telekom v Commission
Joined Cases C-152/19 P and C-165/19 P
Principle
The case involved exclusionary conduct and access conditions in telecommunications infrastructure.
It is particularly useful for AI because telecommunications competition has historically demonstrated that technical infrastructure and access conditions can affect downstream competition.
Relevance to AI inference
AI infrastructure increasingly resembles a vertically integrated stack:
GPU → cloud → inference API → AI application → consumer
A dominant infrastructure provider could potentially operate at multiple levels simultaneously.
For example:
Cloud provider → AI chip → inference service → AI assistant
If the provider gives its own downstream AI service superior infrastructure performance, rivals may face higher effective costs.
Latency as a competitive variable
The relevant disadvantage might consist of:
- longer processing queues;
- lower GPU priority;
- slower networking;
- inferior geographic routing;
- lower maximum concurrency;
- lower throughput.
Thus, latency discrimination can potentially operate as a non-price access restriction.
10. Case 6 — Qualcomm (Exclusivity Payments)
Commission Decision, Case AT.40220 — Qualcomm (Exclusivity Payments), 2018
Principle
The Qualcomm case concerned conduct by a major technology supplier and the competitive effects of arrangements involving downstream manufacturers.
The case illustrates how contractual arrangements involving important technological inputs can affect the ability of competitors to reach downstream markets.
Relevance to AI
AI infrastructure increasingly involves contractual relationships between:
- cloud providers;
- AI developers;
- chip manufacturers;
- model developers;
- application providers.
The FTC's 2025 AI partnership study specifically identified cloud commitments, computing resources, switching costs, exclusivity-related arrangements and access to sensitive technical information as issues potentially affecting competition.
Latency connection
If an AI developer is contractually tied to a particular cloud or infrastructure provider, it may become difficult to obtain alternative infrastructure capable of delivering equivalent latency.
Thus:
contractual lock-in → infrastructure dependency → latency disadvantage → reduced competitive ability
can become a possible theory of harm.
11. Case 7 — Microsoft / Commission
Case T-201/04, Microsoft Corp. v Commission (2007)
Principle
The Microsoft litigation involved interoperability and access to information necessary for competing products.
The case demonstrates the importance of technical interoperability as a competitive condition.
Application to AI
AI ecosystems increasingly require interoperability among:
- models;
- APIs;
- cloud environments;
- vector databases;
- inference accelerators;
- developer tools;
- agent frameworks.
If interoperability restrictions cause competitors to experience substantially worse inference performance, the competitive impact may extend beyond simple denial of access.
Example
A dominant platform could make its own model capable of:
- direct memory access;
- privileged caching;
- native accelerator support;
- priority routing;
while competing models must use a slower interface.
The result may be:
formal interoperability + practical performance discrimination.
This is particularly important because an API that technically works but operates at dramatically inferior latency may not provide meaningful competitive equivalence.
12. Case 8 — Google Android
Case AT.40099 — Google Android
The European Commission's Android decision concerned Google's conduct relating to mobile operating systems, applications and distribution.
Relevance
Android demonstrates the importance of defaults, distribution and ecosystem integration.
AI assistants are increasingly embedded in:
- smartphones;
- browsers;
- operating systems;
- search engines;
- productivity software.
Suppose an operating-system provider:
- gives its own AI assistant direct access to device hardware;
- gives it privileged local inference;
- gives competing assistants only cloud access.
The rival could have inferior latency even if both assistants have similar underlying model quality.
Competition concern
The competitive advantage would arise not from superior AI alone, but from ecosystem-controlled technical privileges.
13. Case 9 — Google Search / Self-Preferencing
The broader Google search litigation also provides an important framework for understanding algorithmically mediated competition.
The central lesson is that where a platform controls an important access point, its technical and ranking decisions may affect competitors' ability to obtain users.
For AI, this can translate into:
model selection → routing → latency → user retention
A platform that automatically chooses which model answers a query can potentially influence competition through both ranking and performance allocation.
14. Emerging AI-Specific Competition Context
Although the traditional cases above do not directly decide "AI latency" disputes, contemporary regulatory work increasingly recognises infrastructure as a potential bottleneck.
The FTC has stated that computational resources are a key input into generative AI and has identified risks arising from concentrated cloud and specialised-chip markets.
The OECD similarly identifies physical AI infrastructure, including advanced computing resources, as an emerging competition-policy concern.
In 2026, the European Commission also identified the importance of low-latency compute capacity in its proposed Cloud and AI Development Act.
The Commission has separately indicated a preliminary view that AWS and Azure should be designated as DMA gatekeepers for cloud computing services, citing entrenched positions, switching costs and AI-related factors in cloud procurement.
These developments are important because they show that low-latency computing is increasingly treated as an infrastructure and competitive-capacity issue, even though a specific legal doctrine concerning "AI inference latency discrimination" has not yet crystallised.
15. Possible Theories of Anticompetitive Latency Manipulation
A. Latency discrimination
A dominant infrastructure provider supplies different latency to:
- its own AI products; and
- competing AI products.
The critical question would be whether there is a legitimate technical justification.
B. Self-preferencing
A platform gives its own AI model:
- priority GPU access;
- superior caching;
- faster routing;
- privileged APIs.
Rivals technically remain available but perform worse.
C. Raising Rivals' Costs
A rival may have to purchase additional:
- GPUs;
- servers;
- geographic infrastructure;
- caching systems;
- networking capacity.
The rival's cost of achieving competitive latency therefore increases.
D. Margin Squeeze
Suppose a dominant cloud provider:
- supplies compute infrastructure to rival AI developers; and
- competes with those developers through its own AI application.
It could theoretically structure wholesale infrastructure prices and downstream AI prices in a manner that makes competitive latency economically difficult for rivals to achieve.
E. Refusal or Degradation of Access
Access could be technically available but commercially inadequate because of:
- artificial throttling;
- low priority;
- restricted throughput;
- inferior geographic routing;
- limited concurrency.
This raises questions similar to infrastructure-access cases.
16. Latency and Market Definition
Latency can also affect relevant-market definition.
For example, the relevant market might not simply be:
"AI model services."
It could potentially be segmented according to use:
General-purpose AI inference
versus
Real-time AI inference
versus
High-performance enterprise inference
versus
Ultra-low-latency inference
The appropriate definition would depend on evidence concerning:
- customer substitution;
- willingness to pay;
- technical requirements;
- switching possibilities;
- latency tolerance;
- model quality;
- throughput;
- geographic requirements.
A 20-second response may be acceptable for batch document analysis but unacceptable for autonomous driving.
Therefore:
The relevant competitive constraint may differ according to latency-sensitive use case.
17. Latency and Consumer Welfare
Latency affects consumers through several mechanisms.
Faster AI
can produce:
- better conversational interaction;
- quicker search;
- faster coding;
- reduced waiting;
- better real-time decision-making.
Artificially increased latency
could produce:
- degraded user experience;
- reduced adoption of rival AI;
- increased developer costs;
- reduced innovation;
- higher prices;
- reduced quality competition.
Therefore, a competition authority should not automatically treat latency as merely an engineering issue.
It can be a quality and innovation variable with economic consequences.
18. Evidence Required to Establish a Latency-Based Competition Concern
A serious investigation would require considerably more than simply showing that one AI model is faster.
Relevant evidence could include:
Technical evidence
- TTFT;
- P50/P95/P99 latency;
- tokens per second;
- throughput;
- GPU utilisation;
- queue times;
- network latency;
- geographic routing.
Commercial evidence
- customer switching;
- churn;
- price differences;
- willingness to pay for lower latency;
- contractual terms.
Internal documents
Potentially relevant evidence could include:
- infrastructure allocation policies;
- GPU scheduling policies;
- routing algorithms;
- API throttling rules;
- internal performance benchmarks;
- product strategy documents.
Comparator analysis
Investigators would ideally compare:
similarly situated customers + same workload + same infrastructure class + same geographic conditions.
That helps distinguish genuine technical limitations from discriminatory treatment.
19. Legitimate Reasons for Latency Differences
A latency difference is not automatically anticompetitive.
It may arise from legitimate factors such as:
- different model sizes;
- different token requirements;
- security controls;
- geographic distance;
- workload intensity;
- customer-selected service tiers;
- hardware architecture;
- optimisation techniques;
- congestion;
- reliability requirements.
Therefore, an antitrust analysis should distinguish:
legitimate performance differentiation
from
strategic performance degradation designed to disadvantage competitors.
20. Possible Remedies
If unlawful discriminatory conduct were established, potential remedies could include:
1. Non-discrimination
Comparable customers receive comparable infrastructure performance.
2. API transparency
Disclosure of:
- throttling;
- rate limits;
- service tiers;
- queue policies.
3. Interoperability
Competitors receive technically meaningful access.
4. Separation of infrastructure and downstream operations
Structural remedies could theoretically be considered in exceptional circumstances.
5. Monitoring
Independent auditing of:
- latency;
- throughput;
- routing;
- access conditions.
6. Data portability and switching
Reducing technical switching costs between cloud/inference providers.
21. Six-Case-Law Synthesis
| Case | Core principle | AI inference-latency relevance |
|---|---|---|
| United Brands v Commission | Quality can be relevant to competition | Latency can be a quality parameter |
| Intel v Commission | Examine competitive effects of exclusionary conduct | Inferior performance can potentially impair rival competition |
| Google Shopping | Platform-controlled technical positioning can affect rivals | Preferential routing/performance may create AI self-preferencing concerns |
| Bronner v Mediaprint | Essential-facilities/refusal-to-supply doctrine | Critical inference infrastructure may raise access questions |
| Slovak Telekom | Infrastructure access conditions can affect downstream competition | Latency/throughput may constitute important access conditions |
| Qualcomm | Technology-related contractual arrangements may affect foreclosure | Cloud/inference commitments can potentially reinforce infrastructure dependence |
| Microsoft | Interoperability can be competitively significant | Formal API access may be insufficient if technical design produces major performance disadvantages |
| Google Android | Ecosystem integration/defaults can reinforce platform power | Privileged on-device inference could disadvantage rival AI assistants |
22. Legal Test for AI Inference Latency
A useful analytical framework is:
Step 1 — Identify the relevant AI market
↓
Step 2 — Determine whether latency materially affects customer choice
↓
Step 3 — Identify control over compute/API/routing infrastructure
↓
Step 4 — Establish whether the undertaking has substantial market power
↓
Step 5 — Measure latency differential
↓
Step 6 — Determine whether similarly situated rivals receive materially different treatment
↓
Step 7 — Examine technical and commercial justification
↓
Step 8 — Determine foreclosure or exclusionary effect
↓
Step 9 — Examine consumer harm, innovation effects and entry barriers
↓
Step 10 — Consider proportionate remedies
23. Conclusion
AI inference latency is capable of becoming an important non-price competitive parameter. It is particularly significant where AI products compete in real-time applications and where a vertically integrated company controls both inference infrastructure and downstream AI services.
The existing case law does not yet establish a standalone doctrine of "AI latency discrimination." Instead, established principles concerning quality competition, exclusionary conduct, essential facilities, interoperability, self-preferencing, infrastructure access and vertical foreclosure provide the legal foundation for analysing the issue.
The emerging regulatory environment reinforces the importance of this question. The FTC has identified concentrated access to computing resources and switching costs as potential AI competition concerns, while European policy work increasingly treats low-latency compute capacity as strategically important infrastructure.
Accordingly, the central competition-law question is not simply:
"Which AI model is faster?"
but rather:
"Is the observed latency difference the result of legitimate technological competition, or is control over scarce inference infrastructure being used to disadvantage competing AI services?"
That distinction will likely become increasingly important as AI markets move from model-training competition toward large-scale, continuously deployed inference competition.

comments