Global Knowledge Graphs And Epistemic Infrastructure Concentration .

 

Global Knowledge Graphs And Epistemic Infrastructure Concentration

Introduction

Global knowledge graphs are large-scale systems that connect entities, facts, concepts, documents, people, organizations, locations, events, and relationships into machine-readable structures. They increasingly underpin search engines, recommendation systems, AI assistants, scientific discovery, financial intelligence, mapping, healthcare, government services, and automated decision-making.

Epistemic infrastructure is broader. It means the technological, institutional, and informational systems through which society determines:

  • what information is available;
  • how information is classified and connected;
  • which sources are considered authoritative;
  • what facts can be discovered;
  • how conflicting information is ranked;
  • which datasets AI systems can access;
  • and ultimately what information is capable of influencing decisions.

The competition-law concern arises when a small number of firms control critical layers of this infrastructure. Concentration may therefore produce epistemic market power: the ability to influence not merely prices or output, but the information environment within which consumers, businesses, researchers, governments and AI systems make decisions.

1. Meaning Of Knowledge-Graph Concentration

A knowledge graph generally contains:

Entities → attributes → relationships → sources → inference → outputs

For example:

Company A → owns → Company B → operates → Platform C → collects → Dataset D.

A large knowledge graph can therefore become a connective layer between enormous quantities of information.

Concentration can occur at several levels:

  1. Data acquisition
  2. Data aggregation
  3. Entity resolution
  4. Knowledge-graph construction
  5. Ranking and retrieval
  6. AI-model access
  7. API distribution
  8. User-interface control
  9. Institutional integration

The competitive problem becomes particularly serious when the same undertaking controls several layers simultaneously.

2. What Is Epistemic Infrastructure?

Traditional infrastructure consists of roads, electricity grids, telecommunications networks and payment systems.

Epistemic infrastructure performs a comparable function for information.

It may include:

  • search indexes;
  • knowledge graphs;
  • digital libraries;
  • scientific databases;
  • citation networks;
  • mapping databases;
  • identity databases;
  • data brokers;
  • cloud repositories;
  • AI training datasets;
  • model-access APIs;
  • content-discovery systems;
  • fact-checking systems;
  • recommendation engines;
  • metadata standards.

The distinction is important because an undertaking can possess substantial market power even where its service is nominally free.

The relevant competitive asset may be control over information flows rather than monetary transactions.

3. Sources Of Concentration

A. Data Scale

Knowledge systems benefit from enormous datasets.

The more entities and relationships a platform contains, the more useful its graph may become.

This can generate:

scale → better results → more users → more data → better results.

That feedback loop may create substantial entry barriers.

B. Network Effects

Users often prefer systems that contain the greatest number of entities and relationships.

A knowledge graph therefore exhibits a form of data/network effect.

An entrant may possess excellent technology but still struggle because it lacks:

  • historical data;
  • entity mappings;
  • user-generated corrections;
  • behavioural signals;
  • source relationships;
  • structured metadata.

C. Switching Costs

Businesses and public institutions may build workflows around a particular:

  • API;
  • identifier system;
  • ontology;
  • database;
  • knowledge graph;
  • cloud environment.

Once integrated, migration becomes expensive.

Switching costs can therefore transform technological superiority into durable market power.

D. Interoperability Barriers

If competing graphs cannot easily exchange:

  • identifiers;
  • metadata;
  • ontologies;
  • provenance;
  • semantic relationships;

then users may become locked into one ecosystem.

Interoperability is consequently an important competition-law remedy.

4. Epistemic Gatekeeping

The most significant concern is gatekeeping over discoverability.

Suppose information exists publicly but a dominant intermediary controls:

  1. indexing;
  2. ranking;
  3. entity classification;
  4. recommendation;
  5. AI retrieval.

The information technically remains available, but its practical accessibility may depend upon the intermediary.

This creates a distinction between:

formal availability and effective availability.

Competition authorities may therefore need to consider whether exclusion from a dominant knowledge system substantially reduces an organization's ability to reach users.

5. Knowledge Graphs As Essential Inputs

A knowledge graph can become an important input where downstream competitors depend upon it.

Potential downstream markets include:

  • AI assistants;
  • search;
  • travel;
  • financial analysis;
  • healthcare information;
  • scientific research;
  • mapping;
  • advertising;
  • cybersecurity;
  • enterprise intelligence.

If competitors cannot realistically reproduce the relevant graph, refusal to provide access may raise essential-facility or refusal-to-deal concerns, depending on the jurisdiction.

The analysis would normally require factors such as:

  • indispensability;
  • absence of realistic alternatives;
  • elimination of effective competition;
  • technical feasibility of access;
  • objective justification;
  • proportionality of the requested remedy.

6. Data Advantages And Competition

Data itself does not automatically constitute market power.

The relevant question is whether the particular data advantage is:

  • difficult to replicate;
  • commercially important;
  • continuously refreshed;
  • exclusive;
  • interoperable;
  • necessary for downstream competition.

A dataset becomes particularly powerful when combined with:

proprietary identifiers + behavioural data + search data + transactional data + computational infrastructure.

This creates a compound data advantage rather than a simple database advantage.

7. Vertical Integration

Knowledge infrastructure can create vertical competition problems.

For example:

Data collection → knowledge graph → AI model → search engine → advertising platform

If one undertaking controls the entire chain, it may have incentives to:

  • favour its downstream products;
  • deny competitors access;
  • degrade interoperability;
  • impose discriminatory API conditions;
  • self-preference its services;
  • use proprietary data unavailable to rivals.

This resembles traditional vertical foreclosure but operates through information infrastructure.

8. Self-Preferencing

A dominant platform may rank its own:

  • knowledge panels;
  • databases;
  • AI answers;
  • products;
  • services;
  • vertical search results

above competing sources.

The competitive concern is not necessarily that the preferred result is inaccurate.

The concern is that the platform may use control over the discovery layer to disadvantage competitors.

This makes self-preferencing especially significant where the platform is also the infrastructure upon which those competitors depend.

9. Epistemic Bias As A Competition Issue

Competition law traditionally asks whether conduct harms:

  • price;
  • quality;
  • output;
  • innovation;
  • consumer choice.

Knowledge infrastructure adds another dimension:

informational quality and diversity.

A dominant graph could influence:

  • which businesses are visible;
  • which scientific claims are discoverable;
  • which sources are treated as authoritative;
  • which products are recommended;
  • which facts AI systems retrieve.

Consequently, competition authorities may increasingly consider plurality of information sources as an element of quality and innovation.

10. AI Makes Knowledge-Graph Concentration More Important

Generative AI systems increasingly rely on structured and unstructured information.

A concentrated knowledge layer can therefore influence multiple AI systems simultaneously.

For example:

Knowledge Graph
↓
Retrieval system
↓
AI model
↓
Answer
↓
User decision

If a dominant infrastructure provider controls the first layer, it can indirectly influence downstream AI competition.

This creates a new form of upstream epistemic leverage.

11. Key Competition-Law Issues

1. Relevant Market Definition

Possible markets include:

  • general search;
  • specialized search;
  • knowledge-graph services;
  • structured data;
  • data brokerage;
  • AI retrieval;
  • enterprise knowledge services;
  • identity resolution.

Market definition may need to account for zero-price services and quality dimensions.

2. Dominance

Indicators may include:

  • data scale;
  • query volume;
  • graph coverage;
  • API dependency;
  • switching costs;
  • interoperability;
  • technological advantages;
  • institutional adoption.

3. Exclusionary Conduct

Potential theories include:

  • refusal to supply;
  • discriminatory access;
  • self-preferencing;
  • tying;
  • bundling;
  • exclusive dealing;
  • data foreclosure;
  • interoperability restrictions;
  • exploitative contractual terms.

4. Merger Control

Acquisitions of:

  • datasets;
  • specialized databases;
  • search technologies;
  • graph providers;
  • scientific information platforms;
  • AI retrieval companies

may create strategic data concentration even where traditional revenue-based thresholds underestimate the transaction's importance.

12. At Least 6 Important Case Laws

1. United States v. Google LLC — Search

The Google search litigation is highly relevant because it concerns control over the search-distribution ecosystem.

The case demonstrates how dominance can be reinforced through agreements affecting access points to search.

Relevance

For knowledge graphs, the analogous concern is whether control over distribution and default access points allows an undertaking to preserve its informational infrastructure advantage and prevent rivals from reaching sufficient scale.

Principle: Control over an important gateway can reinforce an already powerful information intermediary.

2. Google Search (Shopping) — European Commission

The European Commission's Google Shopping decision concerned preferential treatment of Google's own comparison-shopping service within its general search results.

Relevance

This is directly relevant to knowledge infrastructure because a dominant information intermediary can use its ranking architecture to favour an affiliated downstream service.

The case illustrates the competition-law importance of:

  • ranking;
  • visibility;
  • search neutrality;
  • self-preferencing;
  • traffic foreclosure.

Principle: Control over a discovery infrastructure can be used to disadvantage downstream competitors.

3. Google Android — European Commission

The Android case concerned Google's contractual arrangements surrounding Android, including restrictions involving search and application distribution.

Relevance

Knowledge infrastructure often operates through an ecosystem rather than a single product.

Control over:

  • operating systems;
  • app stores;
  • search;
  • defaults;
  • APIs

can reinforce control over information-discovery systems.

Principle: Competition analysis may examine how contractual restrictions across complementary technological layers reinforce dominance.

4. Microsoft Corp. v. United States

The Microsoft litigation remains a foundational technology-antitrust case.

Microsoft's conduct involving Windows and web browsers demonstrated how control over a dominant platform can be leveraged into adjacent technological markets.

Relevance

The same economic logic can apply to knowledge infrastructure:

dominant infrastructure → control over access → exclusion of complementary technologies → preservation of ecosystem power.

Principle: A dominant technological platform can use control over an important infrastructure layer to impede adjacent competition.

5. United States v. Microsoft Corp. — D.C. Circuit

The appellate decision is particularly important for its treatment of exclusionary conduct and network effects.

The court recognized that practices that appear individually small can become competitively significant when they protect a dominant position in a networked market.

Relevance

Knowledge graphs exhibit comparable feedback effects.

A dominant graph can become stronger because:

  • more users generate more signals;
  • more sources improve coverage;
  • greater coverage attracts more users;
  • greater usage increases downstream dependence.

Principle: Network effects can make exclusionary conduct particularly consequential.

6. Aspen Skiing Co. v. Aspen Highlands Skiing Corp.

The U.S. Supreme Court recognized an exceptional refusal-to-deal theory where a dominant firm terminated a previously profitable course of cooperation and thereby harmed competition.

Relevance

The case provides an important conceptual foundation for analyzing refusal to provide access to information infrastructure.

If a dominant knowledge provider historically supplied information or interoperability to rivals and later withdraws access strategically, the conduct may raise questions concerning exclusionary intent and competitive harm.

Principle: A strategically motivated termination of cooperation can, in exceptional circumstances, constitute unlawful exclusionary conduct.

7. Verizon Communications Inc. v. Law Offices of Curtis V. Trinko, LLP

Trinko limited the scope of compulsory-dealing theories under U.S. antitrust law.

Relevance

This is particularly important for knowledge graphs because it prevents competition law from automatically converting every proprietary information system into a compulsory public resource.

The case emphasizes the importance of:

  • preserving incentives to innovate;
  • distinguishing competition law from sector regulation;
  • establishing genuine anticompetitive conduct.

Principle: Dominance alone does not generally create a universal obligation to share proprietary assets.

8. Bronner v. Mediaprint

The Court of Justice of the European Union established a stringent framework for refusal-to-supply claims involving potentially indispensable infrastructure.

Relevance

A knowledge graph seeking essential-facility treatment would need to demonstrate circumstances approaching genuine indispensability.

This prevents competition law from requiring dominant firms to provide access merely because their information resource is commercially valuable.

Principle: Indispensability and elimination of effective competition are central to exceptional refusal-to-supply cases.

9. IMS Health GmbH & Co. OHG v NDC Health

IMS Health is especially relevant to data infrastructure.

The dispute concerned access to a pharmaceutical data structure protected through intellectual-property rights.

The CJEU recognized circumstances in which refusal to license an intellectual-property asset could constitute abuse of dominance.

Relevance

The case is highly relevant to proprietary:

  • ontologies;
  • classification systems;
  • identifiers;
  • databases;
  • knowledge structures.

Principle: Intellectual-property protection does not necessarily immunize a dominant infrastructure from competition law where exceptional refusal-to-license conditions are satisfied.

10. Magill

The Magill cases established the foundational European doctrine concerning exceptional compulsory licensing of intellectual property.

Relevance

Where a proprietary knowledge system becomes indispensable for a downstream market, competition law may examine whether withholding access prevents the emergence of new products or services.

This is particularly significant for AI systems that require access to structured information.

Principle: Exceptional circumstances can justify intervention where refusal to license prevents downstream innovation and competition.

13. Comparative Case-Law Lessons

CaseCore issueKnowledge-infrastructure relevance
Google Search (Shopping)Self-preferencingPreferential treatment within discovery infrastructure
Google AndroidEcosystem restrictionsLeveraging across technological layers
MicrosoftPlatform foreclosureInfrastructure-based network effects
Aspen SkiingRefusal to cooperateStrategic withdrawal of access
TrinkoRefusal to dealLimits of compulsory access
BronnerEssential facilitiesIndispensability
IMS HealthData/IP accessProprietary information structures
MagillCompulsory licensingDownstream innovation

14. Global Competition-Law Framework

The issue can be analysed under several legal systems.

European Union

Key tools include:

  • Article 101 TFEU;
  • Article 102 TFEU;
  • EU merger control;
  • Digital Markets Act.

Article 102 is particularly relevant to:

  • refusal to supply;
  • tying;
  • discrimination;
  • exclusionary conduct;
  • self-preferencing.

The DMA adds a more regulatory approach to major digital gatekeepers.

United States

The principal framework includes:

  • Sherman Act §1;
  • Sherman Act §2;
  • Clayton Act;
  • merger enforcement.

U.S. doctrine generally requires careful proof of competitive harm and is cautious about imposing mandatory access obligations.

United Kingdom

The Competition Act 1998 and UK merger-control regime provide tools for examining:

  • abuse of dominance;
  • exclusionary conduct;
  • digital platform concentration;
  • strategic acquisitions.

The UK's digital-markets regime also provides a more ex ante approach for firms designated with substantial and entrenched market power.

India

The Competition Act 2002 is relevant to:

  • abuse of dominant position;
  • denial of market access;
  • discriminatory conditions;
  • leveraging;
  • tying/bundling;
  • combinations.

Knowledge infrastructure can therefore become relevant where control over data or digital gateways creates an appreciable competitive advantage.

15. Merger-Control Risks

Knowledge-graph concentration may arise through acquisitions rather than conduct.

A large platform could acquire:

niche database + scientific repository + mapping dataset + identity-resolution company + AI retrieval provider.

Individually, each acquisition might appear modest.

Collectively, however, they could create a global epistemic infrastructure conglomerate.

Competition authorities should therefore consider:

  • data complementarities;
  • future competition;
  • innovation competition;
  • interoperability;
  • access foreclosure;
  • ecosystem effects;
  • acquisition of nascent competitors;
  • control of unique datasets.

Traditional turnover thresholds may be insufficient where strategically important data assets generate little current revenue.

16. Data Portability As A Remedy

One potential remedy is enhanced portability.

Users or businesses could move:

  • entity identifiers;
  • metadata;
  • records;
  • relationships;
  • annotations;
  • historical information

between systems.

However, portability alone may not solve the problem where the dominant firm possesses unique proprietary data that cannot be replicated.

17. Interoperability Remedies

A stronger remedy may require technical interoperability.

Possible obligations include:

  • open APIs;
  • standardized identifiers;
  • common metadata formats;
  • machine-readable exports;
  • interoperability protocols;
  • non-discriminatory access;
  • provenance standards.

The objective is not necessarily to make every proprietary graph completely open.

Instead, the objective is to prevent proprietary architecture from becoming an artificial barrier to competition.

18. Data Access Remedies

Competition authorities may consider:

FRAND-style access

Access could be provided on:

  • fair;
  • reasonable;
  • non-discriminatory

terms.

Data trusts

Sensitive or commercially important datasets could be administered by independent intermediaries.

Clean rooms

Competitors could access specific information without obtaining commercially sensitive underlying datasets.

API access

Controlled technical access may allow downstream innovation without transferring the entire database.

19. Epistemic Infrastructure And Consumer Welfare

Consumer harm may occur even where prices remain zero.

Potential harms include:

  • reduced informational diversity;
  • lower search quality;
  • reduced innovation;
  • inaccurate information;
  • manipulation of visibility;
  • reduced choice;
  • discriminatory ranking;
  • weaker privacy;
  • reduced ability to challenge dominant information sources.

Thus:

consumer welfare ≠ price alone.

For knowledge infrastructure, quality, diversity, reliability and contestability become economically significant.

20. Innovation Competition

A concentrated knowledge infrastructure can suppress innovation indirectly.

A start-up may be unable to compete because it cannot obtain:

  • sufficiently comprehensive data;
  • standardized identifiers;
  • reliable entity resolution;
  • API access;
  • historical information;
  • source metadata.

The incumbent may therefore not need to copy the entrant's technology.

It can simply prevent the entrant from obtaining the informational inputs necessary to scale.

21. Epistemic Monoculture

The most serious long-term concern is epistemic monoculture.

If many downstream systems rely upon the same knowledge source:

one graph → many AI systems → many applications → many decisions.

An error or bias in the upstream graph can consequently propagate through multiple markets.

From a competition perspective, this creates concerns about:

  • systemic dependency;
  • common-input concentration;
  • correlated errors;
  • innovation suppression;
  • loss of informational diversity.

22. Relationship With AI Governance

Knowledge graphs may become part of the infrastructure used by:

  • AI agents;
  • autonomous economic systems;
  • regulatory technology;
  • financial AI;
  • healthcare AI;
  • scientific AI;
  • procurement systems.

The entity controlling the knowledge layer could therefore influence the information available to autonomous systems.

This creates a novel competition concern:

control over the information environment of competing AI systems.

23. Possible Competition Theories

Future enforcement could develop around several theories.

Theory 1 — Knowledge Infrastructure Monopoly

A firm controls an indispensable structured information layer.

Theory 2 — Epistemic Gatekeeping

The firm controls which information becomes discoverable.

Theory 3 — Data Leveraging

The firm transfers upstream informational advantages into downstream markets.

Theory 4 — Graph-Based Self-Preferencing

The firm uses its graph to favour affiliated products.

Theory 5 — Interoperability Foreclosure

The firm prevents competing systems from communicating effectively.

Theory 6 — AI Input Foreclosure

The firm restricts access to information required for competing AI models.

Theory 7 — Epistemic Merger Concentration

Acquisitions combine complementary information assets into an infrastructure monopoly.

24. Regulatory Challenges

Competition authorities face several difficulties.

Defining the relevant market

Knowledge infrastructure does not fit neatly into conventional product markets.

Measuring market power

Traditional market shares may underestimate:

  • data advantages;
  • API dependency;
  • switching costs;
  • network effects.

Separating accuracy from competition

A dominant system may legitimately rank information according to quality.

Competition law must distinguish legitimate quality improvement from exclusionary ranking.

Protecting innovation incentives

Mandatory access can weaken incentives to invest in proprietary data systems.

Privacy conflicts

Data-access remedies must comply with privacy and data-protection requirements.

International coordination

Knowledge graphs operate globally while competition authorities remain largely jurisdictional.

25. Future Direction Of Competition Law

The evolution is likely to move from:

market concentration

towards:

infrastructure concentration

and ultimately toward:

control over critical informational ecosystems.

Competition authorities may increasingly examine not merely whether one firm sells the dominant product, but whether it controls the information architecture upon which competing markets depend.

Conclusion

Global knowledge graphs represent a potentially fundamental layer of the digital economy. Their competitive importance derives from their ability to connect vast quantities of information and make that information usable by search engines, businesses, governments and AI systems.

The central competition-law risk is therefore not simply a conventional monopoly over a database.

It is concentration of epistemic infrastructure.

Where one undertaking controls data acquisition, entity resolution, graph construction, ranking, APIs and downstream applications, it can potentially influence both economic competition and the informational conditions under which competition occurs.

The principal legal lessons from Google Shopping, Google Android, Microsoft, Aspen Skiing, Trinko, Bronner, IMS Health and Magill are that competition law already contains many of the conceptual tools required to address this problem—self-preferencing, leveraging, refusal to deal, essential facilities, interoperability and exceptional access to proprietary information.

The emerging challenge is to adapt those doctrines to an environment in which knowledge itself becomes infrastructure. The decisive question for future antitrust enforcement will increasingly be:

LEAVE A COMMENT