Edition 21 asked what happens when AI starts to change what "location" even means. This is that edition. Inference doesn't migrate to one new geography — different workloads have very different answers to where they're actually allowed, or economically sensible, to run.
AI infrastructure has spent the past several years chasing power. Large training campuses made the logic easy to understand: secure enormous power blocks, concentrate accelerators, build dense networking, optimise the factory around throughput. Inference complicates that geography — not because inference simply "moves to the edge." It doesn't. The more important change is that different inference workloads have very different answers to a more basic question: where are they actually allowed — or economically sensible — to run?
1 The deal
That distinction matters to an investor because the cheapest place to build compute is not necessarily a place where every workload can actually run. Each workload has its own placement envelope — the set of locations that can satisfy its technical, regulatory and economic requirements. The wider that envelope, the more locations can substitute for one another. The narrower it becomes, the scarcer an acceptable location can become. A location premium only has a chance to exist where that scarcity starts to bite — but scarcity alone does not guarantee an owner captures it, a distinction this edition returns to throughout.
Geographic processing now has a printed price — but it's a usage fee, not a rent. OpenAI and Mistral each explicitly charge a 10% uplift for eligible regional-processing / regional-inference endpoints (both vendor documentation, retrieved August 2026). That does not establish a 10% data-centre location premium. The surcharge accrues at the model-service layer. The infrastructure question is how much, if any, survives into facility utilisation, contract durability or rent — and that is the question this edition works through.
It matters because inference is no longer the smaller half of AI compute. McKinsey models inference-related data-centre demand rising from roughly 21 GW in 2025 to 93 GW by 2030, against training demand rising from about 23 GW to 62 GW over the same period — though McKinsey's own modelling assumes AI data centres can serve both workloads, so this is not a clean physical split. Inference is projected to overtake training during the forecast period and non-AI workloads by 2029.
2 The engineering read
First, separate two different latencies
One reason the inference debate gets confused is that "latency" is treated as a single thing. It isn't.
Intra-cluster latency: inside an AI factory, accelerators exchange enormous volumes of data at extremely low latency — high-bandwidth, low-latency GPU-to-GPU communication is central to how modern training and inference workloads (including mixture-of-experts models) actually run. Training should not be described as "latency insensitive." It may be relatively insensitive to its distance from the eventual end user while being extraordinarily sensitive to communication inside the compute cluster.
Location latency: a different question — the time required for inference infrastructure to communicate with the user, enterprise application, private data source, or another system the model repeatedly calls. For interactive AI, that can affect user experience. For machine-to-machine or time-critical applications, the constraint can instead be operational: the application itself may have a tight latency budget even when no human is waiting for the response. AWS's Amazon EC2 G7e instances — accelerated by NVIDIA RTX PRO 6000 Blackwell GPUs — became available in Asia Pacific (Tokyo) in February 2026 and Asia Pacific (Seoul) in March 2026. In July 2026, AWS extended G7e support to SageMaker AI inference in both markets, explicitly framing the expansion as allowing inference endpoints to sit closer to Asian end users and reduce generative-AI latency. That is evidence proximity can matter. It is not evidence all inference needs to be local.
Inference is not one workload. Four different workload patterns can produce very different placement envelopes:
Real-time and time-critical inference
This covers two related but distinct cases. Human-facing: voice, video, interactive generative AI, customer-facing applications, where response time affects user experience directly. Machine-facing: applications or systems with tight operational response-time requirements, where location latency can matter even without a human directly waiting for the response. Tokyo or Seoul inference capacity can have a real advantage here even where cheaper accelerator capacity exists farther away.
Sovereign or regulated inference
The relevant boundary isn't milliseconds, it's a national border. Unlike latency or data gravity, this is fundamentally an eligibility constraint: even if another region is cheaper, faster or has more available compute, it remains outside the placement envelope if the workload cannot legally or contractually be processed there. AWS introduced India Geographic cross-region inference in August 2026, initially for OpenAI models on Bedrock, letting customers with in-country processing requirements route requests between Mumbai and Hyderabad while inference stays inside India. Sovereignty does not necessarily eliminate pooling — it changes the boundary within which pooling can occur. The workload is geographically constrained but not necessarily metro-constrained — sovereignty creates a placement envelope without creating an edge requirement. Japan's Digital Agency is running a parallel version: a fiscal-2026 pilot testing three domestically developed foundation models on Sakura Cloud — the Government Cloud's only domestically developed cloud service — with blind evaluation trials running September–November 2026 alongside existing access to Amazon Nova and Anthropic Claude models. The constraint in both cases is eligibility, not latency.
Enterprise inference — data gravity
A different kind of gravity, created by repeated interaction with private databases, applications, security controls and other enterprise systems. Here, the placement constraint is not necessarily end-user latency. Moving inference farther away can increase network dependency, data movement, response time and architectural complexity when the model repeatedly needs to retrieve, process or act on enterprise data. In those cases, compute may be pulled toward the data and systems it depends on — rather than toward the end user.
What Equinix's Inference Exchange is — and isn't
Equinix announced its Inference Exchange on 2 September 2026, combining NVIDIA Enterprise Reference Architectures with Together AI's inference platform to connect distributed inference infrastructure to enterprise data, clouds and networks — but the service is announced for availability from Q1 2027, with no pricing, capacity commitments, or anchor customers yet disclosed.
It's useful evidence of where the industry is investing. It is not yet proof that enterprises will pay data-centre owners a measurable premium for it.
Latency-tolerant inference
Batch workloads, asynchronous processing, background AI tasks. These have much weaker reasons to sit close to the user. The priority instead becomes GPU utilisation, power cost, available capacity, and cost per token. Proximity has little economic value where latency is non-binding, which can pull compute away from expensive metros rather than toward them.
The counterforce: not all inference demand needs local inference capacity
Remote inference itself is not new. A Singapore application could already call compute hosted in another region, and companies could build applications across multiple regions themselves. What is becoming more explicit is who manages the placement. Instead of a customer having to decide where inference capacity is deployed and manually select which region should serve an eligible workload, that placement can increasingly be managed automatically at the managed-model service layer. The service can dynamically draw on model capacity across a broader pool of supported regions.
AWS's Global cross-Region inference makes this visible. Customers in Thailand, Malaysia, Singapore, Indonesia and Taiwan can invoke supported Claude models while AWS's managed inference service dynamically routes eligible requests across 20+ supported commercial regions worldwide. The customer does not have to manually select the serving region for each eligible request. That does not create remote inference. It makes cross-region placement increasingly productised and automatically managed at the model-service layer.
The infrastructure implication is important. If a workload does not have to stay in Singapore, for example, Singapore demand does not automatically require Singapore inference capacity. In other words: demand originating in a market ≠ compute that must be hosted in that market. For the data-centre investor, this changes the scarcity equation. Cloud pooling expands the set of locations that can serve a workload. Latency, sovereignty, data gravity and other constraints shrink it.
This is where a simple "inference moves to the edge" thesis breaks. Inference carries two opposing forces at once: cloud pooling can make capacity across regions increasingly interchangeable for workloads that do not need to stay local, while proximity, sovereignty and data gravity can pin other workloads to a much narrower set of locations. The investment question is therefore not whether inference is becoming local or global. It is which workloads can be pooled — and which ones cannot.
The seam worth marking. A geographic constraint can narrow the eligible pool without eliminating pooling inside that boundary. The more important investment question is whether that constraint creates value the facility owner actually captures, rather than value retained by the hyperscaler, network provider or GPU cloud.
3 The capital allocation read
OpenAI and Mistral both publish a 10% uplift for eligible geographically constrained processing/inference offerings — the market pricing "here, not there" explicitly. (Azure OpenAI and Google Vertex AI also offer region-scoped deployments; a consistent, directly comparable price differential was not confirmed against their own documentation, so no claim is made about them either way.) That surcharge is charged per API call to the model vendor. It is not, today, an offtake payment to the facility hosting the endpoint. A usage fee only becomes an owner's return if it survives the path into facility economics. Otherwise, the premium belongs to another layer of the stack — not the data-centre investor. The underwriting question is never "this facility is close to users." It is: what scarce function does the facility control that the customer cannot cheaply reproduce elsewhere?
The width of that placement envelope determines how many viable substitutes a workload has. A workload with twenty acceptable locations gives the customer bargaining power. A workload with one is where location becomes strategic. Conceptually, location scarcity rises as the number of viable substitutes falls. Managed cross-region inference can expand that placement envelope for eligible workloads because the serving location can increasingly be selected dynamically at the managed-model service layer rather than manually tied to the market where demand originates. But that expansion stops wherever a workload constraint binds — latency may narrow the pool, sovereignty may confine it to one country, enterprise data gravity may pull it toward a particular ecosystem, service availability or application architecture can narrow it again. So the relevant question is not how many cloud regions exist. It is how many of them are actually viable substitutes for this particular workload.
Malaysia is increasingly part of the same cloud architecture that makes Singapore valuable — AWS operates separate Singapore and Malaysia regions, and Microsoft's Malaysia West (Kuala Lumpur) region has been generally available since May 2025, with a second Malaysian region in Johor Bahru ("Southeast Asia 3") announced by Microsoft but not yet given a confirmed launch date. AWS already lets eligible inference invoked from both Singapore and Malaysia draw on globally distributed model capacity through its managed cross-region inference architecture. That makes "can inference move from Singapore to Johor?" the wrong question. The better one: which workloads still require characteristics that Singapore provides once Malaysia itself becomes a genuine cloud and AI region? If a workload's latency, residency and connectivity requirements are satisfiable in Malaysia, the two markets become more substitutable for that workload. If not, Singapore retains scarcity. That is a different investment proposition from simply "Singapore is closer to demand" — and this edition is not claiming a Singapore-versus-Johor IRR ranking either way.
A latency-driven placement requirement can erode as networks improve and as on-device silicon absorbs the short-call tier. A legal in-country requirement doesn't erode simply because fibre improves — it persists until the regulatory or compliance boundary changes. That makes sovereignty a potentially more persistent constraint on where a workload can run. Persistence of the constraint is not the same as proof of an owner-side premium. Owner capture is a function of location scarcity, infrastructure control, and contractual retention together; if another layer of the stack holds the control point, the facility owner may see very little of it regardless of how scarce the location is.
Placement-envelope underwriting pass
- What is this workload's placement envelope — how many locations satisfy its latency, residency, connectivity, reliability and economic requirements?
- Is scarcity driven by latency, sovereignty, or data gravity — or some combination?
- Does pooling economics outweigh locality for this workload — can a managed service satisfy it from elsewhere in an eligible pool?
- What scarce function does the facility actually control that the customer cannot cheaply reproduce elsewhere?
- Does that control survive into rent, utilisation or contractual durability — or does another layer of the stack capture it first?
Investment Lens
4 What this means for the broader market
The AI infrastructure build-out began with a power question: where can we build the next large block of compute? Inference adds another one: where can this particular workload actually run? Those are not always the same place. Some inference will remain geographically flexible, with managed-model services able to route eligible requests automatically across broader regional capacity pools. Some will remain inside national borders. Some will cluster around enterprises and cloud ecosystems. Some will genuinely move closer to users or time-critical systems. Location value therefore comes from disappearing substitutes, not from proximity alone — and value only reaches the facility owner if the owner, not another layer of the stack, controls the bottleneck.
On-device inference is another live counterforce: as capable local silicon absorbs some latency-sensitive workloads, the pool of workloads that actually requires metro data-centre inference could shrink.
The investment opportunity, then, is not simply to own "edge" capacity. It is to identify where a growing workload faces a shrinking set of acceptable locations — and then determine whether the infrastructure owner actually controls that bottleneck.
The most valuable megawatt may not be the closest one. It may be the least substitutable one. And that raises the next question — if a durable constraint doesn't automatically hand its premium to whoever sits inside the border, is the winning asset a data centre at all, or the control point wrapped around it?
That is the question I'm taking forward next.