Training, Inference, and the Geography of AI Compute: Why Workload Type Will Determine Where the Next Data Centers Get Built
AI infrastructure is often discussed as one enormous category of GPU demand. But training and inference place different demands on power, networks and geography. Frontier training rewards enormous, tightly connected compute systems. Inference adds another constraint: the user. As AI shifts from building models to serving them at global scale, workload type will increasingly influence where infrastructure belongs.
Training rewards concentration. Thousands of accelerators working synchronously need enormous power, scalable cooling, contiguous sites and dense GPU-to-GPU fabrics. During model development, proximity to the eventual user is usually secondary.
Inference creates a different constraint. Every request travels to the compute, the model processes it and the output travels back. Inference creates a latency budget.
Workload architecture
Similar accelerators, different centers of gravity
Frontier training
Concentrate compute
- Model development
- Highly coordinated clusters
- GPU-to-GPU fabric
- Power + scale
Production inference
Serve compute
- Respond to demand
- Request-driven systems
- User + service latency
- Power + demand
Training has become data-center scale
Meta described moving from 128-GPU training jobs to multi-thousand-GPU systems, later assembling a 129,000-H100 cluster and developing the one-gigawatt Prometheus cluster across multiple buildings.
4K
GPUs per early LLM job
24K
H100s in each 2023 cluster
129K
H100 cluster
1 GW
Prometheus
Inference geography is measurable in milliseconds
In AWS testing with identical model configurations, moving inference closer to users materially reduced mean time to first token.
135 ms
Los Angeles user / Oregon
80 ms
Los Angeles Local Zone
197 ms
Honolulu user / Oregon
114 ms
Honolulu Local Zone
Interactive analysis
What determines where compute belongs?
Placement variable
Power Availability
Large-scale AI infrastructure must locate where sufficient electricity can be delivered. Training often has greater freedom to move toward power; inference balances energy with proximity and demand.
The geography may become a portfolio
Gigawatt campus
Frontier training
Regional AI hub
Shared training and inference
Metro compute
Latency-sensitive serving
Enterprise edge
Private data and systems
Network edge
Telecom-integrated inference
Device
Local AI without a round trip
Bottom line
The industry is unlikely to converge on one universal AI campus. Training follows the compute. Inference increasingly follows demand, data and latency. The AI infrastructure map will need both.
Verified sources
- 01Meta Engineering — Meta's Infrastructure Evolution and the Advent of AI
- 02Meta Engineering — Building Prometheus
- 03AWS — Reduce Conversational AI Response Time Through Inference at the Edge
- 04NVIDIA — AI Inference Cluster Infrastructure Economics
- 05NVIDIA — Telecom Leaders Build AI Grids to Optimize Inference
- 06AWS Documentation — Local Zones
All factual references used in this article were publicly available on or before August 24, 2026.
Sean Kurz
Expert insights from the Nistar team on energy infrastructure and hyperscale development.