AI Factories and the Token Economy
Architecture, Economics, and Capital Strategy for the Industrialization of Intelligence
Abstract
This publication offers an integrated technical and economic foundation for the artificial intelligence factory, defined as the industrial-scale infrastructure that converts electrical energy into measurable cognitive output denominated in tokens. The token is the unit of account that closes the loop from capital deployment to commercial output. It is measurable, denominated identically across providers within tolerable variance, billable in standardized contractual form, hedgeable through reserved capacity and forward markets, and increasingly tradeable across an emerging brokerage layer. The publication treats the token economy not as a marketing construct but as the operational reality against which capital, engineering, and governance decisions must be tested.
The analytical scope extends from substation to silicon and from capital authorization to operational retirement. Power architecture coverage includes eight-hundred-volt direct current distribution, solid-state transformer integration, and the cumulative loss budget that governs power usage effectiveness and tokens per kilowatt-hour at the factory level. Cooling coverage extends from rear-door heat exchangers through direct-to-chip cold plates to single-phase and two-phase immersion, with explicit treatment of constructability, maintainability, and refresh implications. Inference stack coverage addresses scheduler maturity, key-value cache management, prompt caching, quantization, speculative decoding, and the brokerage layer that increasingly mediates between customer applications and the underlying compute substrate. Capital analysis treats time to power, refresh cadence, depreciation discipline, and the build-versus-buy decision threshold as joint optimization problems rather than as independent line items.
Principal findings include the following. Token economics are dominated by the asymmetry between input and output tokens, with output tokens consuming approximately four times the marginal cost of input tokens at frontier scale and reasoning tokens further amplifying that cost. Tokens per kilowatt-hour have improved approximately thirteen-fold across the five silicon generations from A100 through GB300, with the trajectory continuing to compound at approximately 2.3x per year through 2030. Capital intensity at gigawatt scale is dominated by silicon, with approximately fifty-seven percent of the capital stack concentrated in accelerators and host servers. Operating economics are dominated by power cost in the operating budget and by silicon depreciation in the capital charge, with both lines amenable to architectural discipline. The build-versus-buy threshold sits at approximately ten billion tokens of sustained monthly demand for generic workloads and lower for strategic workloads where workload differentiation drives capability requirements.
Principal recommendations include adopting tokens per watt and cost per token as primary operating metrics, planning power architecture two generations ahead with explicit refresh provisions, standardizing on direct-to-chip liquid cooling for frontier silicon while preserving optionality for two-phase immersion, running a disciplined silicon portfolio with thirty-six-month refresh cadence, investing in inference stack maturity, pricing across at least three service tiers, compressing time to power through long-lead substation orders and modular construction, and establishing architecture governance with explicit decision rights authority. The recommendations are operationally tied; partial adoption produces partial outcomes and partial outcomes do not converge.
The intended audience is executive infrastructure leadership, finance and investment principals, engineering and operations executives, governance and regulatory practitioners, and policy advisors with infrastructure exposure. The geographic scope is global with explicit treatment of North American and European regulatory frameworks; the temporal scope is the 2026-2030 industry trajectory. The analytical posture is independent and evidence-based. The publication treats engineering decisions as capital and governance decisions and treats artificial intelligence factory architecture as the structural commitment it has become.
Executive Summary
Artificial intelligence has crossed from research curiosity to industrial output. The unit of that output is the token. Tokens are now produced, priced, traded, hedged, and consumed at industrial scale, and the infrastructure that produces them is correctly named an artificial intelligence factory. A factory of this kind is measured not by square feet of white space or by megawatts of information technology load alone, but by the rate at which it converts electrons into useful tokens, the unit cost of those tokens, and the capital efficiency with which the conversion is achieved. Operators who have not made this transition are still running data centers. Operators who have made it are running factories.
This publication makes five arguments. First, the token is a measurable and tradeable unit of cognitive output, comparable in economic function to a barrel of oil or a kilowatt-hour of electricity, with the necessary properties of measurability, denomination, billability, and hedgeability for a unit of account. Second, the economics of artificial intelligence factories are dominated by tokens per watt and cost per token, both of which respond directly to architectural and operational choices that executives have within their control.
Third, the architecture that produces competitive tokens at gigawatt scale requires eight-hundred-volt direct current distribution, solid-state transformer integration, direct-to-chip liquid cooling, mature inference stacks, and disciplined silicon refresh, none of which is optional at the frontier. Fourth, capital deployment velocity is the operational constraint that separates leaders from followers, with time to power on the order of thirty to forty-two months as the dominant gating function. Fifth, the operators who win the next decade will not be those with the most square feet or the most megawatts, but those who industrialize the conversion most rigorously.
Principal findings from the analytical work include the following. Output tokens cost approximately four times input tokens at frontier scale, and reasoning tokens, when uncapped, can multiply the per-query cost by eight times or more. Tokens per kilowatt-hour have improved approximately thirteen-fold across five silicon generations from A100 through GB300, with the efficiency frontier compounding at approximately 2.3x per year through 2030. Capital intensity at gigawatt scale runs approximately thirty-eight billion United States dollars per gigawatt, with silicon comprising approximately fifty-seven percent of the stack. Operating margin sensitivity to industrial power cost is approximately eighty-eight million United States dollars per cent per kilowatt-hour at gigawatt scale, making siting decisions one of the dominant determinants of multi-year profitability. The build-versus-buy decision threshold sits at approximately ten billion tokens of sustained monthly demand for generic workloads, and substantially lower where workload differentiation, data residency, or strategic dependency considerations apply.
Principal recommendations include the following. First, adopt tokens per watt and cost per token as primary operating metrics, displacing legacy metrics that no longer reflect the production reality of the artificial intelligence factory. Second, plan the power architecture for the next two silicon generations rather than the current one, with explicit refresh-resilient power and cooling design. Third, standardize on direct-to-chip liquid cooling for frontier silicon, while preserving deployment optionality for two-phase immersion at the highest densities. Fourth, run a disciplined silicon portfolio with a thirty-six-month rolling refresh cadence funded by an explicit refresh reserve.
Fifth, invest in the inference stack , scheduler, cache, quantization, decoding, as if it were a hardware platform. Sixth, price across at least three service tiers and operate at premium, on-demand, and batch posture concurrently to flatten diurnal demand. Seventh, compress time to power through long-lead substation orders at site selection, modular construction, and parallelized commissioning workflows. Eighth, establish architecture governance with explicit decision-rights authority, ownership clarity, and quarterly review cadence at the executive committee level. Ninth, build the workforce capability pipeline as a deliberate certification program rather than as a hiring function. Tenth, report on the right page, tokens, watts, cost, capital, rather than on legacy metrics that obscure the underlying economics.
Forward outlook through 2030 anticipates continued silicon efficiency gains, mass adoption of eight-hundred-volt direct current distribution, two-phase immersion at the highest densities, mixture-of-experts and hybrid model architectures, sovereign artificial intelligence inflection in 2028 and beyond, edge inflection across consumer and industrial applications, and continued pricing compression with achieved blended price moving from approximately one dollar fifty cents per million tokens in 2026 to approximately sixty-five cents in 2030 under base-case forecasts. The operators who win this trajectory will defend revenue through workload mix and service-tier discipline rather than through pricing power, will absorb depreciation through refresh discipline rather than through deferral, and will maintain capital velocity through governance maturity rather than through fire-fighting.
The intended audience is executive infrastructure leadership, finance and investment principals, engineering and operations executives, governance and regulatory practitioners, and policy advisors. The publication is organized for that audience: a definitional foundation, a worked examination of the token economy across canonical use cases, an architectural reference from substation to silicon, a quantitative treatment of capital and operations economics, a strategic recommendations chapter, a vendor and ecosystem environment, reference architectures, an operations playbook, a forward roadmap, an implementation framework, vendor qualification, security architecture, sustainability considerations, talent strategy, customer success, and case studies grounded in the four industry archetypes that dominate the 2026 environment. The throughline is unchanged from the first edition. The operators who win the next decade will be those who industrialize the conversion most rigorously.
Full paper below.

