Abstract
This paper develops the framework for sizing the uninterruptible power supply runtime in the modern AI-factory data center, with focus on the structured decision logic that determines the autonomy window the operator commits the campus to operate against. The framework treats the runtime decision as a five-dimensional engineering object: the workload-class dimension that specifies the failure-tolerance envelope of the running compute, the architecture-class dimension that specifies the power-chain redundancy upstream of the UPS, the operating-model dimension that specifies the failover-and-restoration machinery the operator runs, the capital-allocation dimension that specifies the cost stack the runtime decision drives, and the lifecycle dimension that specifies the battery-chemistry maintenance and replacement schedule across the multi-decade lifecycle.
The analysis specifically treats the contemporary industry shift from the legacy thirty-minute UPS runtime convention toward the modern five-to-ten-minute runtime envelope, with explicit attention to the workload-class and architecture-class conditions under which the shift is operationally bounded versus operationally unbounded. The legacy thirty-minute convention was designed against deployment contexts in which generator startup time was long, in which workload restart cost was high, and in which battery chemistry favored longer autonomy windows. The modern deployment context does not satisfy any of those conditions in the same way, and the convention has begun to drift toward the shorter envelope across multiple major operators.
The framework is grounded in the documents that actually govern stored-energy autonomy and standby power: IEEE 485 for lead-acid sizing, IEEE 1184 for UPS battery selection, IEEE 1188 for maintenance, testing, and replacement of valve-regulated cells, IEC 62040-3 for UPS performance classification, NFPA 110 for emergency and standby power systems, NFPA 111 for stored electrical energy systems, the National Electrical Code, ANSI/TIA-942-C, ANSI/BICSI 002, and the Uptime Institute Tier framework, together with the body of practitioner experience the author has accumulated across hyperscale UPS deployments. The framework is workload-aware to the extent that an AI-training workload, an AI-inference workload, and a mixed enterprise workload each carry distinct autonomy-window requirements and distinct failover-and-restoration requirements.
A theme running through the analysis is that the nominal autonomy figure is a poor proxy for the quantity that governs survival. What governs survival is the transfer gap, the interval between loss of the utility source and the restoration of a stable alternate source, decomposed into detection and start, crank and stabilize, and transfer and synchronize. Stored energy has to cover that interval with margin. Sizing to a convention rather than to a measured transfer gap produces autonomy that is either wasted or insufficient, and the direction of the error is not knowable from the specification sheet.
Recommendations identify the decision logic, the operating-model integration, the capital-allocation discipline, and the lifecycle treatment required to operate the UPS runtime as a sustained engineering commitment. The paper is intended for executives, engineering organizations, capital sponsors, and operating teams whose UPS-runtime decisions over the next two capital cycles will determine whether the campus’s autonomy window remains operationally bounded or perpetually re-litigated.
Executive Summary
The autonomy window of an uninterruptible power supply is one of the few numbers in a data center design that is simultaneously an engineering parameter, a capital commitment, a fire protection input, a floor-loading input, and a contractual promise. It is usually decided the way conventions are decided, by inheritance. This paper argues for deciding it the way engineering decisions are decided, against a stated failure model, and it sets out the framework for doing so.
Finding one. The thirty-minute convention is an artifact of a deployment context that no longer exists. It was set when generator start and transfer sequences were slower, when restarting a workload was expensive and manual, and when valve-regulated lead-acid chemistry made long autonomy the cheapest way to buy confidence. Each of those three conditions has moved. Carrying the convention forward without re-deriving it is an unexamined inheritance rather than a design position.
Finding two. The binding constraint on short autonomy is the standby generation plant, not the battery. Published reliability work on emergency and standby diesel generators shows fail-to-start and fail-to-run probabilities that are material at the single-unit level and that compound across the duration of an outage. A five-minute autonomy behind a plant with an unexamined start reliability is a different risk position from the same five minutes behind an N+2 plant with a disciplined test regime, and the specification sheet does not distinguish them.
Finding three. The quantity that matters is the transfer gap, not the nameplate minutes. Decomposed into detection and start, crank and stabilize, and transfer and synchronize, the transfer gap is measurable at commissioning and drifts over the life of the plant. Autonomy is the margin held over that measured interval. An operator who has never measured it is holding an unknown margin, whatever the battery is rated for.
Finding four. Electrical ride-through is frequently the shorter of the two ride-through windows that matter. At the densities now being deployed, the cooling plant loses its own ride-through on a timescale that can be shorter than the electrical one, and useful compute ends when the thermal envelope is breached rather than when the battery is exhausted. Sizing the electrical autonomy without sizing the thermal autonomy alongside it produces a number that cannot be delivered.
Finding five. Autonomy is a capital allocation, and it competes with the other things that buy reliability. Battery capacity, generator redundancy, transfer automation, and workload-level resilience all draw on the same reliability budget. Extending autonomy consumes capital, floor area, structural capacity, and fire protection scope that could have bought redundancy or automation instead. The allocation is a governance decision and belongs in front of the people who own the capital plan.
Recommendation one. Specify autonomy against workload class and architecture class, and record the derivation. The paper’s decision framework takes the governing inputs, sets a segment baseline, applies modifiers, and produces a defensible figure with the reasoning attached. A number without its derivation cannot be defended at review, cannot be audited, and cannot be revised intelligently when the workload changes.
Recommendation two. Prove the number at commissioning and keep proving it. Design autonomy becomes demonstrated autonomy only through capacity testing, transfer testing, and a maintenance regime aligned to the maintenance and testing practice for the installed chemistry. Telemetry at the cell level, capacity-test cadence, and a documented end-of-life criterion convert the specification into an operating fact.
Recommendation three. Write autonomy into procurement and service-level language with the definition attached. A service level that promises minutes without defining the load, the state of charge, the end-of-discharge voltage, the ambient temperature, and the age of the cells promises nothing enforceable. Factory and site acceptance obligations, capacity-test terms, warranty structure, and remedy provisions should all resolve to the same definition the engineering used.
Recommendation four. Govern the decision through refresh. Battery service life is shorter than building life, so autonomy is re-decided every refresh cycle whether or not anyone convenes to decide it. Naming the owner of that decision, the review trigger, and the documentation of record is what keeps the campus from drifting back to convention by default.
The paper is intended for executives, engineering organizations, capital sponsors, operating teams, and regulators whose decisions over the next two capital cycles will determine whether the autonomy question remains operationally bounded or perpetually contested.
Full white paper below

