The Building the Internet Lives In
A data center works by gathering thousands of computers under one roof, feeding them electricity that never stops, pulling the heat they produce out of the air, and wiring them to the internet so that distant devices can borrow their work. Looked at closely, the building is a power and cooling plant that happens to perform computation: the job of removing heat and the guarantee of unbroken electricity decide where the facility is built, how its halls are arranged, and what it costs to run, before anyone asks about processors, memory, or software. The racks are the visible part; the substations, switchgear, battery rooms, diesel generators, chilled-water loops, and fiber entrances are the reason those racks can keep running at all.

The metaphor of a building holding the internet is less fanciful than it sounds. Every message opened, every video streamed, every card payment authorized sends electrical signals to a concrete structure somewhere, where machines execute the request and return an answer in a fraction of a second. The work cannot happen entirely inside the device in a person’s hand: a phone carries a small battery, has no air moving across its parts, and has nobody standing watch over it. Concentrating the machines in one place solves that problem by centralizing the three things a phone cannot carry: a power supply that does not sleep, a heat-removal system built at industrial scale, and engineers who keep both running. Luiz Barroso and Urs Hölzle gave this idea its canonical name in their 2009 book “The Datacenter as a Computer,” where they described warehouse-scale computing, the treatment of an entire building as one computer; their finding was that at that scale, the economics of electricity and the handling of hardware failure dominate every design decision.
The facilities range from rooms holding a few dozen racks to campuses that occupy hundreds of acres and draw electricity in the hundreds of megawatts. The International Energy Agency estimated that data centers account for about 1 to 1.5 percent of global electricity use, a share that makes the thesis concrete: a structure full of machines whose collective appetite for electricity resembles that of a small industrial region. Yet public conversation about these buildings tends to dwell on the machines, the chips, the brands of processors, while the decisive engineering is the kind visible in a utility yard: transformers humming, cooling towers venting vapor, diesel exhaust stacks waiting for a day that must never come. Ask an operator what keeps them awake and the answer comes back in the two currencies of the thesis, delivered electricity and rejected heat.
Step inside and the first impression is order. Racks stand in long rows, each a steel frame about the width of a refrigerator, holding servers stacked like drawers. Cold air enters the front of every rack; hot air leaves the rear. Rows face each other in repeating pairs so that cold aisles and hot aisles alternate down the hall, a layout that keeps the two air streams from mixing before they have done their work. Nothing about the arrangement is decorative. Every choice, the aisle width, the perforations in the floor tiles where air rises, the direction the doors swing, serves the twin obligations of feeding current to the machines and carrying their heat away. The building earns its keep in the space between those two obligations, and every later section of this article returns to them: the electricity arriving, the electricity surviving failure, the heat leaving, the water spent in its removal, and the wires that let the whole plant serve a world of distant devices.
What a Data Center Actually Is
A data center is a building, or a separately engineered space within one, built for the single purpose of running other people’s computers: it provides them with continuous, conditioned power, engineered heat removal, redundant network connectivity, physical security, and staffed operations. Each clause of that definition does work. Continuous power means the machines do not go dark when the grid blinks. Conditioned power means the voltage arrives clean and steady, since the machines’ supplies tolerate only a narrow band. Engineered heat removal means the air and water systems were sized for the exact thermal output of the installed racks, not borrowed from the building’s general air conditioning. Redundant connectivity means at least two independent paths to the internet, so no single cut cable or failed device isolates the hall. Staffed operations means people walk those aisles, monitor alarms, swap failed parts, and test the backup systems on a schedule. When all of those are present, the room earns the name.
The boundary matters because the name gets used loosely. A locked closet with a few servers under the office stairs is not a data center; it has no second power path, no dedicated cooling plant, and no one watching it at three in the morning. A home laptop is not one, for reasons too obvious to belabor. A telecommunications hut beside a highway, holding radios and a battery string, is not one either: it keeps a signal alive, but it performs no computation for customers and carries none of the redundant systems the definition requires. Even a single rack installed in a company’s back room, running a file server and a printer queue, falls short, since a power outage or a failed air conditioner stops everything, and there is no second anything. The Uptime Institute, the originator of the tier language the industry uses to grade such facilities, draws exactly this distinction in its description of the lowest tier, basic capacity with no redundancy; a room below even that bar has no tier at all.
At the top of the scale sits the hyperscale campus, a cluster of machine halls sharing a substation, a water plant, and a fiber ring, operated by a cloud provider or rented out by a specialist landlord. Between the two ends lies the ordinary enterprise facility: a company’s own machine hall, sized for its own needs, with real redundancy and real staff. Size alone does not settle the question; a small hall with dual power feeds, containment for its hot air, monitored alarms, and tested generators qualifies, while a large room of servers with a single electrical feed and a window air conditioner does not. The line is drawn by engineering intent, and it runs through power, heat, connectivity, security, and watchfulness. Everything else in this article concerns buildings on the right side of that line.
The Problem It Solves: Distance, Scale, and Trust
Three facts about computers force the existence of the data center, and they explain more of its design than any single technology inside it. The first is distance. A device in a person’s pocket cannot do meaningful work against machines halfway across the world without delay, because electrical signals in fiber travel at roughly two hundred kilometers per millisecond, a hard physical ceiling. Computation that must feel instant, a search result, a payment authorization, a video call, has to happen close enough that the round trip fits inside human perception. That alone would justify scattering machines around the planet; the rest of the story is why they gather in buildings instead of living inside the devices themselves.
Ownership adds a further reason the machines gather. A retailer whose peak traffic arrives in one holiday month would have to buy December’s worth of computers and watch them idle through the other eleven; a hospital running quiet research models at night and urgent diagnostics by day faces the same waste. Shared halls let both buy the average and rent the peak, turning fixed capital into a metered service.
The second fact is scale. One machine can serve a few users; millions of users asking at once need millions of times the capacity, and demand spikes without warning, when a product launches, a game goes viral, or a market opens. No single computer holds that much headroom, and buying it in advance would strand most of it idle. The data center answers by pooling: thousands of machines share the incoming load, and capacity is assigned to whichever task needs it in the moment, the same logic that lets an electric grid serve a city without building a power plant for every house. The pooling also works across time, since work that tolerates delay, backups, batch analysis, model training, runs at night or in whichever hall has spare capacity, while urgent work takes the fastest machines. Luiz Barroso and Urs Hölzle made this the center of their 2009 account of warehouse-scale computing: at building scale, the dominant cost is not the machines themselves but the electricity to run them and the logistics of replacing the ones that fail.
The third fact is trust. A computer on a desk can be unplugged, stolen, flooded, or simply switched off by its owner; no business can build on that. Banks, hospitals, airlines, and governments need machines whose operation they can verify and whose uptime someone guarantees in writing. The data center provides the verifiable part through physical security, badge readers, cameras, locked cabinets, and the contractual part through service agreements backed by the tier system the Uptime Institute defined, where a Tier III hall is certified as concurrently maintainable, meaning any part of its electrical or mechanical system can be serviced without stopping the machines. The International Energy Agency’s estimate that these facilities consume about 1 to 1.5 percent of global electricity reads differently against these three facts: the share is the price of moving computation out of untrustworthy, distant, undersized locations and into buildings engineered for exactly that job. Distance, scale, and trust are the problem; the power and cooling plant is the answer.
The Two Constraints That Decide Everything: Power and Heat
Every watt of electricity a computer consumes becomes a watt of heat, with no exceptions and no negotiation; physics keeps the ledger balanced. A facility that draws ten megawatts of electricity must therefore reject ten megawatts of heat into the outside world, continuously, and the machinery that performs the rejection, pumps, fans, chillers, cooling towers, is itself a second consumer of electricity. This is why the thesis leads with power and heat rather than processing: the computers are the easy part of the building to buy, and the hard part is feeding them and cooling them without interruption. An enterprise rack typically draws 8 to 12 kilowatts, enough that a single row of them needs the electrical service of a small apartment block; racks built for artificial intelligence training run at 40 to 100 kilowatts, and the densest designs approach 120 kilowatts, the NVIDIA GB200 NVL72 class of rack-scale system. The density sets every downstream decision, because heat that cannot be removed will destroy the machines it came from, and electricity that cannot be delivered cleanly will crash them.
The industry’s shorthand for the balance between the two is PUE, power usage effectiveness, defined as total facility energy divided by the energy consumed by the computing equipment itself, a ratio rather than a percent. A perfect score would be 1.0, meaning every watt entering the building reaches a processor. Christian Belady, then at HP, coined the metric in 2006, and The Green Grid, an industry consortium, published its formal definition in 2007. The facts the consortium’s members reported give the scale of the challenge: a typical older facility scores about 2.0, meaning it spends a full watt on power delivery and heat removal for every watt the machines use; the industry average in the Uptime Institute’s annual survey sits at about 1.55; the best hyperscale halls reach about 1.07 to 1.10, with Facebook’s Prineville, Oregon facility reporting roughly 1.07 and Google’s fleet averaging roughly 1.10. The gap between 2.0 and 1.07 is not a difference in computers; it is a difference in plant design, in how cleverly the building moves heat out and power in.
Heat sets the site and the shape before a single rack is ordered. A hall planned for evaporative cooling needs a water supply and a climate dry enough to evaporate into; a hall planned for free cooling, using outside air directly, needs cold ambient temperatures for enough of the year to justify the gamble. ASHRAE Technical Committee 9.9, the standards body for thermal environments in technology spaces, defines the envelope the industry designs against: its 2011 update expanded the recommended range for data processing environments to 18 to 27 degrees Celsius, 64.4 to 80.6 degrees Fahrenheit, which tells an engineer exactly how warm the machine hall is allowed to run and therefore how much cooling work is unavoidable. Power sets the site with equal force: the building must sit near transmission lines and a substation with spare capacity, since electricity cannot be trucked in, and the local utility must be able to contract for the load years before the hall opens. Processing power enters the conversation only after these two constraints have fixed the location, the building’s footprint, and the size of its electrical and mechanical plants. That ordering, power and heat first, computation after, is the whole of the thesis.
How a Request Reaches the Building
A request begins at a device and must cross four kinds of distance before it reaches a rack: the last mile to the internet provider, the provider’s regional network, the long-haul backbone, and the building’s own front door. The last mile is the familiar one, a home router, a cell tower, a fiber line into an office, carrying the request to the provider’s local equipment. From there it enters the provider’s core, a mesh of high-capacity routers that forward packets by reading their destination addresses and choosing the next hop, billions of such decisions per second across the network. Long-haul fiber, much of it buried along rail lines and highways or laid across ocean floors, carries the traffic between regions; a packet from Chicago to a hall in Oregon rides glass the whole way, converted between light and electricity at each end.
The building announces its presence to this mesh with a block of addresses and a routing policy. Its border routers speak BGP, the Border Gateway Protocol, the inter-domain routing protocol of the internet backbone that the Internet Engineering Task Force standardized in RFC 4271, telling neighboring networks which addresses live behind them and learning in return how to reach the rest of the world. Multiple providers connect at the building’s edge so that no single carrier’s failure isolates it, and the routers constantly withdraw and re-advertise paths as links fail and recover, which is why an undersea cable cut reroutes traffic in seconds rather than taking the building offline. At the door, load balancers spread arriving requests across the machines inside, so that no single server drowns while its neighbors idle; the internal network then carries each request down through layers of switches to the rack that will do the work, and the answer retraces the path in reverse. Every hop adds delay, which is why the building’s position on the map, near users and near fiber, matters as much as anything inside it. The network often shortens the trip before it begins: copies of popular content already wait in smaller caching facilities nearer the user, and the routing system sends the request onward to the nearest healthy copy, turning one long trip into a short local one whenever the requested content allows it.
Where Does a Website Actually Live?
A website lives as files and programs stored on machines inside a data center, copied to many machines so that no single failure removes it. The domain name is only an address label; the actual content sits on disks in a machine hall, ready to be served.
The question is worth asking plainly because the common picture, a site floating in a cloud, hides the mechanism. There is no cloud in the physical sense, only other people’s computers, and those computers sit in buildings with addresses, utility bills, and security guards. A small site may live on one virtual machine among thousands in a shared hall; a large one is duplicated across halls on different power grids so that a regional outage leaves it reachable. When a person types an address, the domain system translates the name to a number, the number routes to the building, and the building’s machines assemble the page. Deleting a site means deleting those files from those disks; nothing about it is atmospheric.
The Power Chain: From Grid to Server
Electricity reaches a server through a chain of transformations, and each link exists to answer one failure. The chain begins at the utility’s transmission lines, carrying tens of thousands of volts from distant generating stations, far too high to use directly. On the campus, or just outside it, a substation steps that voltage down to medium levels, typically in the tens of kilovolts, through transformers sized for the entire site with room to grow. From the substation the current enters the building’s switchgear, banks of breakers and switches that divide the feed into branches, isolate faults, and select between sources; this is the first point where the design splits into two, because a serious facility brings in not one utility feed but two, from separate substations or separate paths, so that a single line failure leaves the building fed.
The next link is the UPS, the uninterruptible power supply, and it performs two jobs that look like one. First, it cleans the electricity: in a double-conversion design, incoming alternating current is rectified to direct current and then inverted back to alternating current, which erases sags, surges, and frequency wobble from the grid before they can touch a machine. Second, it bridges interruptions with batteries, carrying the load while other sources start. Past the UPS, power distribution units step the voltage down again, from the 480 volts common in American industrial distribution to the 208 or 120 volts the racks use, and they divide it into circuits with breakers sized for rows and racks. At the rack, a rack-level distribution unit feeds each server’s own power supplies, and the servers carry two of them, one wired to the A feed and one to the B feed, so that any single failure anywhere in the chain, a breaker, a distribution unit, a whole UPS line, still leaves every machine powered.
The UPS battery plant deserves a closer look, because it is sized by time rather than by energy. Lead-acid or lithium strings hold enough charge to bridge the minutes until the generators stabilize, and they are monitored cell by cell, since a single weak cell can drag a whole string below its rating; battery rooms carry their own ventilation because charging chemistry releases hydrogen. The switchgear, meanwhile, carries protective relays that watch for short circuits and open the right breaker in milliseconds, so a fault in one branch does not collapse the bus feeding the rest. Each layer narrows the blast radius of the layer below it.
The redundancy has a grammar the industry speaks fluently. N means the capacity needed to run the load; N+1 means that capacity plus one spare module, so any one unit can fail; 2N means two complete independent paths, each able to carry the full load alone. A 2N electrical design duplicates the switchgear, the UPS systems, and the distribution paths, and the two sides are kept electrically separate so that a fault on one cannot propagate to the other. Operators prove the grammar in maintenance windows: in a 2N hall, one entire side can be de-energized, serviced, and re-energized while the machines ride on the other, a drill that turns the redundancy from a drawing into a demonstrated fact. None of this is visible to the software running above it, which is the point: the application sees only unbroken, clean electricity, while below it a small electrical utility performs continuous triage. The chain is also where the facility’s economics are decided, because every duplicated transformer, every battery string, every oversized feeder is capital spent before the first customer rack is sold, and James Hamilton, the Amazon Web Services vice president and distinguished engineer who writes on data center cost economics, has made exactly this the center of his published analysis: at scale, the electrical and mechanical plant, not the computers, sets the cost of the building.
What Happens When the Grid Fails
The grid fails more often than its reputation suggests, and the data center is built on the assumption that it will fail again. When utility voltage sags or vanishes, the UPS notices in milliseconds, because it is already in the current path, and its batteries pick up the load with no switching transient; the machines never see the event. The UPS battery systems carry the load for minutes, typically 5 to 15, while the diesel generators start and stabilize, a window sized deliberately: long enough for the engines to reach speed, synchronize, and accept load, short enough that the battery plant stays affordable. Diesel backup generators start in seconds, and the automatic transfer switch, watching the utility feed, commands the start the instant the outage is confirmed, then shifts the building’s load onto the engines once their voltage and frequency are stable.
The generator plant is a small power station in its own right. Engines the size of locomotives sit in dedicated rooms or yards, each driving an alternator, each with its own fuel day tank fed from bulk storage on site, and the whole plant is sized to carry the full building load indefinitely as long as fuel contracts hold. Redundancy applies here too: the engines are arranged so that the failure of one still leaves enough capacity, and the fuel system is built to run through extended outages, with deliveries contracted in advance for the kind of regional event that keeps the grid down for days. Stored diesel degrades with time, collecting water and microbial growth, so the fuel is filtered and tested on a schedule, because an engine that starts on paper but chokes on bad fuel is no backup at all. None of this machinery earns its keep by running; it earns it by waiting, maintained and tested, for the hour it is needed.
When an engine refuses to start, the building’s defenses narrow to the battery window and the minutes begin to tick. The UPS alarms, staff shed every non-critical load to stretch the remaining charge, and the transfer logic tries the next engine in sequence; if none catches, the only choices are the grid returning or an orderly shutdown of the machines before the batteries empty. That narrow, unforgiving sequence is why the testing discipline and the tier certifications exist: the disaster is never the outage itself, it is the backup that was assumed to work and did not.
Testing is where many facilities prove themselves. A generator that has not started under load is a rumor, so operators run the engines on a schedule, exercising them against real building load or against load banks, resistive heaters that simulate it, to prove they will accept the transfer when the grid drops. The Uptime Institute’s tier language, the origin of the industry’s grading vocabulary, makes this discipline explicit at its higher levels: a Tier III facility is certified as concurrently maintainable, meaning any component, a UPS module, a generator, a switchgear section, can be taken out of service without dropping the load, and a Tier IV facility is certified as fault tolerant, designed to survive any single equipment failure. The tiers are design and operations certifications rather than guarantees, but they encode the lesson the power chain teaches: surviving the grid’s failure is not a feature of the building, it is the building, and everything from the battery chemistry to the fuel contract exists so that the machines inside never learn the grid was gone.
How Heat Gets Out: Air Cooling
Nearly every watt of electricity that enters a machine hall leaves as heat, so the first heat removal strategy ever used in these buildings was also the most obvious: move a great deal of air past the machines and carry the warmth away. Air remains the default medium in most facilities because it is free, nonconductive, and easy to distribute, though it carries relatively little heat per unit of volume. The engineering of air-based heat removal is therefore the engineering of airflow itself, directing cold supply air exactly where machines inhale and capturing hot exhaust before it can mix back into the supply.
The classic layout is the hot aisle and cold aisle arrangement. Racks stand in rows with their fronts facing each other across a cold aisle and their backs facing each other across a hot aisle. Chilled air is delivered into the cold aisle, drawn through each machine front to back by its internal fans, and expelled as hot exhaust into the hot aisle, where it is collected and returned to the cooling plant. Containment sharpens this separation: physical barriers, rigid panels or hanging curtains, seal the cold aisle or the hot aisle so supply and exhaust cannot mingle above or around the racks. Without containment, hot exhaust recirculates into intakes, machines breathe their own waste heat, and the plant must supply colder air than the hardware actually needs.
Raised floors were the original delivery mechanism. A perforated tile floor stands above a structural slab, creating an underfloor plenum that doubles as a duct and a cable pathway. Conditioned air pressurizes the plenum and rises through perforated tiles placed in the cold aisles. The design concentrates cold air where it is needed and keeps power and signal cabling out of the overhead space. Heavier racks and denser layouts have pushed some operators toward solid slab floors with overhead ducting or in-row cooling units instead, trading the plenum’s flexibility for structural load capacity and shorter air paths.
The machines that produce the cold air fall into two families. CRAC units, computer room air conditioning, use direct-expansion refrigeration circuits, compressors and refrigerant, to cool air passing over their coils. CRAH units, computer room air handling, have no compressors of their own; they blow air across coils fed by chilled water from a central plant, usually a chiller or an economizer loop. CRAH units paired with a central chilled water plant dominate large facilities because the refrigeration happens once, at scale, rather than inside dozens of small compressors scattered across the hall.
ASHRAE Technical Committee 9.9 publishes the thermal guidelines that set the envelope for all of this. Its 2011 update widened the recommended inlet range for equipment to 18 to 27 degrees Celsius, a change that let operators raise supply temperatures and lean harder on free cooling without voiding hardware warranties. Warmer supply air means smaller temperature lifts for chillers and more hours per year when outside air alone can do the job.
Why Do Data Centers Run Warmer Than Office Buildings?
Because the machines prefer it. ASHRAE guidelines allow server inlet air from 18 to 27 degrees Celsius, far warmer than office comfort, and every degree of higher setpoint trims compressor work. People feel chilly at 22 degrees; silicon is content at 27, so the thermostat serves the racks, not the staff.
Airflow management closes the loop between design and operation. Blanking panels fill empty rack slots so exhaust cannot short-circuit through gaps, brush grommets seal cable cutouts in raised floors, and variable-speed fans in CRAH units modulate with load instead of running flat out. Pressure sensors hold the cold aisle at a slight positive pressure relative to the surrounding space, which keeps stray warm air from leaking in. Each measure is small; together they decide whether the cooling plant fights the building or works with it. The limit of the whole approach is thermodynamic: air is a thin medium, and when rack heat density climbs past what fans and ducts can move at reasonable cost, the industry reaches for liquid.
Liquid Cooling and Why AI Forced the Change
Air cooling has a ceiling, and AI training workloads pushed straight through it. A typical enterprise rack draws 8 to 12 kilowatts, a heat load that fans and cold aisles handle comfortably. AI training racks draw 40 to 100 kilowatts, and the densest designs approach 120 kilowatts, the class exemplified by rack-scale systems such as the NVIDIA GB200 NVL72. Ten times the heat in the same steel box defeats air: the volume of airflow needed to carry that much heat out of a single rack would demand fan energy and duct sizes that no machine hall can afford. Liquid carries heat far more effectively per unit of volume, so the industry moved the coolant closer to the silicon.
Direct-to-chip cooling is the most common step up from air. A cold plate, essentially a liquid-cooled heat sink, mounts directly onto each processor or accelerator, replacing the finned air sink. A pumped loop circulates water, or a water-glycol mix, through the plates, and the warmed liquid travels to a coolant distribution unit that transfers the heat to the building’s facility water loop. The rest of the rack, memory, storage, power supplies, still breathes air, so direct-to-chip is a hybrid: liquid handles the hottest components, air handles everything else. It keeps the familiar rack and plumbs it into the building’s water system through quick-disconnect fittings that seal automatically when a machine slides out for service.
Immersion cooling goes further by removing air entirely. Machines, stripped of fans and sometimes of their cases, sit submerged in a tank of dielectric fluid that does not conduct electricity. In single-phase immersion, the fluid simply warms as it absorbs heat and is pumped through a heat exchanger; mineral-oil-based fluids are the common choice. In two-phase immersion, an engineered fluid boils at a low temperature on contact with hot components, and the vapor rises to a condenser cooled by facility water, drips back down, and repeats the cycle. Both versions eliminate nearly all machine fans and erase hot spots, but they change service procedures completely: a technician cannot reach into a tank of oil the way they reach into an air-cooled rack, and every component must tolerate the fluid chemistry.
Rear-door heat exchangers split the difference. A standard rack keeps its air-cooled machines, but its rear door is replaced with a finned coil fed by chilled water. Hot exhaust passes through the coil on its way out of the rack and leaves close to room temperature, so the machine hall never sees the heat. Passive versions rely on the machines’ own fans to push air through the coil; active versions add their own fans. The appeal is retrofit: an existing air-cooled hall can adopt liquid heat removal one rack at a time without replumbing every machine.
None of this displaced air cooling everywhere. Liquid systems add pumps, leak detection, water treatment, and a second piping network threaded through the building, all of which cost money and demand new operational skills. Operators deploy liquid where rack density forces it, in AI training halls and high-performance computing, and keep air where 8 to 12 kilowatt racks still dominate. The design question is no longer which medium is better in the abstract but where each rack’s heat density sits on the spectrum between them.
Water quality matters as much as plumbing. The loops that touch cold plates use deionized or treated water with corrosion inhibitors, because mineral deposits or galvanic corrosion inside a processor block would throttle heat transfer long before any leak appeared. Leak detection runs as its own system: sensing cable routed along pipe runs and under manifolds alarms at the first drop, and isolation valves segment the loop so a single fitting failure drains liters, not the whole circuit.
Water: The Hidden Utility Bill
Heat removed from machines has to go somewhere, and in most large facilities its final destination is the sky, carried there by evaporating water. Air-cooled chillers and CRAH units move heat from the machine hall into a water loop, and that loop must reject its heat in turn. Evaporative cooling performs that rejection by exploiting a physical bargain: turning liquid water into vapor absorbs a large amount of heat, so sacrificing some water to the atmosphere cools the rest. Cooling towers are the machines built around this bargain.
A cooling tower is a box engineered for maximum air-water contact. Warm return water from the chillers sprays down over a lattice of fill material while fans draw outside air up through the falling droplets. A fraction of the water evaporates, and the evaporation chills the remaining flow, which returns to the chillers several degrees cooler. The evaporated fraction must be replaced continuously with makeup water, and a second stream, called blowdown, is bled off deliberately to keep dissolved minerals from concentrating as pure water leaves. Chemistry follows: treatment programs dose the loop with biocides and scale inhibitors, because a warm, aerated basin is an inviting habitat for biological growth and mineral deposits alike.
This is why water enters the picture at all. Mechanical refrigeration can reject heat to air directly through dry coolers, but evaporation moves far more heat per unit of equipment and electricity. Operators trade water consumption for energy efficiency: a tower-cooled plant uses less electricity than an air-cooled one of equal capacity but drinks steadily to do it. The Green Grid, the industry consortium that formalized PUE, defines a companion metric, water usage effectiveness, that divides annual water use by IT energy, giving operators a second ratio to optimize alongside the first. In regions where water is cheap and plentiful, the trade favors towers; where it is scarce or politically contested, designers shift toward dry or hybrid heat rejection and accept the higher electricity bill.
Water risk also shapes siting and permitting. A large hall can evaporate volumes comparable to a small town’s supply, so utilities and regulators ask where the makeup water comes from and where the blowdown goes. Some facilities answer with reclaimed or treated wastewater for cooling, keeping potable supplies for people. Others close the loop further with air-cooled designs that consume almost no water but demand bigger electrical service and larger heat rejection equipment on the roof. Either way, water is not an afterthought to the electricity contract; it is the second utility negotiation, and in dry regions it can be the binding one.
Hybrid designs pair towers with dry coolers and switch between them by season, trimming water use in cooler months while keeping peak heat rejection for the hottest weeks.
Why Do Data Centers Consume So Much Water?
Because evaporative cooling consumes it. Cooling towers reject heat by evaporating water, so every megawatt of heat removed this way sacrifices a corresponding flow of water to the atmosphere. Air-cooled designs drink far less but pay in electricity and larger heat rejection equipment.
The Network Inside: Spine, Leaf, and Fabric
If power and heat removal are the building’s body, the internal network is its nervous system, and it looks nothing like the network in an office. Office networks move traffic up and down a hierarchy toward the internet; machine halls move most traffic sideways, between machines. A training job scatters work across hundreds of accelerators that must exchange partial results constantly, storage arrays feed compute nodes, and load balancers spray incoming requests across fleets. This machine-to-machine pattern is called east-west traffic, and the entire internal topology is designed to serve it with uniform, predictable latency.
The unit of construction is the top-of-rack switch. Each rack carries one or two switches mounted at its top, and every machine in the rack connects to them with short copper or fiber cables. The top-of-rack switch is the machine’s only path out of the rack; it aggregates dozens of machine links into a few high-speed uplinks. Redundant pairs are standard, so a single switch failure does not strand a rack.
Above the racks sits the spine-leaf fabric. Leaf switches, which include the top-of-rack layer, connect to every spine switch, and spine switches connect to every leaf, but leaves never connect to each other and spines never connect to each other. The result is a Clos topology, a design borrowed from telephone switching and described by Charles Clos in 1952: any machine can reach any other machine in exactly the same number of hops, usually two, which makes latency predictable regardless of which racks the two endpoints occupy. New capacity scales horizontally; adding racks means adding leaf switches and, when needed, more spines, without rewiring the existing fabric.
Traffic engineering inside the fabric relies on equal-cost multipath routing. Because many identical paths exist between any two machines, switches spread flows across all of them simultaneously, using every spine rather than queuing behind one. This only works if the fabric is non-blocking in practice, which brings in oversubscription: the ratio of total downlink capacity to uplink capacity at each layer. A fabric with no oversubscription lets every link run at full rate at once but pays for switching capacity that sits idle most of the time; operators therefore accept calculated oversubscription, sizing uplinks for realistic aggregate demand rather than theoretical maximums, and they watch utilization to know when the calculation needs revisiting.
Contrast this with the old three-tier hierarchy of access, aggregation, and core, where traffic between two machines in different parts of the building climbed to a central core and back down, adding hops and bottlenecks. Spine-leaf flattens that climb into a single predictable path. The fabric also carries storage traffic, management traffic for out-of-band controllers, and replication streams between halls, often on logically separated virtual networks riding the same physical switches. When operators speak of the network as a fabric, they mean this: not a hierarchy of boxes but a woven mesh where bandwidth is pooled and any endpoint is equidistant from any other.
Two more mechanisms ride on top of the fabric. Network virtualization overlays, VXLAN being the common encapsulation, let many tenants share one physical fabric while each sees a private network; the underlay switches forward encapsulated packets without knowing tenant boundaries, and the overlay handles isolation. For AI training, the east-west pattern gets an accelerator of its own: remote direct memory access over the Ethernet fabric, usually RoCE, lets one accelerator read another’s memory without involving either host’s processor, cutting latency on the collective operations that synchronize a training step. The fabric is therefore tuned twice, once for general traffic with ECMP and oversubscription, and once for the tightly coupled, latency-sensitive exchanges that decide how fast a model trains.
How Data Centers Talk to Each Other
No facility is an island; the machines inside one building constantly exchange traffic with machines in others and with the wider internet. The internet has no central authority directing this traffic, so networks coordinate through a routing protocol that lets each one advertise which address ranges it can reach. That protocol is the Border Gateway Protocol, the inter-domain routing protocol of the internet backbone, maintained by the Internet Engineering Task Force and standardized in RFC 4271. Every network that participates in global routing, from hyperscale operators to regional ISPs, runs BGP sessions with its neighbors, announcing reachability and withdrawing it when links fail. The global routing table that results is the internet’s collective, continuously renegotiated map.
Much of this coordination happens at internet exchange points. An IXP is a physical location, typically inside a carrier-neutral building, where many networks connect to a shared switching fabric and exchange traffic directly with one another. This direct exchange is called peering: two networks agree to hand each other traffic destined for the other’s customers, usually without payment, because both sides save the cost of sending that traffic through a paid transit provider. Peering keeps local traffic local; a request from a user to a service hosted in the same metro area can hop between the two networks at the exchange instead of traveling to a distant transit hub and back.
Enterprises that need private, encrypted paths from their own offices into cloud fabrics, or that need data center buildings to reach one another across the backbone without touching the public internet, typically turn to managed VPN services rather than raw peering, a subject treated in depth in this Azure VPN gateway deep dive.
The longest links in the system are subsea cables. The majority of intercontinental traffic travels as light through fiber-optic cables laid on the ocean floor, and those cables come ashore at cable landing stations, hardened buildings where the subsea fiber terminates and connects into terrestrial networks. From the landing station, capacity fans out to nearby machine halls over terrestrial fiber, often through diverse physical paths so a single fiber cut cannot isolate a region. Choosing a site near landing stations and dense exchange ecosystems is a network-position decision: milliseconds of latency and the number of independent paths into a building are set at site selection and cannot be retrofitted later.
Two more particulars complete the picture. Route security has become its own discipline because BGP was designed on trust: any network can announce any address range, and mistaken or malicious announcements can divert traffic, an event operators call a route hijack. Resource Public Key Infrastructure, RPKI, counters this with cryptographic attestations that bind address ranges to the networks authorized to announce them, letting routers reject invalid origins automatically. Adoption is uneven, so hijacks remain a failure mode of the backbone. Physical interconnection has its own geography too: inside carrier-neutral buildings, the meet-me room is the secured space where networks’ fiber runs terminate on shared patch panels, so a new peering session can be provisioned by cross-connecting two ports in the same room rather than trenching new fiber across a metro area. The economics follow the physics: peering ports and cross-connects cost far less than paid transit, which is why dense exchange ecosystems attract more networks, which attract more peering, in a loop that concentrates interconnection in a handful of buildings per region.
What Sits in the Racks
Open a rack and the first thing to notice is what is missing: no monitors, no keyboards, no decorative cases. The machines inside are built for a different job than a desktop computer, and every difference traces back to density and serviceability. A desktop is designed for one person within arm’s reach; a rack machine is designed to be one of thousands, managed remotely, serviced in minutes, and packed as tightly as heat removal allows.
The most visible difference is management. Every machine carries a baseboard management controller, a small independent computer on the motherboard that runs even when the main system is powered off. Through it, administrators can power-cycle a hung machine, mount virtual installation media, read temperature and voltage sensors, and reinstall an operating system, all over the network, without touching the hardware. This out-of-band management, standardized in interfaces such as Redfish from the DMTF, is what makes a fleet of ten thousand machines operable by a small team: physical presence is reserved for replacing failed parts.
Power supplies are the second difference. Rack machines typically carry two hot-swappable power modules fed from independent electrical paths, so the loss of one supply or one feed does not interrupt the machine. Memory uses error-correcting codes as a matter of course, because at fleet scale the rare bit flip becomes a routine event rather than a freak occurrence. There is no graphics card for display; accelerators appear only where workloads need them, for training and inference, and they serve computation, never a screen.
Storage lives in the racks too, in arrays built for the same density logic. Disk shelves pack dozens of drives, increasingly solid-state, behind redundant controllers that present the capacity to compute nodes over the fabric. The split between compute and storage varies by design: hyperconverged nodes bundle both in each chassis, while large-scale designs separate them so each can scale and fail independently.
Network gear occupies its own share of rack space: the top-of-rack switches of the spine-leaf fabric, aggregation switches in network rows, and patch panels that organize the fiber and copper trunks running overhead or underfloor. Patch panels look unremarkable, rows of ports in a steel frame, but they are the building’s junction points, where every cable run begins, ends, and can be rerouted without disturbing the machines.
Rack power density ties the whole inventory together. A typical enterprise rack draws 8 to 12 kilowatts, which air cooling handles with room to spare. AI training racks draw 40 to 100 kilowatts, with the densest designs approaching 120 kilowatts in rack-scale systems of the NVIDIA GB200 NVL72 class, and those racks are plumbed for liquid. The rack is a standard steel frame, 19 inches wide, but what it demands from the building, electricity, heat removal, water, network capacity, varies by an order of magnitude with its contents. That variance is why the building is designed around the rack’s appetite rather than its computers.
Two supporting systems fill the remaining rack units. In-rack power distribution units take the feeds from the building’s busways or remote power panels and break them into the outlet strips the machines plug into, with metered versions reporting per-outlet draw so operators can balance phases and spot a runaway load. Above the machines, cable managers and overhead trays keep fiber and copper runs separated and labeled; a mislabeled trunk in a hall of ten thousand machines is a fault waiting for a maintenance window. Accelerator trays deserve special mention: GPU sleds mount eight accelerators on a baseboard stitched together by high-bandwidth interconnects such as NVLink, so the eight behave as one compute unit with shared memory bandwidth far beyond what the PCIe bus could provide. The rack thus holds three populations, compute, storage, and network, plus the power and cabling infrastructure that lets all three be serviced without disturbing their neighbors.
How Thousands of Machines Act as One
Ten thousand machines are only useful if they behave like one machine, and three layers of software make that illusion real. The first is virtualization. A hypervisor sits between the hardware and the operating systems, carving each physical host into multiple virtual machines that each believe they own a computer. This decouples workloads from hardware: a virtual machine can move from one host to another while running, and a failed host’s guests restart elsewhere. Utilization rises because the gaps between workloads, idle memory here, idle cores there, get filled by neighbors sharing the same host.
Containers go a step further by sharing the host’s operating system kernel instead of virtualizing hardware. A container image packages an application with its libraries and configuration, and it starts in milliseconds rather than the tens of seconds a full virtual machine needs. The tradeoff is weaker isolation than hardware virtualization provides, which is why the two are often stacked: containers running inside virtual machines, combining fast startup with strong boundaries.
Orchestration is the layer that schedules work across the fleet. A scheduler watches every host’s available cores, memory, accelerators, and network, matches each unit of work to a host with room, starts it, monitors its health, and restarts it elsewhere when a host fails or is drained for maintenance. It packs workloads tightly like cargo in a hold, spreads replicas across failure domains so no single rack outage takes a service down, and rolls out new software versions gradually, replacing old containers with new ones while the service stays up. Luiz Barroso and Urs Hölzle gave this pattern its name in 2009, in “The Datacenter as a Computer”: warehouse-scale computing, the entire facility programmed and operated as a single machine.
Managed offerings that hide the control-plane machinery make this fleet-wide scheduling model consumable without running the scheduler oneself, as this Azure Kubernetes Service explained walkthrough shows. Whether self-built or consumed as a service, the effect is the same: the operator stops thinking in machines and starts thinking in capacity, and the building’s ten thousand computers answer, for most purposes, as one.
The control plane is the orchestrator’s brain, and Kubernetes supplies the canonical example. An API server accepts declarative descriptions of desired state, how many copies of a service should run and what each needs; a distributed store, etcd, holds that state durably; controllers drive reality toward the description; and an agent on every host, the kubelet, starts and monitors the containers assigned to it. Declarative scheduling inverts the operator’s job: instead of placing work, the operator describes the outcome, and the system converges on it, healing around failures without human intervention.
A second scheduling tradition serves batch work rather than services. AI training and scientific computing run on schedulers in the Slurm family, which allocate whole nodes to a job for its duration and tear the allocation down when it finishes, optimizing for throughput and fair sharing of expensive accelerators rather than for keeping services reachable. The two traditions are converging as training moves onto orchestrated platforms, but the distinction explains the fleet’s behavior: some machines run long-lived services that must never go down, others run jobs that must finish as fast as possible, and the scheduler’s policy decides which goal wins when they compete for the same hardware. Virtualization, containers, and orchestration together are what turn a building full of interchangeable machines into something an operator can reason about as one computer, which is exactly the warehouse-scale idea Barroso and Hölzle named.
The Efficiency Number: PUE and What It Hides
Every data center pays two electricity bills: one for the computation and one for everything that makes computation possible. The chillers, the air handlers, the transformers, the lighting, and the conversion losses between the utility feed and the rack all draw current that never touches a processor. Power Usage Effectiveness, PUE, compresses the split into a single ratio: total facility energy divided by IT equipment energy. A perfect score of 1.0 would mean every watt entering the building reaches the machines, with zero overhead. A typical legacy building lands near 2.0, meaning the infrastructure consumes as much electricity as the computers it protects.
The ratio came from a practitioner, not a committee. Christian Belady, then an engineer at HP, proposed it in 2006 as a way for operators to compare overhead across buildings that differed in every other respect. The Green Grid, an industry consortium, published the formal definition in 2007, and the number became the industry’s common currency for efficiency claims.
The published ranges tell the story of the industry’s overhead. The Uptime Institute’s annual survey puts the industry average at about 1.55. Older buildings, with constant-speed fans and over-provisioned chillers, cluster near 2.0. Hyperscale buildings sit at the other end: best in class runs from about 1.07 to 1.10, with Facebook’s Prineville, Oregon facility reporting roughly 1.07 and Google’s fleet averaging roughly 1.10. The gap between 2.0 and 1.07 is the gap between a building designed for computation and a building designed as a heat-removal plant.
What Is PUE in a Data Center?
PUE stands for Power Usage Effectiveness: total facility energy divided by IT equipment energy. A score of 1.0 is perfect, about 2.0 is typical of older buildings, the industry average is about 1.55 according to the Uptime Institute, and the best hyperscale buildings reach about 1.07 to 1.10. It measures overhead, never total consumption.
The ratio also hides as much as it reveals. PUE says nothing about water: a building can reach 1.1 by evaporating river water through its cooling towers, trading electricity overhead for water overhead. It says nothing about the embodied carbon in the concrete, steel, and machines. It says nothing about utilization: a building running its machines at a fraction of capacity spreads fixed overhead across a small IT load, while a building that packs its racks earns a flattering ratio by keeping the denominator large. And it says nothing about the source of the electricity: coal-fed and hydro-fed buildings can post identical scores. The Green Grid later introduced companion metrics for water and carbon, but PUE remains the headline number, and the headline is incomplete.
That incompleteness is why operators treat PUE as a diagnostic rather than a grade. Tracked over months against a stable IT load, a rising ratio points to fouled coils, misbehaving controls, or creeping inefficiency somewhere in the chain. Engineers who keep the measurement honest alongside the rest of their instrumentation often keep a technical reference index at hand when they compare overhead readings across designs and generations.
A Short History: From Machine Rooms to Hyperscale
The lineage begins in the machine room: the raised-floor halls of the mainframe era, where water-cooled cabinets the size of wardrobes sat behind glass and tape libraries lined the walls. Through the 1980s the minicomputer spread the same pattern into corporate basements: dedicated rooms, dedicated air handlers, dedicated staff. The phrase “data center” entered common use with the commercial internet in the 1990s, when the first buildings existed solely to host other companies’ machines: the telecom hotels and dot-com-era colocation halls that rented floor space, power, and network access by the rack.
The next break came from software. In the first decade of the twenty-first century, virtualization let one physical machine pretend to be many, and utilization (the share of installed capacity actually doing work) climbed out of the single digits where it had languished for years. Consolidation followed: dozens of half-idle machines folded into a few dense ones. The building changed with it. Racks grew denser, the heat per square foot rose, and the air handlers that had sufficed for mainframe rooms started to strain.
The industry also invented a shared vocabulary during these years. The Uptime Institute’s tier language, Tier I through Tier IV, began as a way to classify how much redundancy a building carried, from basic capacity to fault tolerance, and it became the shorthand buyers used to compare facilities. ASHRAE Technical Committee 9.9 codified the thermal side: its guidelines define the recommended envelope for data processing environments, expanded in the 2011 update to 18 to 27 degrees Celsius. Together they turned data center design from craft into specification.
Then the cloud reorganized the business. Amazon Web Services’ 2006 launch proved that computation could be sold as a metered utility, and the hyperscale building became the unit of competition: not a room of machines but a campus of them, designed as one system. Luiz Barroso and Urs Hölzle gave the pattern its name in 2009 with “The Datacenter as a Computer: An Introduction to the Design of Warehouse-Scale Machines,” arguing that the entire facility, power distribution, cooling, networking, and software included, should be engineered as a single computer, the way a processor’s cache and cores are designed together.
Facebook’s 2011 Open Compute Project attacked the cost structure by publishing open hardware designs, and Google’s 2016 application of DeepMind machine learning to cooling controls, reported as cutting cooling energy by up to 40 percent, showed that even mature buildings still held algorithmic savings. Each episode tightened the same loop: cheaper compute, denser racks, harder thermal problems.
Microsoft’s Project Natick ran from 2014 to 2020 as an experiment in changing the environment itself: a sealed steel cylinder of machines was submerged off Orkney, Scotland in 2018 and retrieved in 2020, testing whether the ocean’s cold and the absence of human hands changed failure rates and operating cost. And the training of large AI models rewrote the density curve: typical enterprise racks draw 8 to 12 kilowatts, while AI training racks draw 40 to 100 kilowatts, with the densest designs approaching 120 kilowatts in the NVIDIA GB200 NVL72 class of rack-scale system. The nineteenth-century coal question returns in each era: William Stanley Jevons’ 1865 observation that efficiency gains raise total consumption has held for compute, and the history of the building is the history of demand outrunning every gain.
Running beneath the whole chronology is the efficiency curve. Jonathan Koomey of Lawrence Berkeley National Laboratory documented Koomey’s law: computations per kilowatt-hour double roughly every 1.57 years. The machines got radically more efficient across these decades, and the buildings got radically bigger anyway, the paradox that section 27 takes up at full strength.
The Open Compute Break: Redesigning the Rack
In 2011 Facebook did something the hardware industry had never seen: it published the blueprints. The Open Compute Project released open specifications for the machines, the racks, and the Prineville, Oregon building that housed them: every drawing a manufacturer needed to build identical hardware, free of licensing. The first Open Compute hardware ran at Prineville, and the facility reported a PUE of roughly 1.07, a number that reset the industry’s sense of what overhead was achievable.
The designs embodied a philosophy the project called vanity-free engineering. Server makers had competed on cosmetic bezels, proprietary connectors, and chassis tuned for the showroom; the Open Compute specifications stripped all of it. Power supplies merged into shared shelves serving a whole row. Airflow paths were redesigned so the building’s cold air passed through the machines with minimal obstruction. Components that added cost without adding computation or reliability disappeared. The rack stopped being a vendor’s product line and became an engineering drawing anyone could manufacture.
The Prineville building applied the same logic at the facility level. Filtered outside air cooled the machine hall for much of the year, with the open racks specified for the wider inlet temperatures that made that possible, and the power distribution was designed for the same elimination of waste that governed the server sleds. The building and the rack were engineered as one system, which is exactly the warehouse-scale argument Barroso and Hölzle had made two years earlier.
The deeper change was organizational. By separating the design from the brand, the project let the buyers of machines (the operators of the largest buildings) set the specification, and let contract manufacturers compete on building to it. That inversion, sometimes called disaggregation, moved innovation from the server box to the rack and the facility: the unit of design became the row, the power shelf, the cooling zone.
The approach spread beyond its origin. Cloud operators contributed their own open designs; chipmakers, banks, and telecom firms joined the project and published specifications for storage, networking, and facility infrastructure. The open rack became a common language: an operator could mix sleds from different builders in one cabinet because the interfaces were public. The proprietary chassis, once the margin engine of the server business, became one option among many.
The second-order effect was the rise of the white-box market. With the designs public, original design manufacturers could sell identical machines to any operator, and the buyers, cloud providers, banks, and telecom firms, gained bargaining power they had never held against the branded vendors. The server commoditized faster than it otherwise would have, and the margin migrated from the chassis to the building and the operator.
The project’s lasting effect was on cost structure. When the hardware design is public, the differentiator moves from the box to the building: to how cheaply a facility can deliver power, remove heat, and keep machines fed. That is the thesis of this article wearing overalls: the computation commoditized, while the plant did not.
Where They Get Built and Why Location Is a Strategy
Site selection starts with a map of constraints, not a map of customers. The building must sit where electricity is cheap and abundant, where heat can be rejected without heroic machinery, where water can be drawn or conserved, where fiber already runs, where taxes and land make the math work, and where technicians can be hired. No site offers all six; every project is a negotiation among them.
Electricity dominates because it is both the largest operating cost and the hardest constraint. Developers seek regions with surplus generation and industrial rates, and they sign long-term contracts for blocks of capacity measured in tens of megawatts. Climate matters for the same reason heat matters: cool regions let the building use outside air for much of the year, a technique called free cooling, while hot regions force the chillers to run year-round and erase the PUE gains the designers chased.
Water is the quiet second utility. Evaporative cooling trades electricity for water, so buildings that chase low PUE numbers often chase large water withdrawals too, a trade the local watershed feels before the grid does. Fiber is the third: the building must land near long-haul routes and internet exchange points, because every millisecond of delay between the building and its users is delay the operator cannot engineer away.
Grid interconnection is the hidden seventh factor. A developer cannot simply buy electricity; the project must be granted capacity on the local substation and transmission system, a process that runs through utilities and regulators and stretches across years in congested regions. Developers therefore scout for sites where the grid has headroom. Retired industrial land with an existing substation is prized, because a site without deliverable power is only land.
Tax policy, land cost, and workforce close the list. States and counties court these projects with abatements and fast permitting, and supporters describe the result in the language of economic development: construction jobs, a broadened tax base, and upgraded grid infrastructure that outlasts the project. Critics use a different vocabulary for the same buildings: they describe rising electricity rates for neighboring households, strain on local water supplies, and industrial neighbors that employ few people once construction ends. Both descriptions refer to real effects; the dispute is over which ledger matters.
Water disputes follow the same two-vocabulary pattern as the tax debate. Operators describe evaporative systems as efficient heat rejection and point to closed-loop designs and municipal reuse partnerships. Critics describe the same withdrawals as competition with households and farms during dry years, and ask why a profitable industry should draw from the same aquifer as its neighbors. The engineering fact that the water leaves as vapor rather than as pollution does not resolve the allocation question.
The pattern repeats wherever large campuses are proposed. Supporters point to the property-tax revenue and the fiber and substation upgrades that follow the building. Critics point to the power contracts that reserve generation capacity and the water permits that allocate a shared resource. The building itself does not settle the argument; it concentrates it. A single campus draws more electricity than the town around it and returns more tax revenue than any other single taxpayer: both statements can be true at once, and the public argument is usually about the ratio between them.
The Economics: Build, Rent, or Borrow
The decision of who builds the building is really a decision about whose balance sheet carries the risk. A company that builds owns an asset that depreciates whether the machines are busy or idle; a company that rents converts the same capacity into a monthly bill that scales with use. Utilization decides which side wins: a steady, predictable load amortizes owned infrastructure beautifully, while a spiky or uncertain load punishes it. Colocation grew into a real-estate industry of its own on this logic, and the hyperscale cloud grew into the default choice for workloads nobody wants to forecast. That is why the industry converged on four operating models rather than one.
Should a Company Build Its Own Data Center or Rent Space?
A company builds when it needs total control of physical security and regulatory compliance and can keep the machines busy year-round. It rents when demand is uncertain, capital is scarce, or scale must change quickly. Building buys control; renting buys flexibility; each choice punishes the buyer who chose it for the wrong reason.
Which operating model fits which kind of buyer is the question the four-way comparison below settles: ownership, buyer, scale, and the economic tradeoff that defines each choice.
| Who owns the building | Who buys the capacity | Typical scale | The defining economic tradeoff |
|---|---|---|---|
| Enterprise on-premises | The company itself, through its internal IT organization | One site or a few, from a rack row to a small hall | Full control over security and compliance in exchange for capital locked in owned infrastructure |
| Colocation | Tenants who lease cages or suites from a specialist provider | A few racks to a private suite inside a shared building | Operating expense in place of capital expense, traded for rent that rises with power density |
| Hyperscale cloud | Customers who rent compute by the hour or the month | Effectively unlimited, billed in granular increments | No upfront construction cost, traded for usage bills that grow with consumption and exit friction |
| Edge | Network operators and regional providers serving nearby users | Dozens of racks near users, not thousands in one hall | Low latency for nearby workloads, traded for the highest cost per unit of compute |
The pattern across the rows is the migration of risk. The enterprise model keeps every risk (construction, utilization, refresh) inside one company in exchange for total control. Colocation moves the building risk to a specialist while the tenant keeps the machine risk. The hyperscale cloud moves nearly all of it to the provider and bills by the minute, which suits workloads that grow unpredictably and punishes workloads that run at full tilt forever, where the meter never stops. Edge reverses the geography: small buildings near users trade the economies of the giant campus for latency, and the cost per unit of compute runs highest where the distance is shortest.
The refresh cycle sharpens the choice. Server hardware turns over typically every 3 to 5 years, so a built facility faces a recurring capital decision the renter never sees: every few years the machines inside the owned building must be bought again. The renter pays for that churn inside the monthly rate; the owner pays for it in lump sums. Neither is cheaper in the abstract: the answer depends on how steady the load is and how cheap the company’s capital is.
Who Actually Owns and Operates Them
Three kinds of owners hold the world’s machine halls. The hyperscalers (Amazon, Microsoft, Google, Meta) design, build, and operate their own fleets, and as the four-way comparison above shows, they are both the landlord and the tenant of the largest campuses: they own the building and they are the buyer of the capacity. Their scale lets them negotiate power contracts and hardware prices no smaller operator can match, and their buildings are optimized for their own software stacks.
The hyperscale build model also changed what counts as the machine. Designing their own servers, and in some cases their own processors, let the largest operators tune hardware to their software and strip out everything else, the vanity-free lesson of the Open Compute era absorbed in-house. Their campuses are therefore not simply bigger versions of enterprise buildings; they are different artifacts, closer to factories than to offices.
The second kind is the colocation industry, organized largely as real estate investment trusts: companies such as Digital Realty and Equinix own the buildings and lease conditioned space, power, and connectivity to tenants. The REIT structure exists because the economics resemble property more than technology: long leases, metered power, and tenants who stay for years. These operators sell what the enterprise buyer wants without the construction risk: the middle row of the table, industrialized.
The third kind is the enterprise owner: banks, hospitals, universities, and governments that still run their own halls for workloads bound by regulation, latency, or institutional habit. Regulatory regimes in finance and health care often require data to remain within specific jurisdictions or under direct institutional control, and those rules keep enterprise halls alive even when the economics favor renting. Their buildings are smaller, older on average, and harder to fill efficiently (the PUE numbers near 2.0 cluster here), but they persist where control outweighs cost.
The financial structure follows the physical one. Colocation REITs exist because a building full of tenants on long leases behaves like commercial real estate: predictable rent, appreciating land, depreciable structures. Some hyperscalers have sold buildings and leased them back, separating the property asset from the operating business. The enterprise holdouts, meanwhile, shrink in number as each hardware refresh cycle, typically 3 to 5 years, forces the owner to re-justify the capital in lump sums, the exact discipline that renting was invented to avoid.
Ownership has concentrated as the buildings have grown. The largest campuses belong to a short list of cloud operators and colocation firms, while thousands of small enterprise rooms quietly shrink as their workloads migrate to rented capacity. Seen through the table’s four rows, the map of ownership is really a map of who can afford to build power plants that compute. Governments complicate the picture only slightly: sovereign and military facilities follow the enterprise logic of total control, with security clearances layered over the same redundancy mathematics. A reader keeping the terminology straight will find the supplementary notes a useful companion to the ownership breakdown.
Failure Mode One: Power
Power failure inside a data center is a choreographed sequence, not a single event. When the utility feed drops, the automatic transfer switch detects the loss within cycles and the uninterruptible power supply (rooms of batteries and inverters) takes the full load instantly, with no interruption the machines can perceive. The batteries are sized to carry the building for minutes, typically 5 to 15, which is the window the diesel generators need: they crank, reach speed, stabilize voltage and frequency, and signal readiness, usually within a minute of the outage. The transfer switch then moves the load to the generators, and the building runs on its own fuel until the grid returns. The return journey is equally sequenced: when utility power stabilizes, the switch transfers the load back, the generators cool down unloaded, and the UPS recharges its batteries for the next event.
Every step has a backup, and every backup has a failure mode. Generator plants are built N+1 or better, meaning the design tolerates the loss of one unit, and the fuel tanks hold days of diesel. But the most common generator failure is mundane: the starter batteries that crank the engine are dead, or the fuel has absorbed water and grown microbial contamination during months of sitting idle. That is why operators test-start their generators on schedule and run them under load: an untested generator is a rumor of a generator.
The transfer switch is the single most consequential device in the chain: a switch that fails to transfer strands a building with healthy generators and dark racks. Switches get stuck, their control wiring corrodes, and a mistimed transfer can briefly parallel the generator with a still-energized grid, a fault the protective relays exist to prevent. UPS systems fail differently: a single bad battery cell in a long string drags the string down, and the monitoring system must find the weak cell before the outage does. Capacitor aging in the inverter, another slow failure, shows up as heat and noise months before it shows up as darkness.
When the chain breaks, the failure is rarely graceful. A generator that starts and then dies under load drops the building back onto batteries already partially drained; a second transfer attempt may find the UPS in bypass, feeding raw generator power (or nothing) to the racks. Protective breakers trip in sequence, shedding load to save the bus, and the operators watch the machine hall go dark in zones. The servers do not wind down; they stop. File systems and databases written for orderly shutdown meet sudden silence instead, which is why the recovery after a power event is measured in hours even when the outage lasted minutes.
The entire power architecture (redundant feeds, battery rooms, generator yards, switchgear lineups) exists to make one promise: that the machines never notice the grid. Every component in the chain is a bet about which failure is most likely, and the history of outages is the history of bets that lost.
Failure Mode Two: Cooling and Water
A machine hall runs on a bargain with physics. Every watt of electricity a processor consumes returns as a watt of heat, and that heat must leave the room as fast as it arrives. Power has elaborate redundancy, with batteries, generators, and duplicate feeds. Heat removal has redundancy as well, but it fails differently and on a shorter fuse. When the chillers stop, the clock starts immediately.
A typical air-cooled hall holds its inlet air inside the envelope recommended by ASHRAE Technical Committee 9.9, the group that writes the thermal guidelines for data processing environments. The committee’s 2011 update set that envelope at 18 to 27 degrees Celsius. The machines themselves decide the deadline. Processors begin reducing their clock speed as inlet temperatures climb past design limits, trading performance for survival. Past a second, higher threshold, firmware shuts the machine down to protect the silicon. Those protections save hardware and kill service. The building’s thermal management therefore works with a failure budget measured in minutes rather than hours.
The most instructive failure is a quiet one. The pumps stop. Chilled water loops depend on circulation as much as on the chillers that cool the water, so a pump failure, a valve left closed after maintenance, or an air lock in a pipe can turn a working chiller plant into a stranded asset. Buffer tanks buy time. A large chilled water storage tank gives operators a window measured in tens of minutes while heat accumulates in the water, and control systems use that window to shed load automatically. When the window closes, supply temperatures rise and the hall’s air handlers push warmer and warmer air across the racks.
Heat removal is only half of this failure mode. Many facilities cool their condensers and chiller plants with evaporative towers that trade water for electricity savings, letting evaporation carry heat into the air. Those towers need continuous make-up water. A municipal supply interruption, a failed intake pump, or a drought restriction can starve them. The design answer is on-site storage, tanks sized to cover a defined outage, sometimes backed by a second source such as a well or reclaimed supply. Water storage is heavy and expensive, so it is sized in hours rather than days.
The physics punishes delay. A hall drawing 10 megawatts of electricity produces 10 megawatts of heat with nowhere to go. Stop removing it and the room temperature climbs on a curve that steepens while the machines keep working at full load. Operators who notice early can throttle or shut down nonessential systems to slow the climb. Operators who notice late face a cascade. The hottest racks trip first, their workloads fail over elsewhere, and the surviving machines run harder, generating still more heat.
Water adds a slower failure that cooling engineers plan for explicitly. Evaporative cooling consumes water for as long as the towers run, and in regions with limited supply the question is not whether a single pump fails but whether the water is there at all. Facilities in dry climates sometimes choose dry coolers that use no water, accepting higher electricity consumption and larger equipment as the trade. The choice between water and electricity is a design decision made once and lived with for the life of the building.
The honest summary is that this failure mode is a race against accumulation. A power failure is binary. The electricity is present or it is not, and the batteries and generators are sized to bridge the gap. A cooling failure is gradual and unforgiving. The heat keeps arriving, the removal capacity shrinks, and every minute of delay narrows the remaining options. Buildings survive it by layering time: storage tanks, redundant pumps, and automatic load shedding, each layer buying minutes for the next. When every layer is spent, the machines protect themselves by switching off, and the facility’s perfect power record means nothing.
Failure Mode Three: Network and Human Error
The machines can have perfect power and perfect heat removal and still be unreachable. The network is the third system that must not fail, and it fails in ways that have nothing to do with electricity or temperature. A fiber cable cut by excavation equipment, a misconfigured router, or a flawed software push can isolate a building whose every machine is healthy.
Fiber cuts are the bluntest form. The cables that carry traffic between facilities run along highways, rail corridors, and utility rights of way, and excavation remains a persistent cause of damage. A single cut rarely severs a well-designed facility, because traffic enters through multiple physically separate paths. Two cuts on the same route, or a cut that lands during a maintenance window on the backup path, can partition the building from the rest of the internet. Repair is physical work. Crews must locate the break, splice new fiber, and test the span, a process measured in hours rather than minutes.
More subtle are the failures of the control plane, the routing system that decides which path traffic takes. The Border Gateway Protocol, the inter-domain routing protocol of the internet backbone standardized by the Internet Engineering Task Force in RFC 4271, works on trust. Networks announce which address blocks they can reach, and their neighbors believe them. A misconfiguration can withdraw correct announcements or advertise wrong ones, and the error propagates worldwide in minutes. Traffic bound for a facility can be drawn into a dead end or a loop while every cable stays intact. Because the protocol has no built-in authentication of who may announce which addresses, recovery depends on engineers identifying the bad announcement and filtering it, while the rest of the internet’s routers slowly converge on the corrected picture.
Cascading failures add a third shape. Modern facilities are managed by software that configures thousands of devices at once, and a flawed automation push can misconfigure the entire fleet simultaneously, including the management network engineers would use to fix it. The machines that should carry the repair commands are the machines the bad push broke. Recovery then requires physical access to consoles, device by device. That is why out-of-band management networks exist, physically separate from the production network, so that a failure in one cannot strand the tools needed to repair it.
Underneath all of this sits the most persistent cause of all, which is people. Outage investigators, including the analysts at the Uptime Institute who publish the industry’s annual outage surveys, treat human error as a standing category in the failure taxonomy rather than a surprise. The error is rarely a single careless act. Investigations typically find a chain: a procedure that was ambiguous, a change window that overlapped with another, a validation step skipped under time pressure, a configuration that was correct on the test system and wrong in production. The industry’s response has been procedural rather than technological. Change management boards review risky work, peers review configurations, automated validation rejects dangerous changes before they deploy, and blameless postmortems treat the outage as evidence about the system rather than about the person.
The pattern across all three network failure types is that redundancy handles the failures it was designed for, and human process handles the rest. Diverse fiber paths survive one cut. Careful routing policy survives one misconfiguration. What survives a bad automation push or an overlapping maintenance window is discipline: the checklist, the second pair of eyes, the rollback plan tested before it is needed. The building is a fortress against physics, and its most common invader walks in through a keyboard.
Physical Security: The Building as a Vault
The facility treats its own perimeter the way a bank treats a vault door. The logic is layered delay. Every barrier is designed to slow an intruder long enough for detection and response to catch up, and every layer assumes the one outside it has already failed.
The outer layer is distance and discouragement. Sites are often set back from public roads behind fencing, with vehicle barriers and bollards rated to stop a truck. Landscaping doubles as standoff. The building itself usually carries no sign identifying the tenant or the purpose of the site, because obscurity is cheap and effective. Loading docks and staff entrances sit behind controlled gates, and every approach is watched by cameras whose feeds land in a security operations center staffed around the clock.
Entry to the building runs through a mantrap, a small vestibule with two interlocking doors where only one can open at a time. A person badges in, the first door closes behind them, and the second opens only after the system verifies the credential. Biometric checks, fingerprint, palm geometry, or iris scans, commonly sit alongside the badge and a personal code, so that a stolen card alone buys an intruder nothing. Anti-tailgating sensors and turnstiles make it difficult for a second person to slip through on one badge. Visitors face a stricter version: identity verification, a recorded purpose, an escort for the entire visit, and a visitor badge that opens nothing on its own. Logs record every entry and exit, and those logs are retained and audited.
Inside, the building is subdivided. The machine halls are not open floor space that any badged employee can wander. Tenants in shared facilities rent locked cages or private suites, fenced enclosures within the hall with their own access controls, so that one customer’s staff cannot reach another customer’s racks. Even within a single operator’s space, the most sensitive areas, the rooms holding network cores, backup media, and key management hardware, carry their own access lists. Camera coverage follows the same logic, watching corridors, hall entrances, and cage perimeters rather than blanketing every aisle, with retention policies measured in weeks or months.
The hardware holding the cryptographic keys gets the innermost treatment. Encryption keys live in tamper-resistant hardware security modules that erase their contents if they detect physical attack, and those devices sit inside the most restricted rooms of an already restricted building. Facilities chosen to host key management infrastructure are evaluated partly on the physical protections surrounding those devices, a subject examined in depth by Azure managed HSM explained. The principle is simple. Logical security assumes the attacker is already inside the network, so physical security exists to make sure that assumption stays hypothetical.
The final layer is procedural. Guards patrol, drills rehearse intrusions and evacuations, and background checks screen the staff who hold the broadest access. Equipment leaves the site only through documented destruction or decommissioning processes, because a discarded drive is a breach that walks out the loading dock. The building is a vault, but a vault is only as strong as the people who hold its combinations, which is why the procedures matter as much as the concrete.
Logical Security: Segmentation and Encryption
If physical security keeps intruders out of the building, logical security assumes they are already inside the network and limits what they can reach. The core technique is segmentation. The network is divided into zones, and traffic between zones passes through filters that decide what is allowed.
The zones follow function. Public-facing systems that must accept connections from the internet sit in one segment. The machines that run customer workloads sit in another. The management network that configures switches, servers, and power equipment sits in a third, isolated from production traffic so that a compromise of a workload cannot reach the controls. Backup systems sit in a fourth, because the data that restores the business after an attack must survive the attack itself. Between every pair of zones stand filters, firewalls and access control lists that permit only the specific conversations each zone needs. A web server may talk to its database on one port, and nothing else in either direction. The details of how those filtering devices compare are the subject of Azure firewall comparison, which walks through the tradeoffs between managed firewall services, virtual appliances, and security groups for traffic segmented and filtered between security zones inside the network.
Encryption provides the second wall. Traffic between facilities and between the facility and the outside world travels encrypted, so interception yields ciphertext rather than content. Data at rest, on drives and in backup media, is encrypted as well, so a stolen disk is a brick. Modern practice encrypts by default rather than by exception, removing the decision from individual engineers. The ciphers themselves are standard and public, which is a strength rather than a weakness, because open algorithms survive scrutiny that secret ones do not.
Key management is where encryption succeeds or fails. Keys are generated, distributed, rotated, and destroyed according to policy, and access to them is logged. The most sensitive keys never leave tamper-resistant hardware, and no single person can retrieve them alone. Rotation matters because a key that never changes is a key whose theft never expires. Facilities automate rotation schedules so that the practice does not depend on anyone remembering.
Monitoring closes the loop. Intrusion detection systems watch for traffic patterns that match known attacks, and anomaly detection flags behavior that deviates from the baseline, such as a server that suddenly begins copying large volumes of data at an unusual hour. Logs from network devices, servers, and access controls flow into central analysis, where correlation can reveal an attack that no single sensor would catch. The response playbook is rehearsed: isolate the affected segment, preserve evidence, and restore from backups that the segmentation was designed to protect.
The honest limit of logical security is that it defends the network as designed, and networks are redesigned constantly. Every new service, every migration, every merger of two companies’ systems redraws the zone boundaries and the filter rules. Misconfiguration is the recurring failure, a filter left open during a migration and never closed, a test segment accidentally bridged to production. That is why the discipline matters more than any single device. Segmentation, encryption, and monitoring are not products that are bought once. They are practices that are maintained for as long as the building operates, against an adversary that studies every published defense.
The Strongest Case Against Them
The critics have a case, and it deserves to be stated at full strength before anything else is said about it. Their argument is not that the machines are useless. It is that the costs are real, local, and unevenly shared, while the benefits are diffuse, distant, and captured elsewhere.
Start with electricity. The International Energy Agency estimates that data centers account for about 1 to 1.5 percent of global electricity use, a share that sounds modest until it is translated into the language of grids. Supporters describe this demand as the cost of digital infrastructure, the price of an economy that runs on computation. Critics describe the same megawatts as a strain on grids that were planned before these loads existed, forcing utilities to keep older plants running, build new generation, and upgrade transmission. When a single campus draws power on the scale of a small city, the question of who pays for the grid upgrades becomes a live dispute. Supporters point to the rates and connection fees the facilities pay. Critics point to residential ratepayers and ask why households should underwrite infrastructure for some of the most valuable companies in the world.
Water carries the same structure. Supporters describe evaporative cooling as an efficiency measure, a way to cut electricity consumption by letting water carry heat away. Critics describe it as consumption of a public resource, measured in millions of gallons, in regions where farmers, households, and ecosystems compete for the same aquifers and rivers. The dispute sharpens in dry climates, where critics ask what a community gains by allocating scarce water to buildings whose output leaves town on fiber. Supporters answer with jobs and tax revenue. Critics answer with a number: the permanent jobs at a finished facility are few relative to its footprint, and the tax abatements offered to attract construction can run for decades.
Then there are the costs that never appear in an energy report. The low hum of cooling equipment carries across fence lines, and neighbors describe it as a sound that never stops. Supporters call it the normal operation of industrial equipment, comparable to any factory. Critics call it a permanent change to the character of a place, imposed without the consent of the people who live next to it. Land follows the same pattern. A campus occupies acreage that could hold housing, farms, or businesses that employ more people per acre, and supporters describe that land use as productive while critics describe it as a poor trade for the community.
Underneath the particulars sits the fairness question, which is the strongest part of the case. The benefits of computation flow to users everywhere and to shareholders anywhere, while the electricity demand, the water consumption, the noise, and the land use land in one specific place. Supporters frame this as the ordinary geography of infrastructure, no different from a power plant or a highway. Critics frame it as an extraction relationship, in which a locality absorbs the burdens and exports the value. The question of who decides sharpens it further. Local zoning boards and utility commissions approve projects whose effects will be felt nationally, and supporters describe that local approval as sufficient democratic process while critics describe it as a mismatch between the scale of the decision and the scale of its consequences. Both vocabularies describe the same buildings. The dispute is about which description is honest, and it is not settled by engineering.
The Complication: Efficiency Has Not Shrunk the Footprint
There is an objection that cuts deeper than any complaint about noise or water bills, because it is aimed at the logic of progress itself. It holds that every gain in efficiency has made the problem larger, and it arrives with a pedigree that engineers ignore at their peril.
The argument begins with William Stanley Jevons and his 1865 book The Coal Question. Jevons observed that improvements in the efficiency of coal use did not reduce coal consumption in nineteenth-century England. They raised it. Cheaper steam power made steam power economical in more places, and total consumption climbed. The mechanism is now called the Jevons paradox. Efficiency lowers the effective price of a resource, demand expands to absorb the savings, and the total grows.
Applied to compute, the record is stark. Jonathan Koomey of Lawrence Berkeley National Laboratory formulated the regularity known as Koomey’s law: the energy efficiency of computing, measured as computations per kilowatt-hour, doubles roughly every 1.57 years. That is one of the most sustained efficiency gains in the history of technology. Facility efficiency improved in parallel. Power usage effectiveness, the ratio of total facility energy to IT equipment energy coined by Christian Belady in 2006, fell from about 2.0 in older facilities toward 1.07 at the best hyperscale sites, with Facebook’s Prineville, Oregon facility reporting roughly 1.07 and Google’s fleet averaging roughly 1.10. Cooling itself got smarter. In 2016, Google applied DeepMind machine learning to its cooling controls and reported cooling-energy reductions of up to 40 percent.
And total electricity consumption by the industry kept growing. Every one of those gains was outpaced by demand growth, as cheaper and more efficient computation opened uses that had been uneconomical before: streaming video, cloud services, and then AI training workloads whose rack densities, 40 to 100 kilowatts and beyond, would have been unthinkable a decade earlier. The efficiency literature, carried by energy researchers who study this pattern across sectors, states the finding plainly. Gains in efficiency per computation have never, in the industry’s history, translated into a shrinking total footprint. The industry’s real product, on this reading, is ever more electricity converted into heat.
That is the objection at full strength. The verdict is two-sided, because the thesis is a claim about buildings, not about the industry’s total size.
Where the thesis survives: the power-and-cooling constraint still governs every design and siting decision. No efficiency gain has repealed the physics. A rack drawing 100 kilowatts still needs 100 kilowatts of delivery and 100 kilowatts of heat removal, and the engineers who site, build, and operate these facilities still answer to the utility contract and the cooling plant before they answer to anyone else. Jevons explains why there are more buildings. It does not change how any single building works, and this article’s claim was always about the building.
Where the thesis narrows: efficiency does not equal sufficiency. The article’s account explains how these buildings are built and run. It does not explain, and cannot promise, that the industry’s total footprint stops growing. Anyone who reads the efficiency numbers as a story of a problem being solved is misreading them. They are a story of a constraint being managed ever more cleverly while the totals rise. The thesis survives intact within its scope, and its scope is smaller than a casual reader might wish.
Reading Any Data Center Story: Three Questions
The payoff of understanding the physical plant is not the ability to build one. It is the ability to read. Every public story about these buildings, the announcement of a new campus, the dispute over a permit, the boast about an efficiency record, is an argument about a physical constraint wearing the clothes of a business story. Three questions strip the clothes off.
The first question is: which constraint is actually being argued about. Power delivery, heat removal, water, and network position are the four candidates, and a story is usually about exactly one of them. An announcement that emphasizes job creation and tax revenue is often, underneath, a story about power delivery, because the utility contract and the grid interconnection are what make the campus possible. A dispute about permits is often a story about water or about heat, because the cooling design is what the neighbors will live with. A story about latency and new markets is a story about network position, about where fiber lands and how many milliseconds separate the building from its users. Naming the constraint correctly is most of the work. The rest of the story usually follows.
The second question is: what do the location and the utility arrangements reveal. Buildings of this kind are not placed by accident. A site near cheap, abundant electricity and a cool climate is a bet on power delivery and free cooling. A site near a fiber landing or a network exchange is a bet on position. A site that required new transmission lines or a dedicated substation tells you that power was the binding constraint and that someone paid to unbind it. The press release will talk about partnership and growth. The substation tells you what the partnership cost and who needed it. Readers who learn to look past the announcement to the interconnection queue and the water permit will rarely be surprised by what happens next.
The third question is: who absorbs the constraint and who captures the value. Every one of these buildings converts local inputs, electricity, water, land, quiet, into a product consumed everywhere. The arrangement can be fair or unfair, but it cannot be understood without asking the question. When supporters describe economic development and critics describe extraction, they are giving different answers to this question about the same building. A careful reader does not pick a vocabulary in advance. The reader asks who signed the utility contract, who holds the water rights, whose rates funded the grid upgrades, and where the profits land, and then decides which description fits.
Those three questions are the whole of the practitioner takeaway. They require no engineering background, only the habit of looking for the physical facts underneath the business language. Apply them to any story about a new campus, a permit fight, an outage, or an efficiency claim, and the story will resolve into an argument about one of the four constraints. That is a useful skill in an era when these buildings are among the largest new loads on the world’s grids, and it is the reason the physical plant matters more than the computers inside it. You now know what to look for. The rest is practice.
Strip away the announcements, the renderings, and the ribbon cuttings, and what remains is a building organized around two unglamorous facts. Electricity must arrive without interruption, and heat must leave without pause. Everything else, the fiber, the racks, the software that makes thousands of machines act as one, is arranged in service of those two facts. The industry calls the buildings data centers, a name that foregrounds the information and hides the plant. A more honest name would foreground the plant. They are power and cooling facilities that happen to compute, and their location, their architecture, and their economics are decided by the utility contract and the cooling tower long before processing power enters the discussion.
The carryaway is a single habit of mind. When you encounter any claim about these buildings, ask what it says about power and heat. An efficiency record is a statement about how cleverly the heat is removed. A new campus is a statement about where power was available. An outage is a statement about which layer of the power or cooling or network design gave way. A community dispute is a statement about who bears the cost of the power and the water. The computers are interchangeable. The constraints are not, and they are visible to anyone who knows to look.
Frequently Asked Questions
Q: What Do the Four Uptime Institute Tier Levels Mean?
The four levels are design and operations certifications from the Uptime Institute, the origin of the tier language. They describe how power and cooling paths are built and maintained, not a guarantee of uptime. Tier I, called Basic Capacity, provides a single path for power and cooling with no redundancy and carries 99.671 percent availability. Tier II, Redundant Capacity Components, adds spare components such as an extra UPS module and reaches 99.741 percent. Tier III, Concurrently Maintainable, uses dual independent distribution paths so any single component can be serviced or replaced without shutting down the IT load, reaching 99.982 percent. Tier IV, Fault Tolerant, tolerates the failure of any one element anywhere in the power or cooling chain without interrupting computing, reaching 99.995 percent. Each step up costs substantially more because redundancy means duplicated switchgear, generators, UPS systems, and pipework that sit idle most of the time.
Q: What Is the Difference Between a UPS and a Backup Generator?
A UPS, or uninterruptible power supply, and a backup generator serve different moments in the same outage. The UPS is a bank of batteries that sits between the incoming grid power and the servers, delivering power continuously and instantly. When the grid drops, there is no switchover delay because the servers are already drawing from the batteries. That battery reserve lasts minutes, typically 5 to 15, which is enough time for the diesel generators to start and stabilize. The generators then take over and can run for hours or days, limited only by fuel supply and refueling logistics. Put simply, the UPS bridges the gap and cleans up power quality problems such as surges and sags, while the generator provides the long-duration supply. A facility that has one without the other has an incomplete strategy: batteries alone die in minutes, and generators alone leave a dark gap before they reach speed.
Q: What Is Immersion Cooling?
Immersion cooling removes heat by submerging entire servers in a tank of non-conductive dielectric fluid rather than blowing air across them. The fluid absorbs heat directly from chips and boards, then transfers it to a heat exchanger where water or another loop carries it away. Two variants exist: single-phase immersion, where the fluid stays liquid and is pumped to a cooler, and two-phase immersion, where the fluid boils at the hot chip surface and condenses on coils above the tank, carrying away far more heat per unit of fluid. Liquid touches heat sources thousands of times more effectively than air, which is why immersion gained attention as AI training racks rose toward 40 to 100 kilowatts, far beyond what conventional air handling was designed for. The tradeoffs are real: tanks and fluid cost money, hardware must be designed or modified for submersion, and maintenance means lifting servers out of liquid.
Q: What Is Hot Aisle and Cold Aisle Containment?
Server racks are arranged in alternating rows so that the fronts of the racks, which draw in cool air, face one aisle and the backs, which exhaust hot air, face the next. The aisle receiving cool air is the cold aisle and the one receiving exhaust is the hot aisle. Containment adds physical barriers, typically roof panels and end doors over the aisles, so the two streams cannot mix. Without containment, hot exhaust drifts back into server intakes, forcing cooling equipment to work harder and produce colder air than the servers actually need. With containment, supply air goes straight into the machines and exhaust goes straight back to the cooling units, which raises return temperatures and lets chillers run more efficiently. The arrangement also allows higher room temperatures within the ASHRAE recommended envelope of 18 to 27 degrees Celsius, since what matters is the air entering each server, not the average of the room.
Q: What Is Free Cooling and Where Does It Work Best?
Free cooling means using the outside environment to reject server heat instead of running energy-hungry refrigeration compressors. The simplest form is an air-side economizer that draws cool outside air directly into the building when conditions allow; water-side economizers use cool water from a tower or loop to chill the data hall without a chiller stage. It works best where the climate is cool for most of the year and humidity stays within limits, which favors high-latitude regions, high altitudes, and coastal areas with steady cool air. Sites chosen for free cooling can turn off compressors for thousands of hours annually, cutting a large share of cooling energy. The limitation is that free cooling is partial: designers still install mechanical cooling for hot spells, and air quality matters because dust, salt, or smoke drawn into a data hall can damage equipment, so filtration and controls become part of the bargain.
Q: How Does Edge Computing Differ From Hyperscale Cloud on Latency?
The difference comes down to distance. Light travels fast, but a round trip across a continent still takes tens of milliseconds, plus switching and queuing delays at each hop. Hyperscale cloud concentrates enormous capacity in a small number of giant campuses, which keeps computing cheap and efficient but puts the servers far from many users. Edge computing reverses the tradeoff by placing smaller facilities close to where people and devices are, cutting the round-trip path and shaving latency down to a few milliseconds. Applications that react in real time, such as factory automation, vehicle coordination, or interactive video, gain from that shorter path, while workloads that are not time sensitive gain little and would pay more per unit of compute at the edge. In practice the two are complements: edge handles the urgent slice locally and hands heavy, patient work back to the hyperscale core.
Q: What Happens to Old Servers When They Are Retired?
Servers are usually replaced on a refresh cycle of 3 to 5 years, driven by warranty expiry, efficiency gains, and the arrival of workloads the old machines cannot handle. Retirement starts with secure data destruction, because disks leaving the building must carry no customer information; certified wiping or physical shredding handles that step, and chain-of-custody records track each drive. After sanitization, working machines are often resold on the secondary market, where they serve smaller operators or laboratories at a fraction of new prices. Components such as memory, power supplies, and chassis can be harvested for spare parts. What remains goes to electronics recyclers that recover copper, aluminum, gold, and rare earth materials, though some fractions still end up as waste. Large operators run formal circularity programs to keep as much material as possible in use, because the sheer volume of retired hardware makes disposal a significant cost and a reputational matter.
Q: Who Inspects or Certifies a Data Center?
Several different bodies certify different things, and their stamps are not interchangeable. The Uptime Institute audits facilities against its Tier standards and issues design documents and operations certifications for Tier I through Tier IV, which speak to power and cooling architecture and how the site is run. Separately, independent auditors assess information security management: SOC 2 reports examine controls over security and availability, and ISO 27001 certification covers the management system for protecting information. Facilities that handle payment card data face PCI DSS assessments, and sites serving government customers may need FedRAMP authorization. Commissioning is another layer entirely, where engineers test every system under load before the building accepts its first customer, verifying that generators start, UPS units transfer, and cooling responds as designed. Buyers read these certifications together because no single one covers architecture, operations, and information security at once.
Q: What Is Dark Fiber?
Dark fiber is optical fiber cable that has been installed but is not carrying traffic, meaning no light pulses travel through it and no electronics are attached at either end. Telecommunications companies and infrastructure owners routinely lay far more strands than they need, because trenching and permitting dominate the cost and extra strands are cheap insurance. They then lease the unused strands to other organizations, which attach their own transmission equipment and light the fiber themselves. For data center operators and large networks, dark fiber offers dedicated capacity between sites without sharing bandwidth with anyone else, plus control over upgrades: when faster optics become available, the lessee swaps the endpoint equipment rather than renegotiating a service contract. The arrangement suits organizations with predictable, high-volume traffic between fixed points, such as two data centers that must stay tightly connected.
Q: What Is a Meet-Me Room?
A meet-me room is a secure, carrier-neutral space inside a data center where different networks physically connect to each other. Racks in the room terminate fiber from many carriers, internet service providers, and cloud on-ramps, and technicians run short cross-connect cables between any two parties that agree to exchange traffic. Because the room is neutral, no single carrier controls access, which lets a customer in the building reach dozens of networks without building separate facilities for each. Internet exchange points often operate inside meet-me rooms, allowing many networks to peer with each other in one place. The economic logic is density: every additional network in the room raises the value of being there, which attracts more networks in turn. For a colocation customer, proximity to a busy meet-me room can matter as much as the quality of the power and cooling.
Q: How Do Subsea Cables Connect to Data Centers?
Subsea cables come ashore at landing stations, buildings on the coast where the undersea fiber terminates and its signals are amplified and converted for terrestrial networks. From the landing station, high-capacity terrestrial fiber, often called backhaul, carries the traffic inland to data centers, sometimes over hundreds of miles. Operators choose landing station locations partly for this reason: a cable is only as useful as the fiber paths connecting it to the buildings where content and cloud services live. Redundancy matters enormously at this junction, because a single cable cut or a single backhaul route represents a concentrated point of failure for traffic between continents. Large data centers near cable landings therefore advertise diverse backhaul paths, meaning the fiber leaves the building by multiple geographically separated routes, so that one construction accident or fiber cut cannot isolate the facility.
Q: What Is a Power Distribution Unit (PDU)?
A power distribution unit is the device that takes bulk electrical power arriving at a rack row and divides it safely among individual racks and servers. In its simplest form it is a sophisticated power strip mounted inside the rack, but in data centers PDUs are usually intelligent: they meter power draw per outlet or per circuit, report consumption to monitoring systems, and can switch outlets on or off remotely. That metering is operationally important because each rack has a power budget, often tied to the cooling capacity assigned to it, and exceeding the budget risks tripping breakers or overheating. PDUs sit at the end of a long chain that runs from the utility feed through switchgear, UPS systems, and distribution panels, stepping voltage down at each stage until it reaches the form servers can use. They are the last organized stop before electricity becomes computation.
Q: Why Are Data Centers Built Where Electricity Is Cheap?
Electricity is the largest recurring operating cost of a data center, dwarfing the amortized cost of the building itself over a facility’s lifetime. A campus planned in the hundreds of megawatts draws as much power as a small city, and it draws it around the clock, so even a small difference in the price per kilowatt-hour compounds into an enormous annual gap. That arithmetic explains why operators hunt for regions with abundant, inexpensive power, whether from hydroelectric dams, wind and solar resources, or simply uncongested grids. Analysts of data center economics, including James Hamilton of Amazon Web Services, have long argued that power price and availability outweigh almost every other siting factor except network connectivity. Cheap electricity also pairs with cooler climates in many favored regions, which trims the second great cost, cooling, at the same time.
Q: What Is Water Usage Effectiveness (WUE)?
Water usage effectiveness is the companion metric to PUE, measuring how much water a data center consumes per unit of computing energy. The Green Grid defined it as the annual water usage, in liters, divided by the annual IT equipment energy, in kilowatt-hours, so the result is expressed as liters per kilowatt-hour and lower numbers are better. It exists because PUE alone can mislead: a facility can achieve an excellent PUE by using evaporative cooling, which rejects heat by evaporating large volumes of water, effectively trading electricity savings for water consumption. WUE makes that tradeoff visible. Like PUE, it is a ratio that describes efficiency rather than total impact, so a very efficient facility can still use a great deal of water in absolute terms if it is large. Operators report both metrics together when they want an honest picture of resource use.
Q: What Does N+1 Redundancy Mean?
N+1 is a shorthand for how much spare capacity a system carries. N is the number of components needed to handle the full load, and the plus one is a single extra component that can take over if any one of the others fails. A data center needing three UPS modules to carry its load, for example, would install four under an N+1 design. The notation extends naturally: 2N means a complete duplicate of everything, so the facility can lose an entire independent system and keep running, while 2(N+1) doubles an already redundant arrangement. Designers apply these formulas to generators, UPS units, cooling pumps, and network paths, choosing the level based on how much downtime each workload can tolerate. Higher redundancy raises both construction cost and the amount of equipment sitting idle, which is why not every system in a building carries the same rating.
Q: What Is a Modular Data Center?
A modular data center is built from prefabricated units, often the size of shipping containers, that arrive with racks, power distribution, and cooling already installed and tested. Instead of constructing a bespoke building over many months, the operator prepares a pad with power and network connections, then places the modules and links them together. The approach suits situations where speed matters, capacity needs are uncertain, or the location is remote: a mining operation, a military deployment, or an edge site serving one metropolitan area can all get computing capacity in weeks rather than years. Modules can also be added incrementally as demand grows, which avoids paying for an oversized building on day one. The tradeoff is density and efficiency; a purpose-built hyperscale campus still beats modular designs on power effectiveness at very large scale.
Q: Why Do Some Data Centers Sit Underground or Underwater?
Unusual locations solve specific problems. Underground sites, such as converted mines or bunkers, offer physical security, protection from weather, and naturally stable temperatures that reduce cooling work. Underwater placement goes further: Microsoft’s Project Natick, which ran from 2014 to 2020, sealed servers in a pressure vessel and submerged it off Orkney, Scotland in 2018, retrieving it in 2020. The ocean provided free, constant cooling and an environment free of oxygen corrosion and human interference, and the experiment reported fewer hardware failures than a comparable land facility. The drawbacks are access and scale. Servicing submerged or deep-underground equipment is slow and expensive, expansion is constrained by the site, and any serious failure can mean retrieving the whole installation. For those reasons, exotic locations remain experiments or niche deployments rather than the industry’s standard pattern.
Q: What Is a Carrier Hotel?
A carrier hotel is a large building, usually in a major metropolitan area, that houses an exceptional concentration of telecommunications carriers, internet service providers, and interconnection infrastructure under one roof. The term comes from the early days of telecom, when long-distance carriers needed neutral ground to hand traffic to one another. A carrier hotel earns its name by density rather than size: dozens of networks terminate fiber there, internet exchange points operate inside, and meet-me rooms provide the cross connects. Being inside one gives a network cheap, short paths to many potential partners, which is why cloud providers, content companies, and financial firms lease space in them even when they keep their main computing elsewhere. The building’s value lies in who else is in it, so owners compete on neutrality and on the richness of the interconnection ecosystem.
Q: How Do Data Centers Suppress Fires Without Damaging Servers?
Water sprinklers would save the building and destroy the servers, so data centers use clean agent suppression instead. These are gases, such as FM-200, Novec 1230, or inert blends like Inergen, that extinguish fire by absorbing heat or displacing oxygen without leaving residue on electronics. Detection comes first: very early smoke detection systems continuously sample air through pipes and can sense combustion particles long before a visible flame appears, giving operators time to investigate. If a fire is confirmed, the system floods the affected zone with gas, often after a short delay that lets personnel evacuate. The room must be reasonably sealed for the gas to hold its concentration long enough to work, which is one reason data halls are built tight. After discharge, the gas ventilates away and equipment can often return to service quickly.
Q: What Is the Difference Between Colocation and Managed Hosting?
The difference is who owns and operates the servers. In colocation, the customer owns the hardware and rents space, power, cooling, and network connectivity inside someone else’s facility; the customer’s own staff, or a contractor, installs the servers and manages the software. In managed hosting, the provider owns the servers and rents them as a service, handling hardware maintenance, operating system patching, and monitoring, while the customer manages only its applications. Colocation suits organizations that want control over their hardware and already have operations staff, trading capital expense for facility quality. Managed hosting suits those that want computing without a hardware team, trading some control and higher monthly cost for convenience. Both differ from hyperscale cloud, where resources are virtual and billed by the minute rather than tied to specific physical machines.