A rack of Nvidia's GB200 NVL72 is rated somewhere between 120 and 140 kilowatts, and nearly all of that power leaves the building as heat. The colocation halls most companies rent were designed for 10 to 12 kilowatts per rack. I write agent software, and in that work a model is a function that costs money and answers in a second or two. Nothing in my code mentions watts, which is exactly why those numbers caught my attention.

These figures come from vendor blogs and trade press, so read them as estimates, and rack numbers shift with configuration. Schneider Electric reports that average rack density went from about 16 kilowatts in 2025 to 27 in 2026, while the AI deployments operators are being asked for sit at 50 to 70. Air cooling runs out somewhere around 80 to 100 kilowatts per rack. That is the reason direct-to-chip liquid loops stopped being exotic and became the default for these machines.

Start from the chip and it makes sense. A B200 GPU is listed at up to 1,200 watts, which is space heater territory, and one NVL72 rack carries 72 of them next to the CPUs, switches and power conversion. Nobody can move that much heat with a fan and good intentions. The coolant runs in pipes to the chip itself, and the facility needs coolant distribution units, a heat rejection loop and a power feed that did not exist in the original building plan. Air handlers and raised floors were built for a different kind of load, and no amount of tuning stretches them to cover this one.

The next generation raises the stakes. Digitimes quotes a cooling vendor executive saying Vera Rubin NVL72 racks will need 130 to 140 kilowatts of cooling capacity and that requirements are doubling in under a year. Schneider's reference design for the same platform scales up to 227 kilowatts, and Vera Rubin Ultra is projected at 200 to 300 kilowatts per rack. Those last numbers are projections from secondary reporting, not Nvidia's specifications.

None of this reaches my code as a temperature. It arrives as a price, a rate limit, or a region that is simply full this afternoon. So I treat model capacity as scarce infrastructure instead of an API that always answers. In kognios the fallback chain runs across nine providers and tells you which model actually answered, and I added that after tracing bad output in one of my pipelines to a free-tier model quietly serving almost every request. I only found it because I went looking.

The tool error guard in kognios comes from the same instinct. An agent gives up after three failures from the same broken tool instead of looping until the bill arrives. Every pass through that loop is an inference call on a chip that draws close to its rated power under load, whether or not the answer is any use. A runaway retry loop shows up as a cost problem first, and at the volume these systems run it is also a heat problem for somebody else's cooling plant.

That changes what I count as sloppy. Pointing the largest model at every step of a workflow is the software version of cooling a server closet with the building's chiller. Routing the simple classification steps to a small model, caching what has already been asked, and stopping early when the last two attempts agree are all dull choices. They are also the choices that decide how many of those racks you need to rent. None of them show up in a demo, and all of them show up on the invoice.

In regulated work the pattern around outages is predictable. The first question after a provider has a bad hour is rarely about the outage itself. It is about what the fallback did, which model handled the customer's data while the primary was down, and whether anyone can prove it from logs. A chain that records the answering model per request makes that question cheap, and a chain that does not makes it a week of archaeology.

I started this piece from rack numbers I cannot verify myself, and the Rubin figures in particular are reported rather than specified, so I would check Nvidia's launch material before quoting any of them in a design review. What I can check today is my own fallback chain, which writes down the model that answered every request.