Chroma — leaderboard

Batteries in the AI factory: from spare tire to shock absorber

The GPUs in an AI training cluster don’t draw power the way a hall full of cloud servers does. They wake at the same time, run the same operation, pause, hand their results to the next set of chips, and do it again. Thousands of processors moving in lockstep can demand more power than the utility can provide, and the surge lasts anywhere from microseconds to a few seconds.

Shahar Noy is vice president and general manager of the data center business unit at QuantumScape, a group chartered with getting the company’s solid-state lithium-metal cells out of the car and into the rack. He uses an automotive analogy to describe what a data center power profile does to the battery. The battery backup unit used to be the spare tire: insurance against an event that might never happen. In an AI factory, it becomes part of the suspension, absorbing every bump so the GPU never feels the road.

The second shift is physical. As the industry moves from 50 V distribution toward 800 V, conversion stages are being deleted from the power path, and the battery migrates from the front of the facility, just behind the utility feed, to the rack itself. Noy points to a SemiAnalysis estimate that cutting those stages saves roughly 5 percent of a data center’s power consumption, matching NVIDIA’s own figure. On a gigawatt-scale build he puts that at tens of millions of dollars a year, or hundreds of millions over the life of the building.

We sat down with Noy for The Data Center Engineer to talk about what breaks in the old electrical architecture at gigawatt scale, what “last meter” power delivery actually means, why buyers keep asking about exactly two battery attributes, and what still stands between a working cell and a shipping system.

Watch the full interview

The conversation has been lightly edited for length and clarity.

QuantumScape has been building a battery for cars since 2010. How did it end up in a data center?

Shahar Noy: QuantumScape was established in 2010, and the first application for the solid-state battery was automotive. We worked on changing the chemistry. What was known as lithium ion, how do you rethink lithium ion and introduce a battery which is denser, safer, can charge fairly quick, and make EVs more accessible to people?

Now, when we see the growth of data centers, we start thinking: why not use all of those great attributes there? If you look into how energy is being delivered, stored, and used in a data center, each and every one of those steps is being disrupted right now, and a solid-state battery can be one of the critical components to enable better energy delivery inside.

What actually makes an AI factory different from a traditional data center?

Noy: The AI factory needs to generate intelligence. Intelligence is measured in tokens, and tokens come out of what we call today GPUs. When you look into the history of data centers, the last decade was all about cloud compute. In a cloud compute use case, you don’t generate intelligence, you make transactions. You process data, you move data, you store data. Moving into more complex intelligence creation requires much more compute, much denser compute, working in tandem with other compute.

We used to build facilities of 10 to 20 megawatts, and that would be a very, very big number. Now we’re talking in gigawatts. If you read the announcements, Meta as an example wants to expand one of their facilities into five gigawatts or even 10 gigawatts. So all of this power demand is an order of magnitude bigger than what we’ve seen only up to five years ago, and it essentially disrupts the whole market.

Is the electrical bottleneck an engineering problem or a strategic one?

Noy: Both. If you’re moving into a one gigawatt facility and you can slash 1 percent of your energy cost, just by having a more efficient energy flow, we’re speaking of 10 megawatts. You can calculate it based on different utility prices, and you’re in the millions of dollars of saving per company.

We are a small company. We didn’t come up with the 800 V transition innovation. If you go and follow NVIDIA announcements and most recently the SemiAnalysis report, the shift from 50 V to 800 V is a combination of a strategic effort and also an engineering effort. If I can bump up the voltage by 16x, and let’s say the power per chip goes up 2x, I can still save on my current by 8x. And if I save on the current, I can use less copper wire, which is expensive, and copper wires are also not highly efficient in the way they transmit electrons.

Which parts of today’s electrical architecture simply don’t scale from 20 MW to a gigawatt?

Noy: It’s almost all of the above, because think about how we got to this point. We always had to take an AC input in, and make the first conversion to DC that goes through a traditional UPS. The UPS, other than just giving you energy backup, would also help with shaping the AC, making it cleaner, putting it back out. Then this AC would go and transform into what we call regulatory needs. In Europe you can be 200 V plus of AC; in the US it’s 110 V AC. And then there was another step down to a 48 or 50-ish volt, which is considered to be safe. It has some roots in how telcos used to support voltage, and also some safety to be under 60 V. Then we took it down to 12 V, and you say, why do we have 12 V? Because we’re using motors inside the data center, fans, hard disk drives.

Now fast-forward to the generation of AI. Why do you need 50 V? We proved in cars that we can build 800 V platforms. They’re out there, they’ve been deployed, and now architects inside data centers are taking advantage of this. And 12 V, do you really need to step down to a spinning device? I don’t think we use as much HDDs or even fans inside those compute racks today. We’re moving to liquid cooling, flash storage. So a lot of those steps are being eliminated.

That leads us into what we believe is the biggest challenge, which is the last meter of power delivery. If we’re truly shifting into 800 V, you need to put that voltage really close to the GPUs, and rethink how it gets distributed, how much power density you can pack around them, and how you deliver it safely.

So what do AI power profiles look like, and why are they harder to support?

Noy: AI workloads rely on multiple GPUs working at the same time. When you train a model, all of those GPUs wake up together, they do a similar transaction, have to wait for a second, make some sort of decision, send this data to another set of GPUs, change weights, and then do this whole transaction again. So it’s a very synchronous type of workload.

Because all of those GPUs by definition are more power hungry, you have more of them working at any given time. Then when they wake up, it can demand more power than what the utility company can provide to you. This is momentarily, not for multiple minutes, but microseconds to a few seconds. So this introduced an interesting challenge, because batteries became the shock absorbers of this phenomenon. To protect the utility company from those spikes, the battery backup units are providing this extra energy momentarily so the GPUs can continue to work at peak performance.

If you can make batteries more resilient, they’re no longer just a backup unit. To your point, this is where they become an active unit. So think about it like you have a car. In your car you have a spare tire. You will use this tire only if there is an incident, like a flat tire or an accident. This was the concept. But now batteries inside the data center actually need to mitigate all of those energy spikes. So from a spare tire, you actually become the shock absorber of the suspension. You need to constantly deliver quality of drive to the driver. Now you need quality of power constantly delivered to the GPU.

Tell us more about last meter power delivery. What does it mean in practice?

Noy: We used to build data centers in separate ways. The utility room got designed by the data center owner, and someone else designed the compute rack. But now, because data centers are becoming so expensive, we see a modular approach, which means every rack is being built as a self-sustainable unit. If you want a GPU rack, or what we call an IT rack, fully independent, it needs to have its own power supply and its own battery attached to it.

In the past we also used to over-provision. The UPS area, the gray space of the data center, would account for certain growth aspects. But now you can refresh those GPUs on an annual cadence. If you see the NVIDIA roadmap, it’s crazy, refreshing every 12 to 18 months. You cannot truly say how much UPS you would need.

The H100, one of the first popular GPUs that started this whole AI revolution, was order of tens of kilowatts per rack. Now we’re shifting into hundreds of kilowatts. The Grace version that they’re deploying right now is around 200 kW. As we move further along, we are projecting those racks to be 600 kW of power needs. So when you design your power and energy to fit exactly the needs of the rack, you’re creating a more optimal solution.

Why are power conversion stages such a problem?

Noy: The steps of voltage are basically legacy. SemiAnalysis did an amazing analysis where the transition to 800 V, cutting a lot of the transitions along the way, saves 5 percent on average from data center power consumption. Multiply that by a couple of years and we’re going into savings of either tens of millions of dollars per year to hundreds of millions over the life of a gigawatt type of data center.

So all of this requires rethinking how 800 V can really be transitioned into the rack without too many steps. Which means that if the battery used to be at the forefront of the data center, the first step after the AC comes in, now it can move further inside, so you get the exact size of energy you need as you build your data center rack by rack.

What is it about solid-state lithium metal that suits AI infrastructure? Density, response time, safety, cycle life?

Noy: The more we talk to partners, to power system ODMs, to data center builders, it seems like the energy density is becoming a key, key factor. We used to live in the 50 V area, and for 50 V we had certain battery modules, we call them BBUs. Now, as we move to 800 V, you need to multiply it by 16. Put that many more traditional cylindrical batteries inside and you slowly start running out of space. Instead of having space for those GPUs, now you need more and more space for batteries. We announced our first QSE-5 as above 800 watt-hours per liter, and we continue to push the envelope. This is super compelling to data centers, because they can now use less space to store energy than they anticipated. So that’s item number one.

Item number two is safety. We read about fires in data centers once in a while. When we design safety into solid-state batteries, the first premise, this goes into EV. EVs drive people, drive families, we need to make it safe. You can argue data centers are autonomous, who cares about safety? But depending on what market research you follow, Bernstein, Gartner, even NVIDIA made some argument in their earnings calls that a single gigawatt costs 50 billion dollars to deploy. So you want a safer environment to protect your CapEx. As those data centers get denser and higher in power and now higher in voltage, the risk profile continues to grow.

You talk about intelligence density. What does that mean from an engineering perspective?

Noy: We’re used to thinking about compute as a single chip. You have one chip in your iPhone. You used to have one chip in a server, and in the times of virtualization multiple people could connect to a single processor and do a lot of work. But at the age of AI, compute is measured at a rack level. NVIDIA has NVL72 moving into NVL144, moving into NVL576, which means how many compute dies they can put in a rack. In their definition, the denser the rack, the more performance they can provide to you.

Those racks are becoming more power hungry and warmer, so they need a liquid cooling rack next to them, and a power and energy rack next to them. If we as an industry can compress the cooling racks and the energy racks, we leave more space to those GPU clusters to stay close to each other, reduce latency, provide you a high level of compute. Everything we can do to reduce the size of the peripherals that support those IT racks is a big plus.

How far away is this from production? Are there pilots running?

Noy: Unfortunately, due to NDAs with customers and partners, I cannot share too much. We did announce availability of our cells, and we did announce Eagle Line, our pilot production line. So as you can imagine, the components are out there. We’re in the process of finalizing system design around it, so hopefully you will see it in the market in the very near future.

What engineering challenges are still open before it’s commercially ready?

Noy: It’s the system. It’s no longer a cell question. Nominal cells, the industry knows, it’s public information, are around 4 V. You try to create 800 V out of them. So the system aspect, how you put them together, how you control them, how you balance them, how do you manage their thermal challenges, that’s the key next hurdle before those things go into production.

Beyond batteries, what else reshapes AI power infrastructure over the next decade?

Noy: Solid-state transformers, also part of the 800 V disruption, will come to the market. They’re fairly essential to make this full transition to DC-DC high voltage very efficient. I can see them arriving by 2030.

On the battery piece, as we completely eliminate the AC and get more efficient in how we deliver 800 V all the way to the GPU, the battery technology would need maybe to go into another step as voltage continues to go up, maybe above 1,000 V. And as those racks scale to a megawatt plus, you might see battery racks replacing the power rack.

What’s next for QuantumScape in AI infrastructure?

Noy: I call the AI infrastructure step one of a much bigger revolution. We’re in the first inning of AI as an agent, as an advisor, as a consultant. But there’s another thing happening in parallel, and it’s the physical AI world. From a battery perspective that’s interesting, because physical AI, when I talk robotics, drones, they need denser and safer energy as well. So we see even more opportunities for this solid-state battery to penetrate new markets.

Chroma — below posts

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

Chroma — tower

Popular Posts:

Quantumscape
Batteries in the AI factory: from spare tire to shock absorber
Johnson-Controls-has-launched-an-Absorption-Chiller-Reference-Design-Guide-aimed-at-data-centers-running-on-on-site-power-generation
Johnson Controls' new absorption chiller guide targets 44% lower cooling power in AI data centers
Why data center power validation is moving to real-time simulation
Why data center power validation is moving to real-time simulation
39D1CAB6-7D1C-4CA7-A2CF-D0919F6EA523_1_105_c
EdgeSites uses existing buildings to deploy 1 MW immersion-cooled AI compute
20260803015848EDT_image_1
Water quality sensors enable real-time monitoring for liquid-cooled AI data centers

Share Your Data Center Engineering News

Do you have a new product announcement, webinar, whitepaper, or article topic? 

TDK — tower ()

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

TDK — leaderboard ()