Surendhar Somasundaram, principal engineer at Infineon
Executive interview · Video
Free to watch

Vertical power delivery: Infineon’s VRM density roadmap and what it does to your thermals

Executive interview · Video
Free to watch

Vertical power delivery: Infineon’s VRM density roadmap and what it does to your thermals

Put a voltage regulator module directly under an AI accelerator and it inherits the accelerator’s problems before it switches a single pulse. The motherboard beneath a multi-kilowatt ASIC already sits at 80° to 85° C with the cold plate running. The silicon inside the module is rated to 125° C, but customers want it held under 110° C, ideally closer to 100° C. That leaves about 20 degrees of thermal headroom to absorb the module’s own losses, which double when a 90 mm² footprint carries four phases instead of two.

That is the bargain behind vertical power delivery (VPD), and Surendhar Somasundaram knows both sides of it. He is a principal engineer at Infineon in San Jose, working on the concept, definition, and development of the VRM power modules that sit in the shadow of the ASIC. The case for moving them there is strong: a vertical path through the motherboard cuts the power distribution network impedance to a quarter or a fifth of a lateral layout, which trims I²R loss and tightens the voltage gradient across a GPU die that can measure roughly 100 mm on a side.

The cost is space, and the space was already spoken for. The decoupling capacitors that carry an ASIC through a load step live under the processor too, and the modules have to fight them for the same square millimeters. Meanwhile the load steps are getting worse. Think 2,000 A in under a microsecond, and a few tens of nanohenries of inductance holding energy that has nowhere to go when the chip suddenly stops asking for it.

Somasundaram’s answer is to make the module do more of the capacitor’s job. A controller that briefly goes nonlinear during a step. Trans-inductor voltage regulator (TLVR) magnetics that couple the phases electrically so phase two reacts before its own control loop has noticed anything, cutting the required output capacitance to a third, or holding capacitance constant and shrinking the peak-to-peak voltage excursion to as little as 30% or 40% of a conventional design. Capacitors built into the module footprint. And an admission that density has a ceiling: “If you can’t cool them, then there is no point in having a high density.”

The Data Center Engineer sat down with Somasundaram ahead of the OCP Global Summit in San Jose to talk about the physical limits of lateral power delivery, why losing capacitance is an energy problem, how TLVR and coupled magnetics buy back transient response, and the solder-joint reliability question starting to shadow the density push.

The conversation has been lightly edited for length and clarity.

Most engineers know the basic advantages of VPD. As AI processors scale, what are the biggest implementation challenges they’re running into?

Surendhar Somasundaram: As the AI silicon scales to multi-kilowatt systems, we are talking about two primary challenges. Number one: how to deliver that power in the given space. The system is not actually growing. The ASIC is growing in size, which means higher currents are being asked to be delivered. Two, what comes with that is higher current density in a small space, which leads to a thermal challenge. So we are fundamentally dealing with a power density and a thermal flux problem. The major focus is vertical power delivery, but there are other architectures too. You can still do a backside lateral or a topside lateral, but each has its own challenges. Determining the right system architecture depends on what your critical trade-offs are, and it has to be a co-design with the power products in mind.

Is the concept well understood at this point, and it’s the implementation at these power levels that gets difficult?

Somasundaram: Up until the CPU generations, current levels were in the range of a few hundred amps. You have a motherboard, you have the ASIC sitting on top of it, and you have enough space around the ASIC for these multiphase power stages to sit on the periphery and feed power in. But when you’re talking about thousands of amps, you have a physical space constraint. Are you going to be able to put them around the ASIC?

Sure, you can. There are some GPU platforms where you would see an ASIC roughly 100 by 100 mm, and the power stages spread across the entirety of the motherboard, literally all the way to the edge of the board. But fundamentally you have an I²R problem. The closer the power stages are to the end solution, the smaller the power distribution network, the PDN. The impedance is smaller, which means less heat dissipated. Any heat you avoid dissipating is energy you can turn into compute.

That is the way VPD has evolved. You take the lateral solutions that sit next to the ASIC and put them underneath it, so the power goes from the module, through the motherboard, to the ASIC on the top side. You are cutting that impedance down to at least one-fourth or one-fifth of what it would have been with a lateral solution.

Processor currents keep rising while the ASICs themselves get larger. What does that combination mean for the power delivery architecture immediately around the processor?

We can’t treat the power as a single node element to those ASICs. You have multiple different nodes spread through the entirety of the silicon, and the silicon is getting bigger and bigger and the currents are getting larger and larger.

Somasundaram: The fundamental limitation for any of the GPU or TPU architects is that as the silicon gets bigger, they have to deal with the voltage gradient that comes with it. We can’t treat the power as a single node element to those ASICs. You have multiple different nodes spread through the entirety of the silicon, and the silicon is getting bigger and bigger and the currents are getting larger and larger. So they have to pay very close attention to how the voltage gradient is distributed across the whole silicon area.

Vertical power delivery lets you place the modules right underneath where you actually want them. You can distribute them through the entirety of the shadow of the ASIC and find the shortest path from the power solution to the die. In a lateral solution, the currents all come from the edge. The first point of entry is the perimeter, and the die sitting in the middle sees a gradient. That puts a limitation on compute efficiency and compute throughput.

On top of that, these ASICs have very high load transients, which means there is a high di/dt event being asked for. A higher PDN may result in missing the V-min or V-max specification for a certain process node. With vertical power, you distribute the power stages locally and cut down that PDN, so when there is a load transient, the effect of the PDN is very, very minimized. For the same solution, you get a much better peak-to-peak voltage response. Hence, for such high-current solutions, VPD is becoming mainstream in the industry.

As current rises you need more power stages in a very limited board area. What are the most difficult electrical trade-offs that creates for the designer?

If you push the density too high, you risk losing too much electrical performance, leading to thermal problems.

Somasundaram: Now we get into the trade-offs. There is no free lunch here. The fundamental trade-off is that there is only limited real estate underneath the ASIC. The way it has been approached is that the power stages sit around the ASIC, and all the decoupling capacitors required to supply the energy needed during the transients, and to absorb that energy during the load release, sit right underneath it. Now the modules are going underneath the ASIC too, so you’re fighting for the spot along with the capacitors.

What can you do about it? You can make the solution higher power density. But then you are dealing with smaller magnetics, which means you have to switch at a much higher frequency, which means slightly lower efficiency and lower electrical performance, indirectly leading to a thermal challenge. That’s number one.

Number two, you will be dealing with significantly less capacitance than in the default scenario, and the transients are not getting any easier. So how do you balance how many capacitors you need against what density of solution you need? If you push the density too high, you risk losing too much electrical performance, leading to thermal problems. That is the fundamental trade-off I see.

Why does losing available capacitance become such a problem when you’re dealing with the extremely fast load transients of AI accelerators?

Somasundaram: Fundamentally, it’s an energy problem. When a severe load step comes up, you need some local help for the system before the controller can react. There’s a limitation on how fast the control bandwidth can act and how fast the module can respond. There are some fancy techniques we are employing, but whatever we do, at the end of the day you need some capacitance next to the module, or through the power plane, to supplement the energy until you get help from the modules.

The energy in a capacitor is ½CV². How much energy is being demanded by the ASIC, and how much you can allow the voltage to dip before you get help from the control architecture and the module, is determined by that C value.

Likewise for load release. Let’s say 2,000 A is going into the ASIC, and suddenly the load says, “Stop. I don’t want any more of the energy.” All the energy already in the path of being delivered to the ASIC is dominated by inductance: the inductance in the VRM itself, plus the parasitic inductance of the PDN. So ½LI² gets added. It’s a few tens of nanohenries, and then I² is again the problem. You have to take that energy and dump it somewhere.

So during both the load step-up and the load release, you want an energy bank to supplement what the control architecture and the modules can do. You can look into the magnetic architectures, the control architectures, reducing the PDN path. All of these can be employed, but fundamentally you have to have a fair mix of capacitors next to the VRM to solve these transient problems.

So the faster the processor changes its demand, the more important capacitance becomes, while the space available is still shrinking.

If you’re asked to deliver 2,000 A in less than one microsecond, that is fundamentally different from 100 A in less than one microsecond.

Somasundaram: Correct. It’s the load demand, and also how steep the load demand is. If you’re asked to deliver 2,000 A in less than one microsecond, that is fundamentally different from 100 A in less than one microsecond. Your slew rate is different, and the energy needed to take on that one-microsecond, thousand-amp load step is significantly different. The new generations of processors are demanding more current in a shorter span. So we have a two-pronged problem: faster slew rate and an even higher load step.

How are engineers addressing the transient challenge when simply adding more capacitance is no longer practical?

Somasundaram: Solving that problem falls upon Infineon, and that’s what we do through our vertical power architectures, through the module and the magnetic design. It’s a multifaceted problem, so let me spell out a few of the techniques.

The first is a smart digital controller where we can enable some nonlinear behavior. You want your system to be linear, stable, and very predictable. But we can have a controller that briefly takes the whole system into a nonlinear mode if it can take advantage of that. If you fire the pulses more rapidly as the load step-up happens, you start providing energy before the voltage dips too deep.

Number two is the way we deal with magnetics. TLVR is one approach, which has been in the industry for the last decade. We are also looking into advanced coupled magnetics. Ideally, we want an inductor with a good steady-state inductance, but during the moments of transient, we want that inductance to go lower, so you get a much faster response from the VRM to your load demand.

Number three is capacitor integration within the module itself. We have products where we integrate a significant amount of input and output capacitance in the same footprint as the module, so it becomes a fully vertically integrated solution. We don’t use just one of these. We use a mix of them.

How does TLVR change the way the voltage regulator responds to large, fast current transients, and where does it provide the greatest value?

You can save on the required output capacitors, down to one-third of what you’d need without TLVR. Or in other words, with the same amount of capacitance, your peak-to-peak voltage can be as low as 30% or 40% of a traditional non-TLVR setup.

Somasundaram: Fundamentally, TLVR is a transformer architecture. The secondary windings of a multiphase regulator are daisy-chained to one another and terminated by a compensating inductor, and that secondary and compensating inductor adjust the impedances. At a high level, TLVR is an electrical coupling.

Say there is a load step-up, and phase one has seen it and started reacting. In a typical multiphase design, phase two and phase three come up next, but there is a wait before the other phases start contributing. In the case of TLVR, as soon as the first phase starts to react, that gets coupled into the secondary, and the secondary is all daisy-chained together. The incremental current gets reflected back into phase two, phase three, phase four, and so on. Even before the other phases react from a control standpoint, they are starting to react to the increase in load current through the information shared by the very first phase. That fundamentally opens up the bandwidth of the control loop.

We have seen in empirical data that you can look at the savings two ways. You can save on the required output capacitors, down to one-third of what you’d need without TLVR. Or in other words, with the same amount of capacitance, your peak-to-peak voltage can be as low as 30% or 40% of a traditional non-TLVR setup. We want as many modules as possible to support these high-current challenges, but we are fighting for space with the capacitors. If a magnetic architecture can help during the moments of transient, the required capacitance is lower in the first place. That’s exactly the benefit of TLVR.

There are other magnetic architectures as well, like negatively coupled magnetics, and we have some fancier concepts where we employ multiple of these together. We have a product that can do negative coupling and TLVR together, so you have the magnetic side of negative coupling and the electrical side of coupling from the TLVR, and we are seeing significant improvements in transient response.

Power density isn’t the only problem. What new thermal challenges appear as more power conversion is concentrated directly underneath or around the processor?

Without even having to turn on these modules, these modules by default start at 85° C.

Somasundaram: Say you’re asked to deliver 1,000 A in a given area. You can have 25 phases of 40 A each, or you can say, “I want 40 phases of 25 A each.” Those are two different things. With 40 phases of 25 A, your efficiency is going to be significantly better, because the current in each phase is only 25 A, and I²R is usually the dominant loss: through the FETs, through the magnetics, through the package. We have seen that 20 to 25 A is where you get peak efficiency. But the trade-off is that you have to accommodate 40 phases.

That’s why over the last four years there has been a constant push to increase the power density of VRM solutions. We started at 1 A per mm², today we are well north of 2 A per mm², and our goal is to get to 3, and in the next couple of years to 4 A per mm². The higher the module density, the more phases you can do in a small space.

Now, the other problem: on a per-phase level, your efficiency is more or less going to be the same regardless of the solution size. Four to five watts of loss per phase is a number we can say. Today we have two phases in 90 mm², and we also have four phases in 90 mm². In case one, I have 10 W of loss in two phases. In case two, I have 20 W in four phases. You’re looking at double the power dissipated in the same 90 mm².

On top of that, the motherboard is going to be very, very hot, because the ASIC is getting bigger and the power is getting bigger. All those kilowatts of compute end up as heat. Even with the cold plate in place, we see the motherboard already at 80 to 85° C. When these modules are sitting right underneath the shadow of the ASIC, without even having to turn on these modules, these modules by default start at 85° C. Then you add the 10 W or 20 W of loss within the module.

How do you take that heat out effectively? You don’t want to dump more heat into the motherboard and make that 85° go even higher. And you can’t let your module overheat, because we have silicon in there. Even though it’s rated for 125° C, we have seen customers want to operate below 110° C, ideally around 105° or 100° C. So we are dealing with about 20 degrees of margin at most, and we have to pull our own heat effectively from the module. We want the module to be a good electrical solution, but it also has to be a good thermal solution. If you can’t cool them, then there is no point in having a high density. It’s meaningless.

To make it even more interesting, lately I’m seeing many conversations around electromigration and solder reliability. The whole motherboard ecosystem relies on soldering. SAC 305 solder is still used prevalently, and that solder has a certain amount of Joule heating, and the elements within it have a long-term reliability impact when the temperature is too high and you’re pushing a lot of current through it. As an industry, we have seen concerns that as we push the density too high and the temperature gets too hot on the motherboard or within the module, we might get into long-term reliability problems, five to seven years out. So it’s a trade-off between the right power density, enough capacitance to meet the transients, and dealing with the thermals so that we don’t overheat it and cause long-term solder reliability issues.

When you’re balancing electrical performance, transient response, thermal management, and physical space, where are the biggest system-level trade-offs?

Somasundaram: The trade-off is pretty much across the board, and usually it’s a trade-off made by our customers. We have certain customers who don’t have that much in the way of transients. They have occasional transients, but mainly they deal with the highest currents running continuously. We have another part of the spectrum who have lots of transients, but their TDC, the thermal design current, which determines the heat, is not that much. And we have customers somewhere in between. It depends on whether you’re talking about top-of-rack ASIC switches, XPU ASICs, inference-based chips, or other networking chips. The workloads are fundamentally different, so each power architect has a different challenge. Overall, they deal with the same thing: I’m only given this much space, how do I maximize it? Power density versus thermals is a classic. Another one would be mechanical versus manufacturing risk. There are some new methodologies in the industry, but they are very complex to make, and we hear they are having some yield issues.

As current demand continues to climb, do you see VPD itself evolving, or is the bigger evolution going to come from the technologies and architecture around it?

Somasundaram: My take would be both. Today, when we talk about VPD, that means either these modules, two phases or four phases, go directly in the shadow of the ASIC, or there is an interposer on top of which the modules go. In the very near future, we see these types of modules evolving into a solution that goes into the motherboard, and then eventually into the substrate of the ASIC. The goal is very simple: you are cutting that PDN, the current path from the edge of the module to the ASIC, as low as possible, so the PDN losses go away.

But you are also moving next to the hotspot. You are getting closer and closer to the ASIC, which is the real hotspot, so the thermal and power density challenges might look different from what we are dealing with now. That’s an evolution I’m expecting to happen in the next two to five years.

On the other side, to fuel this evolution, the foundational pieces also have to evolve: the magnetics, the packaging, the capacitor integration, the ways to cool them, and even the silicon. Today’s silicon is optimized for less than 1 MHz. To get into the motherboard or into the ASIC, you have to be able to switch at a much higher frequency. That’s what the industry as a whole is working towards, and I see it coming through in the next two to five years.

What VPD innovations are you personally hoping to see at OCP, and what should attendees look for from Infineon?

Somasundaram: I come from a non-thermal background, but we have discussed quite a bit about thermal, so I always look forward to seeing what thermal innovation the industry is moving towards. I’m always fascinated by immersion cooling, even though there are some challenges to it. I would love to see whether it’s close to mass adoption.

Then, whether there is any forum where a common denominator is being achieved. Right now the industry is very fragmented, because each customer is dealing with their own problem. I would love to see technical committees coming together and making some sort of standardized approach to solving these vertical power and thermal problems.

As for the Infineon booth, we are excited to show our next-generation four-phase TLVR module, where we used some of the magnetic architectures like coupled magnetics. We are excited to demo it to people. It’s booth E26.

Unlock the Documents and Reserve Your Panel Seat

This interview is free to watch. One free registration unlocks all three documents and reserves your seat at the live expert panel on Thursday, October 29, 2026, 12:00 PM EDT.

Data center selection guide41 pages
The future of powering AI24 pages
SST whitepaper5 pages
Expert panel seatThu, Oct 29, 2026

We’ll email you a confirmation link. The three documents and your panel seat unlock as soon as you confirm.
Infineon

This resource hub is sponsored by Infineon Technologies. Documents are provided by Infineon; interviews are produced by The Data Center Engineer.

Back to the hub