Giuseppe Bernacchia, distinguished engineer at Infineon
Executive interview · Video
Free to watch

Infineon on protecting the 800 V rack: fast enough to save it, slow enough not to trip it

Executive interview · Video
Free to watch

Infineon on protecting the 800 V rack: fast enough to save it, slow enough not to trip it

Short a bus in an industrial plant and the long cable feeding it buys you time. Its inductance slows the rise of the fault current, and a mechanical breaker that needs milliseconds to open gets its milliseconds. An AI rack offers no such favor. The cables are short, because every extra meter is parasitic inductance and I²R loss, and the bus sits at 400 V or 800 V DC. Put high voltage across almost no inductance and the current heads for a few kiloamps fast enough that customers ask for protection that reacts in microseconds. Give it milliseconds, the way a relay does, and Giuseppe Bernacchia’s estimate is hundreds of kiloamps. “And then you destroy everything.”

Bernacchia is a distinguished engineer at Infineon and the system architect for its AI data center work, a brief that runs from the power entry of the IT rack down to the voltage regulator controllers feeding the GPUs. He has spent 25 years at the company. The last couple have gone to high-voltage systems, protection and DC/DC conversion alike, and what he has found is that the big end customers want the protection behavior they have always had at 12 V or 50 V, scaled up to 800 V. The physics does not cooperate.

The trouble starts inside the semiconductor. A protection switch on an 800 V bus has to survive the fault, which means safe operating area (SOA), and it has to carry the full load current with low on-resistance, which means a small, dense die. Those two requirements pull in opposite directions, because packing the cells closer to cut RDS(on) is exactly what erodes SOA. Infineon’s answer is silicon carbide, which Bernacchia calls the best fit for the job, and in particular the SiC JFET, a simpler device than a SiC MOSFET.

Then come the system problems. A hot-swap controller that trips on a GPU training transient shuts down a system that was never in danger, so Infineon’s digital controllers add glitch timers and filters that let a pulse of a couple of times nominal current pass if it ends quickly enough. A solid-state breaker switches in microseconds, but so far the power devices fail closed, so human-safety standards still call for a mechanical relay and its air gap. And a megawatt rack will need tens of devices in parallel, where a few hundred nanoseconds of turn-off skew can leave one device carrying most of the current.

The Data Center Engineer sat down with Bernacchia ahead of the OCP Global Summit in San Jose to talk about why SOA and RDS(on) fight each other at 800 V, how to tell a training transient from a short, where solid-state protection still loses to a relay, and his two hot topics for high-voltage DC: paralleling tens of devices, and sensing current across a huge dynamic range. It continues Infineon’s run in this series, which began with the case for shifting to DC grids.

The conversation has been lightly edited for length and clarity.

AI data centers are moving to higher power levels and new distribution architectures. How are the requirements for electrical protection changing as a result?

Giuseppe Bernacchia: On one side, the big end customers would like to keep the requirements pretty much as they are used to having for the 50 V or the 12 V legacy. Moving to high-voltage bus distribution brings a lot of different challenges. The first is human safety. You have 800 V or 400 V, so if you touch one of the cables, the risk is that you get electrocuted. But that is solved by the end customers at system level on their own.

If you look at the components Infineon is providing, some of the requirements are really challenging in terms of the power devices. Currents are supposed to be lower, because that’s the whole purpose of moving from 50 V to high voltage. Nevertheless, the combination of a high voltage even with low currents, and we’re talking about a few amps, is really destructive for the devices. So having suitable devices available is becoming an issue. But in general, customers would like to keep more or less what they had and simply scale it up.

What are the biggest protection challenges engineers are encountering as power density and available fault current continue to increase?

Optimizing SOA normally contradicts an RDS(on) improvement, in particular the figure of merit of RDS(on) times area.

Bernacchia: When we look at the semiconductors, the main challenge is the SOA capability of the devices. This could be very high current for a very short time, full protection, for things like circuit breakers up front in the feeder lines. Or it could be inrush current control and short-circuit protection for hot swaps on the tray itself. The combination of high voltage and these current requirements is quite stressful for the devices. There are not many devices which can support that. Infineon is very strong with silicon carbide technology, which we think is the best fit for this kind of application because it’s a very rugged technology and offers really amazing performance, in particular the JFET, which is a simple device compared to a more complex silicon carbide MOSFET.

SOA is the first thing, but the second point is that these devices also need to provide very low RDS(on), because we still need to support very high power levels. These devices are in the current path, and you want low RDS(on) to avoid high power dissipation. Typically you have a certain requirement in efficiency, even when the power goes up. These two things together are very difficult to achieve at the same time, because optimizing SOA normally contradicts an RDS(on) improvement, in particular the figure of merit of RDS(on) times area. When you optimize for RDS(on), typically you shrink the size of the device, in the sense that the cells making up the power device become closer and closer to reduce the path and so reduce the losses. And this goes totally against the SOA capability.

Then there is telemetry. This has not changed from the legacy systems, but now you need to do sensing at very high voltage. This requires area, and accuracy becomes more difficult to achieve. And if you want density, you still need to deal with the clearances and creepages defined by the high-voltage standards. At the same time, the packages used for these power devices are big, because you need mass to do a proper job in dissipation when you handle the faults or the inrush current control. But we have very tight requirements in height or in area.

Protection has often been treated as something added around the power architecture. Is it becoming more important to consider it part of the architecture from the very beginning?

Some customers at least tend to underestimate the role of protection.

Bernacchia: My experience in the past couple of years, since we started looking at these high-voltage systems, is that some customers at least tend to underestimate the role of protection. In the past this has been relatively simple. At 50 V or 12 V, you put a controller and a MOSFET and pretty much it works. But now more people are realizing that the design of these protection systems, more than devices, is quite challenging.

Because we have protection as well as the next stage, the DC/DC conversion following it, we see that starting the analysis soon is absolutely important, in particular when you look at the combined behavior during faults or during startup. Take the hot swap at the power entry. Its function is to charge all the capacitors in the system before the system starts. These devices have very limited SOA capability, which means when you plug the system into the bus, the devices see very high voltage, and the current they can provide is very limited, so it takes time. If you start the DC/DC converter too soon, you are taking away most of the current these devices are providing for charging the capacitors. You start charging and then you discharge, start charging and discharge. So you need to properly sequence the startup, and if you consider both of them together, this becomes much easier.

Another example is EMI. In high voltage we need to fulfill EMI requirements, and that means EMI filters, which are essentially inductances and capacitors. When you have transients caused by the GPU being activated, the transient causes oscillations in those filters. If you don’t consider this properly by looking at the system as a whole, you may have a setup for the protection which is absolutely not fitting the requirements. You set an overcurrent protection level at some point, but because the EMI filter is reacting, you may have a current which goes beyond that, and you start stopping the system when you should not.

We are providing customers with a reference design where they have all the pieces together, while most of the time they see one block at a time: the hot-swap block, the DC/DC block, and so on. Looking at the whole system up front is absolutely key to proper behavior in different conditions.

Fault detection and isolation speed are becoming more important. Why does response time matter so much in these next-generation power systems?

Milliseconds would mean currents going to hundreds of kiloamps, and then you destroy everything.

Bernacchia: Consider standard circuit breakers in industrial applications. There you have long cables connecting one part of the equipment to the source, and these inductances slow down the currents in case of a fault. In this system, everything is super packed. The cables are very short, because a long cable means not only parasitic inductance but also I²R losses. Small inductances and high voltage are a really destructive combination. You apply a very high voltage on an inductance, and this creates a very high di/dt. The current is increasing to extremely high levels, in the range of a few kiloamps, extremely fast. That’s why fast detection of the fault is absolutely key. Normally customers are requesting reaction in the order of microseconds, while standard mechanical relays typically take milliseconds or more to react. Milliseconds would mean currents going to hundreds of kiloamps, and then you destroy everything.

Being able to handle extremely high currents means the caps need to be huge to absorb the energy, and the cables need to be thick. This is all against density, and in the end also cost. We all know that the important thing in these systems is to pack together as many GPUs as possible. Space is absolutely king in this application.

Engineers have to balance fast protection against nuisance trips and overall system availability. How do they do that?

Bernacchia: The GPUs experience incredible current transients when they go into training, really multiple times the nominal power these GPUs are rated for. You have a switch which is protecting your system. It’s normally on, just carrying the current as it should, and all of a sudden the current it sees goes from a nominal level to a couple of times that level very, very fast. What the controller does is try to detect these conditions and shut down. We have overcurrent protection limits exactly to save the external load. But if you open the device, you stop the transient, and essentially the whole system stops. You don’t want to do that.

That’s where Infineon, with our protection controllers, is really providing state-of-the-art in protection. We have our family of digital hot-swap controllers, and thanks to the digital control we can set different thresholds, different timings, and different filtering schemes. These incredibly fast transients at the beginning of training normally last a very short time. On the other side, you may want a tripping level for the overcurrent condition that saves the load, which is just above the nominal power level. We can set filters or timers so that the controller lets these very fast pulses go through without any action, because it recognizes them as transients. But if the overcurrent condition lasts for more than a prefixed time, which the end customers can pre-program, the controller takes action, flags a fault, and shuts down everything. Short timings typically don’t create such a big harm, but prolonged overcurrent conditions damage everything. That’s where these filters and glitch timers are extremely important.

We can also tune the reaction of the power device. A very fast shutdown typically creates very high overshoots at the output, and you want to avoid that unless it’s a destructive condition. And there are conditions like output undervoltage. When the output voltage sags, is that because the power device is deteriorating? Or is it a load transient? Here we monitor the power device health and report these conditions to the main controller so it can take action.

Where do solid-state approaches provide advantages over electromechanical protection, and what trade-offs still need to be considered?

Until now at least, the power devices normally fail on close. If the device is failing, you are not isolating at all.

Bernacchia: Solid-state solutions are very important for speed. You can react in microseconds instead of milliseconds. Also, you don’t have mechanical parts, so the wear and tear is far less. Reliability can be extremely high, in particular with silicon carbide MOSFETs.

Where is the trade-off? In devices like circuit breakers, and for human safety. There you still need an air gap, and you still need galvanic isolation. Until now at least, the power devices normally fail on close. If the device is failing, you are not isolating at all. You’re creating a short between input and output, and this is extremely dangerous. That’s why for human safety the standards foresee a mechanical relay anyway, which provides galvanic isolation. And the standards foresee a certain volume for these devices exactly to provide the safety, and unfortunately this is extremely difficult to reduce.

What we are trying from Infineon’s point of view is to work with the standardization committees, because these standards were made first and foremost for industrial applications, where you can have people running in the fab who are able to touch parts at high voltage. In the data center it’s a closed environment. Only certified people typically can access it, and there are guarding mechanisms like safe-touch designs for the connectors. So we’re trying to see whether there is a way to adapt the standards to data centers. But this is a very long process, because when human safety is involved it’s a very big topic.

There are still places where standard mechanical relays are used as a bypass, and high-power resistors are used to pre-charge instead of using solid state. Solid state provides a much, much smaller solution in terms of volume, so it strongly depends where you want to put it. If the footprint is not really a big deal, those standard solutions offer a very standard way to address the topic, a standard UPS pre-charging circuit, for example. But on the trays, where you want to use most of the size for the DC/DC converter and not for protection, which is still a necessary nuisance for the customer, solid-state solutions become really interesting.

Architectures are already moving to higher-voltage DC distribution and DC microgrids. What new protection challenges do those systems introduce?

The protection is really the backbone of high-voltage DC.

Bernacchia: Honestly speaking, this is a bit outside my sphere of competence, because I’m not working with the microgrids. But definitely the high-voltage DC bus architecture poses a lot of challenges, in particular on the protection. The protection is really the backbone of high-voltage DC. Human safety, but also conditions like arcing in the system when you unplug things.

And the increasing power is making the systems humongous. If you need to handle kiloamps of current, it’s not one single semiconductor device. You need many of them in parallel. How do you synchronize the turn-on and turn-off? If you have an overcurrent condition and a misalignment in the turn-off time, not nanoseconds but some hundreds of nanoseconds, which does occur in the system, then the risk is that one device is carrying most of the current, and that blows up. One or two devices in parallel can still do the job at lower currents. But if you look at the megawatt racks we plan in the future, we definitely need tens of devices working in parallel.

The other hot topic is telemetry. Customers are implementing power capping, using the telemetry from different places in the system to understand when the GPUs are firing and trying to distribute the loads so that you don’t overload the grid. But if you need to handle a huge dynamic range for the current, how do you achieve the required accuracy without exploding the size? You need to sense current in a high-voltage domain and then send the telemetry to a different ground domain. This means isolation, again creepage, which means area. High-voltage distribution lives thanks to the protection devices, for safety and for protection in the proper sense, to make sure nothing blows up when you operate.

How important is coordination between power devices at different levels of the architecture, from facility distribution down to the rack and the processor?

Bernacchia: To some extent yes, and to some extent no. From a system point of view, if you can detect a fault in the GPU and communicate upstream to the protection systems, you can preemptively address the fault without it propagating throughout the system. But this is more difficult to do when you consider how many elements are in the system. In a standard rack you may have 18 trays with the processors, then network switches of different types. Then you go up and you have the power distribution unit, which is a collection of these circuit breakers, and then perhaps, in the future, a solid-state transformer and the microgrid. How much wiring do you need, just for synchronization between the different stages? We already know that in the IT rack, cabling is becoming a nightmare. Adding extra cabling just for the sake of communicating backwards is a little bit difficult.

Since I’m working on the rack topics, we see some good opportunities in synchronization across the rack, between the GPU, the VR, the DC/DC, and the hot swap. Very simply, the VR knows very quickly when the GPU is undergoing a transient. The VR can immediately communicate to the DC/DC converter, what we call the intermediate bus converter, to say, “Be careful, there is a transient happening, so do something to address that.” The same at the hot swap. Instead of tuning the knobs all the time because we don’t know when the transients are happening, if I know the transient is happening and I can recognize its amplitude, I can tell my controller, “Look, don’t trip overcurrent, because the transient is happening.”

We are trying to do that to some extent on the tray, because as Infineon we have the possibility to build the entire system from power entry to the VR, so we are exploring these possibilities. Going outside the tray is partly possible, because customers are also looking at power capping across multiple racks, to avoid firing all the racks at the same time, which would essentially collapse the grid. Of course, they are dealing now with the more urgent topic, which is making the system work. I believe this is an area that will expand as systems become more stable in operation.

What is the biggest opportunity for innovation for a system architect in this space?

Bernacchia: Silicon carbide is the material of choice for this. We, as well as some of our competitors, offer JFETs or silicon carbide MOSFETs, so there is definitely room for improvement in pure technology development. In a sense this is the low-hanging fruit. It’s not easy, but it’s the first thing somebody can think about.

But there are nice opportunities at the control level, in the control of the power devices. How to turn off in a proper way to avoid overshoots. How to parallel multiple devices properly and synchronize them. If the input is collapsing, can you prevent the output of the protection from collapsing as well, or at least prolong the hold-up time so the system can shut down properly and not abruptly? Vice versa, if you have an input overvoltage condition, can you prevent the rest of the system from seeing it? We are developing new dedicated controllers for this high voltage, and we are trying to put in all the learnings we had from the past two years working in this area.

Without getting ahead of any announcements, what should engineers watch for from Infineon at OCP in next-generation protection?

Bernacchia: One of the biggest topics will be related to our controllers, our next generation. I think there will be some communication about that. Further out in time, there will very likely be communication about our power technologies.

For the first time, Infineon has a booth at OCP. I encourage everybody to visit. On protection, we will definitely have some of the reference designs we have for the hot swap, so for the power entry of the IT trays. We have collected a lot of experience designing for the different grounding schemes and configurations of the hot swap. We should be able to showcase some new-generation circuit breakers and solutions for the high voltage, though I’m not sure yet, because we are still in discussion with the colleagues dealing with that. But I think it will be interesting to see how the different pieces are coming together, and our customers will see that we really offer something across the whole distribution.

Unlock the Documents and Reserve Your Panel Seat

This interview is free to watch. One free registration unlocks all three documents and reserves your seat at the live expert panel on Thursday, October 29, 2026, 12:00 PM EDT.

Data center selection guide41 pages
The future of powering AI24 pages
SST whitepaper5 pages
Expert panel seatThu, Oct 29, 2026

We’ll email you a confirmation link. The three documents and your panel seat unlock as soon as you confirm.
Infineon

This resource hub is sponsored by Infineon Technologies. Documents are provided by Infineon; interviews are produced by The Data Center Engineer.

Back to the hub