Chroma — leaderboard

Runware ships 1 MW containerized AI inference pods with closed-loop cooling

Runware has introduced its Sonic Inference Pod, a containerized, modular AI inference “data center” built to bring GPU capacity online faster than a conventional facility build. The company describes each pod as a self-contained 1 MW deployment packaged in a 20-foot shipping container, aimed at model providers that need inference infrastructure without waiting on multi-year data center timelines.

Runware’s Sonic Inference Pod integrates custom servers, PCBs, and a closed-loop cooling system that the company says is engineered specifically for inference performance. Runware also claims a cost advantage versus traditional, gigawatt-scale builds, citing “facilities costs” up to 100x lower than a conventional data center build and saying it passes those savings through its GPU compute pricing.

What’s clear is the design choice: treat inference as its own workload class and build around it. For operators, containerized GPU infrastructure can change the planning conversation from “new building” to “how quickly can we secure power, place equipment, and operate it reliably,” but it also shifts the burden onto fleet operations, remote monitoring, and consistent thermal and power behavior across many small sites.

Cooling, siting, and deployment model

Runware describes the cooling approach as “closed-loop, water-free cooling,” with pods recirculating 1.5 m3 in a closed loop while holding GPU temperatures within 2° C of target conditions. The company contrasts that with evaporative cooling, and also says the pod format allows siting “wherever power exists,” including renewable generation sites such as solar parks. Runware says the pods are self-contained and don’t require changes to the power grid.

Runware also frames the pod format as a fit for “sovereign deployments,” saying pods can be placed within specific national borders or jurisdictions for organizations with data residency or regulatory requirements.

Scale-out plans and availability

Runware says it plans to deploy “dozens of pods” and “tens of thousands of GPUs” across the US and Europe in H2 2026 based on customer demand, and that this opens access to its custom-built hardware beyond its own platform for the first time. The company says Sonic Inference Pods are “live today” across the US and Europe and “ship ready to deploy.”

Runware says each pod joins a distributed inference network managed “as a single fabric,” with automatic rerouting of requests if a pod goes offline to optimize for latency, cost, and availability. Capacity is available through Runware’s serverless infrastructure, with Runware handling provisioning, scaling, and operations.

“So we redesigned from first principles: only the components inference needs, liquid-cooled GPUs, and a modular format we can place wherever power already exists,” said Flaviu Radulescu, co-founder and CEO of Runware.

Source: Runware

Chroma — below posts

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

Chroma — tower

Popular Posts:

Quantumscape
Batteries in the AI factory: from spare tire to shock absorber
20260803015848EDT_image_1
Water quality sensors enable real-time monitoring for liquid-cooled AI data centers
Emerson's new DeltaV automation platform unifies data center power and cooling control
Emerson's new DeltaV automation platform unifies data center power and cooling control
CPC-TRM-Tech-Guide-Liquid-Cooling-Reliability-(1)-4
In liquid cooling reliability, tubing choice creates another layer of protection
73
Tellus Energy AI model targets 20–40% lower data center power costs

Share Your Data Center Engineering News

Do you have a new product announcement, webinar, whitepaper, or article topic? 

TDK — tower ()

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

TDK — leaderboard ()