Runware has introduced its Sonic Inference Pod, a containerized, modular AI inference “data center” built to bring GPU capacity online faster than a conventional facility build. The company describes each pod as a self-contained 1 MW deployment packaged in a 20-foot shipping container, aimed at model providers that need inference infrastructure without waiting on multi-year data center timelines.
Runware’s Sonic Inference Pod integrates custom servers, PCBs, and a closed-loop cooling system that the company says is engineered specifically for inference performance. Runware also claims a cost advantage versus traditional, gigawatt-scale builds, citing “facilities costs” up to 100x lower than a conventional data center build and saying it passes those savings through its GPU compute pricing.
What’s clear is the design choice: treat inference as its own workload class and build around it. For operators, containerized GPU infrastructure can change the planning conversation from “new building” to “how quickly can we secure power, place equipment, and operate it reliably,” but it also shifts the burden onto fleet operations, remote monitoring, and consistent thermal and power behavior across many small sites.
Cooling, siting, and deployment model
Runware describes the cooling approach as “closed-loop, water-free cooling,” with pods recirculating 1.5 m3 in a closed loop while holding GPU temperatures within 2° C of target conditions. The company contrasts that with evaporative cooling, and also says the pod format allows siting “wherever power exists,” including renewable generation sites such as solar parks. Runware says the pods are self-contained and don’t require changes to the power grid.
Runware also frames the pod format as a fit for “sovereign deployments,” saying pods can be placed within specific national borders or jurisdictions for organizations with data residency or regulatory requirements.
Scale-out plans and availability
Runware says it plans to deploy “dozens of pods” and “tens of thousands of GPUs” across the US and Europe in H2 2026 based on customer demand, and that this opens access to its custom-built hardware beyond its own platform for the first time. The company says Sonic Inference Pods are “live today” across the US and Europe and “ship ready to deploy.”
Runware says each pod joins a distributed inference network managed “as a single fabric,” with automatic rerouting of requests if a pod goes offline to optimize for latency, cost, and availability. Capacity is available through Runware’s serverless infrastructure, with Runware handling provisioning, scaling, and operations.
“So we redesigned from first principles: only the components inference needs, liquid-cooled GPUs, and a modular format we can place wherever power already exists,” said Flaviu Radulescu, co-founder and CEO of Runware.
Source: Runware


















