Chroma — leaderboard

SwiftInference launches edge AI inference platform with sub-100ms latency

SwiftInference has launched an edge AI platform designed to run voice, vision, LLM, and agentic inference closer to end users by deploying GPU compute at telecom carrier sites. The company is targeting sub-100 ms P90 latency and describes the service as an alternative to cloud or data center infrastructure for teams that need predictable response times and local data handling.

The platform’s key performance claims center on latency and jitter: SwiftInference cites sub-100 ms P90 latency, “faster than cloud,” and 3.7 times lower variance. It also emphasizes data sovereignty by keeping inference within the local metro area and not crossing jurisdictional boundaries, positioning the platform for data residency and compliance-driven use cases.

From an infrastructure perspective, pushing inference into carrier edge sites is a straightforward way to attack tail latency, because the network path is often the biggest variable once a model is warmed and serving is stable. But edge deployments shift the engineering tradeoffs: you’re swapping centralized fleet efficiency for a more distributed footprint where GPU placement, routing policy, and operational maturity decide whether “low latency” holds up under real traffic.

Alongside the platform launch, SwiftInference introduced a Bring Your Own GPU (BYOG) program that accepts idle NVIDIA GPUs, deploys them at carrier edge sites, and provides what it calls a production-ready inference endpoint in five business days. The company frames BYOG as a way to turn existing hardware into live inference capacity in about one week and says it offers a 50/50 revenue share on third-party traffic.

SwiftInference says its API is fully OpenAI-compatible, and that existing applications can migrate by changing an endpoint URL without code modifications. The platform includes a proprietary routing layer that directs each request to an “optimal” node based on model availability, load, and geographic proximity. “We’ve built the fastest AI inference platform on the planet, running at the edge of telecom networks, right next to users. You get the speed of local compute with all the simplicity of a cloud API,” said Kendall Ananyi, co-founder and CEO.

The company says the platform is live across major US metropolitan markets and is available via swiftinference.ai. “Our customers have hit the ceiling of what cloud inference can deliver. Our data says the infrastructure layer can do better,” said Isaac Buwembo, co-founder and COO.

Source: SwiftInference

Chroma — below posts

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

Chroma — tower

Popular Posts:

Quantumscape
Batteries in the AI factory: from spare tire to shock absorber
Johnson-Controls-has-launched-an-Absorption-Chiller-Reference-Design-Guide-aimed-at-data-centers-running-on-on-site-power-generation
Johnson Controls' new absorption chiller guide targets 44% lower cooling power in AI data centers
Why data center power validation is moving to real-time simulation
Why data center power validation is moving to real-time simulation
39D1CAB6-7D1C-4CA7-A2CF-D0919F6EA523_1_105_c
EdgeSites uses existing buildings to deploy 1 MW immersion-cooled AI compute
20260803015848EDT_image_1
Water quality sensors enable real-time monitoring for liquid-cooled AI data centers

Share Your Data Center Engineering News

Do you have a new product announcement, webinar, whitepaper, or article topic? 

TDK — tower ()

Get Data Center Engineering News In Your Inbox:

By subscribing, you agree to our Privacy Policy and Terms of Use. You can unsubscribe at any time.

TDK — leaderboard ()