Cerebras and Callosum are partnering to integrate Cerebras silicon into the Callosum software platform, with the goal of supporting ultra-low-latency, heterogeneous agentic inference workloads. The companies are targeting multi-agent systems where different stages of reasoning and action have different compute requirements, pushing operators toward mixed infrastructure instead of a one-size-fits-all stack.
The integration centers on Callosum’s software layer, which decomposes and orchestrates heterogeneous agentic workloads. Cerebras is contributing its Wafer-Scale Engine for inference, which the companies tie to low latency for enabling “new categories of agentic AI.” Cerebras also says the work expands its European reach, following its previously stated plan to deliver 200 MW of European AI compute data center capacity.
For data center and platform engineers, the practical takeaway is that “agentic” workloads are being framed as orchestration problems as much as raw accelerator performance problems. If an inference pipeline is genuinely heterogeneous—multiple specialized models, long-horizon planning loops, and action steps—then scheduling, data movement, and matching stages to the right hardware can become the limiting factor just as quickly as FLOPS.
Commercially, Cerebras capacity will be exposed through Callosum APIs, positioning the Callosum platform as an access layer for Cerebras-backed inference. “Agentic AI demands infrastructure that can keep up with the speed of reasoning, not just the size of the model,” said Andy Hock, Chief Strategy Officer at Cerebras. Callosum CEO Danyal Akarca added, “The next-generation of AI will be defined by how intelligently compute is orchestrated, not simply how much compute is available.”
Source: Cerebras


















