The emergence of agentic AI in datacenters is introducing new architectural challenges that have remained largely unexplored. A new research paper, featuring a production study at Microsoft Azure and an analysis of open-source frameworks, provides the first architectural characterization of how these agentic systems strain existing infrastructures.

Fragmentation of Execution

According to the study, agentic workflow execution is characterized by intense heterogeneity. Each request expands into a sequence of Large Language Model (LLM) inferences, tool invocations, and orchestration decisions. This process forces the system to repeatedly cross the CPU-GPU boundary.

A key finding is that the CPU is now on the "critical path," as orchestration and tools run on the host. Resource demand is not steady but consists of low average loads with sudden spikes, making conventional uniform servers inefficient for these workloads.

The Agora Prototype

To address these mismatches, researchers presented Agora, a prototype designed for commodity servers. Agora implements several strategies:

  • Dynamic harvesting of idle CPU cores for high-throughput background work.
  • Protection of agent tail latency against sudden spikes caused by tool invocations.
  • GPU memory oversubscription, allowing more agents to be placed on each GPU.
  • Use of prefetching to hide swap latency during agent state transitions.

The results indicate that pooling cores by role and applying affinity-aware scheduling can restore microarchitectural locality and significantly improve server utilization and throughput while preserving latency targets.