Go deeper with this research in CS Atlas: Subscribers with access can use Atlas to explore the technical companion, compare CPU architectures, and work through the assumptions behind our forecast. Start with a question: “When do agentic workloads need more cores versus faster cores or more memory?” “Which CPU vendors are best positioned as execution expands around GPUs and other accelerators?” or “What assumptions get us to a $221B server CPU market in 2030, and what would change that forecast?” Atlas connects these questions with our broader infrastructure research, so you can follow the parts of the analysis most relevant to your work.
Agents create more work for the rest of the computer
Late last year, when most industry discussion centered on GPUs, we began making the case that the CPU was about to regain strategic importance. We formalized that argument in March with Secret Agent CPU, which set out our preliminary TAM forecast and explained why agentic AI changes the CPU-to-GPU ratio in the data center, and we raised the model in Secret Agent CPU, Revisited. The view came from close study of the workload, which is how we approach every compute cycle.
The simplest framing is that the GPU does the thinking and the CPU carries out the actions. An agent runs tools, modifies files, queries databases, and evaluates results before deciding on its next step, and that execution runs on CPUs. A coding agent illustrates the pattern: it writes a change, runs the tests, reads the failure, and iterates. Faster inference shortens the reasoning steps, faster CPUs return test results sooner, and additional cores allow more of those jobs to run concurrently rather than queue.
We think it is also important to note that this work extends beyond new AI infrastructure. An agent calling a business service creates work in that service’s software and databases, so part of this demand lands on server CPU infrastructure of every kind, including general-purpose fleets running workloads that have nothing to do with AI. Agents also operate at machine speed, issuing requests far faster than a person can click through an application, and we think that pace means general-purpose infrastructure will need to be upgraded as well as expanded to keep up. Some of the new work fits within existing headroom. The remainder, together with that upgrade cycle, is where the structural expansion of this market begins, and it is the portion this guidebook is designed to measure.
More workers and faster workers
In the AI factory, GPUs run the model's reasoning, CPUs carry out the resulting work, and memory holds work in progress. We describe CPU design along three dimensions. Wide is the ability to run more independent jobs, fast is the speed of an individual worker, and fed is how much memory bandwidth, cache, and capacity each active worker can draw on. More cores serve concurrent users and parallel tools, faster cores shorten dependent steps, and a design that cannot keep its workers fed delivers the full benefit of neither.
Most server CPUs in service today were tuned for a different workload, with cloud parts optimized for core density and cost per virtual machine and accelerator hosts optimized for feeding GPUs. These we have defined in the past as cloud native CPUs. Agentic work asks for all three at once, and because every design trades width, speed, and memory against a fixed budget of power, area, and cost, no single balance point fits every agentic workload. That is why we expect CPU designs built specifically for agentic execution (agentic native CPUs), and why the vendor approaches in this guidebook differ as much as they do.
Exhibit 1. CPU work extends from the agent environment into existing business systems
Conceptual workflow; connections do not imply one coherent hardware domain. Source: Creative Strategies analysis.
Public runtimes already separate model serving from execution, and the always-on designs differ in where that execution happens. Meta describes an isolated environment for Muse’s browser and tools, and Anthropic separates its harness, execution containers, and durable session state (Meta, Anthropic). In our testing of OpenAI’s dots, the agent keeps its own cloud Linux environment available and can also delegate work to a user’s linked computer, which moves part of the execution load out of the data center. None of these designs implies a physical core permanently reserved for each user. Where the work runs, however, changes what each operator needs from cores, memory, storage, and networking, and we examine those differences in the full report.
More activity does not automatically mean more purchases
The CPU opportunity depends on the amount of useful work growing faster than efficiency gains and existing capacity can absorb it. Sharing and suspending environments can support a large agent population with modest purchases, while memory limits can require more servers yet favor lower-cost CPUs. Competition divides along similar lines. A host that exchanges data constantly with an accelerator benefits from close integration, an independent tool pool rewards execution performance, memory, and software support, and different suppliers can win each within the same deployment. We also expect every implementation of personal agents to vary by provider, shaped by the strengths of its models and the infrastructure it can access, as the contrast between Muse and dots already shows. There is no standard way to build an agent computer, and that variety is why CPU vendors are taking such different approaches to cover the breadth of use cases.
Our measure is completed tasks at consistent quality and response time, relative to full-system cost and energy. The full report applies that measure to compare vendors, then separates the capacity customers require from the CPUs they purchase each year.
Inside the Full Report
Our server CPU forecast through 2030, with value nearly tripling to about $220 billion, merchant sales separated from captive silicon, and the split across general-purpose servers, accelerator hosts, and agent execution.
Where NVIDIA, AMD, Arm, Intel, Qualcomm, and the hyperscalers are strongest, architecturally and commercially, and which agent workloads each design fits best.
Why dedicated CPU racks are moving into the scale-up domain, the CPU-to-GPU ratios they imply, and where mixed-vendor configurations still need proof.
How Muse and Dots place agent work in different parts of the stack, from our own testing, and what each design demands of cores, memory, storage, and networking.
One agent workflow traced step by step, showing where CPU speed reaches the user and when it raises GPU utilization.
A sizing method that links task activity to cores, environments, and racks, with a worked comparison of wide, fast, and fed designs under queues and memory limits.
Atlas integration companion packet for subscribers to go deeper with the research.




