Modularity Threatens Etched Rack Thesis
Etched
This points to a market where the winning system may be the one that lets buyers mix the fastest part for each job, not the one that sells the whole rack. In practice, inference is breaking into separate steps, prefill for reading the prompt, decode for generating tokens, and CPU work for tool calls and orchestration. That makes a modular stack easier to justify than a single chip and rack design that has to win every step at once.
-
AMD and Cerebras now market a disaggregated setup where AMD Helios handles prompt prefill and Cerebras handles decode. That is a direct example of a buyer getting rackscale throughput from one vendor and token speed from another, instead of standardizing on one integrated box.
-
Intel and SambaNova are pushing the same pattern one step further. Their blueprint uses GPUs for prefill, SambaNova RDUs for decode, and Intel Xeon 6 CPUs for agent tools and orchestration, which turns inference into a chain of specialist engines rather than one monolithic machine.
-
Etched is taking the opposite bet. Its current positioning is to co design chips, racks, software, and manufacturing into one rack scale product optimized for both prefill and decode. That can simplify procurement and tuning, but it also means buyers must accept one vendor across every stage of the workload.
The next phase of inference infrastructure is likely to look more like a composable factory line than a single appliance. If prefill, decode, and agent execution keep separating, vendors that slot cleanly into mixed fleets will widen their addressable market, while integrated rack players will need to prove they are materially better across the full pipeline, not just in one stage.