Disaggregated Inference Weakens Etched's Pitch

Diving deeper into

Etched

Company Report
its AMD partnership introduces a disaggregated inference model that could reduce demand for Etched's single-vendor integrated approach.
Analyzed 6 sources

This partnership matters because it turns inference hardware into a mix and match pipeline instead of a full stack rack purchase. AMD handles the heavy front half of a request, where long prompts are ingested and turned into KV cache, while Cerebras handles the back half, where tokens are generated one by one at very high speed. That gives buyers a way to keep existing GPU style infrastructure and add a specialist decode engine only where it helps most.

  • Etched is selling the idea that one vendor can own both prefill and decode inside the same system boundary. The AMD and Cerebras design weakens that pitch by showing customers can buy the two stages separately, and optimize each with different silicon without replacing the whole cluster.
  • This is becoming a pattern, not a one off. Cerebras is also bringing a split prefill and decode stack to AWS with Trainium on prefill and CS-3 on decode, and SambaNova is working with Intel on a similar division where GPUs handle prompts, RDUs handle token generation, and Xeon runs agent actions.
  • The practical appeal is simple. Most enterprises and cloud operators already have CPUs and GPUs for prompt processing, orchestration, and tool use. A disaggregated design lets them insert a specialist inference box into that workflow, which is an easier sale than adopting a new monolithic rack from a single startup.

The market is moving toward best engine per stage. If that keeps spreading, winners will be the vendors whose chips become the default decode layer inside mixed clusters, and integrated systems like Etched will need to prove that one box is so much simpler or cheaper that buyers should give up that flexibility.