Wafer's Continuous Optimization Moat
Wafer
Wafer is defensible only if it behaves less like a static inference host and more like a fast moving performance lab. Its margin comes from finding small stack level gains before others do, then pushing those gains into production across kernels, schedulers, model configs, and cheaper AMD capacity. Once AMD and rival hosts copy any single trick, the real moat is how quickly Wafer finds the next one and rolls it out across live workloads.
-
The evidence already shows the edge is operational, not theoretical. Wafer has published workload specific improvements on AMD, including 11.3x higher Kimi 2.5 throughput, DeepSeek V3.2 output speed rising from 38.5 to 200.8 tokens per second, and GLM 5.2 on MI355X reaching about 80% of Nvidia B200 performance at less than half the cost. Those wins came from many fixes across the stack, not one breakthrough.
-
This is the same lane larger inference platforms are racing into. Baseten, Fireworks AI, and Together AI all sell developer friendly model serving, but at far larger scale, with estimated run rate revenue of about $600M, $800M, and $1B versus Wafer at $8M. That means any durable edge for Wafer has to come from sharper optimization loops, especially on hardware and models the bigger platforms have not tuned deeply yet.
-
AMD is both an opportunity and a clock. Wafer benefits today from ROCm immaturity and AMD alignment, including AMD Ventures backing and access to roadmap and capacity, but AMD is also pushing Helios rack scale systems and broader heterogeneous inference. As AMD software gets easier to use, more of the raw hardware discount will pass through to the market, reducing standalone arbitrage.
The next phase favors companies that can turn inference optimization into a continuous shipping process. As model routing spreads across Nvidia, AMD, and other accelerators, the winners will be the platforms that ingest new chips, tune them quickly, and expose the gains through simple APIs and gateway partnerships before the rest of the market reprices.