Wafer sells inference efficiency
Wafer
This cost structure means Wafer is really selling inference efficiency, not just API access. Because most of its expense is rented Nvidia and AMD capacity, every software gain that makes a GPU produce more tokens per hour drops straight into gross profit. That is why Wafer focuses on kernels, schedulers, quantization, and hardware routing, and why a 50% gross margin is plausible even without owning chips.
-
The product is built around changing the output of a rented GPU, not lowering the rental bill itself. Wafer profiles live traffic, tests engine and hardware configurations, and deploys the fastest setup, so the same paid GPU hour can serve more requests or lower latency enough to support higher value dedicated contracts.
-
AMD is the clearest margin lever. Wafer reported GLM-5.2 on AMD Instinct MI355X reached about 80% of Nvidia B200 performance at less than half the cost, and it also reported an 11.3x Kimi 2.5 throughput gain on AMD. That turns weaker default software support into an opening for software arbitrage.
-
This is the same economic game larger inference platforms are playing. Fireworks also runs at about 50% gross margin and targets 60% through better utilization, while Together combines per token APIs with GPU rentals at about 45% gross margin. Wafer is attacking the same spread, but from a narrower and more optimization driven starting point.
The next step is turning these one off optimizations into a compounding control layer across more chips and more customer traffic. If Wafer keeps finding workload specific gains before open source frameworks and cloud vendors absorb them, it can widen margins, move further into dedicated enterprise endpoints, and become the software layer that decides where open model inference runs.