Together ATLAS validates Wafer thesis
Wafer
This shows that workload aware inference is quickly becoming the core moat in managed model serving, not a nice extra. Together is pushing the same basic idea as Wafer, use live production traffic to keep making the serving stack better after deployment. In practice that means the platform is not just renting GPUs and exposing an API, it is watching real token patterns, adapting decode behavior, and turning customer traffic into a performance data flywheel.
-
ATLAS is not a one time speed trick. Together describes it as an adaptive speculative decoder that updates from historical patterns and live traffic, and ATLAS 2 extends that online learning loop as a general template for systems that improve under production use.
-
That is close to Wafer's product thesis. Wafer is built around continuously optimizing open source LLMs plus kernels, engines, and hardware for lower latency and better efficiency, which is the same layer of value creation, improving inference after the model is already in production.
-
The overlap matters because Together is doing it at much larger scale. Together says API volume grew to more than 400 trillion tokens per month and offers serverless APIs, dedicated reasoning clusters, batch inference, voice agents, and private deployments, while Wafer is much earlier at an estimated $8M annualized revenue as of August 31, 2026.
The market is heading toward inference platforms that behave more like self improving operating systems than static hosting layers. The winners will keep absorbing more of the optimization stack, from scheduling and decoding to model routing and hardware placement, and real workload data will become the training set for infrastructure itself.