AWS Mantle commoditizes baseline inference
Wafer
AWS turned managed open model serving into a default cloud feature, which makes Wafer compete less against DIY infrastructure and more against a bundled checkbox inside existing AWS accounts. Once Bedrock let teams hit open models through OpenAI style APIs on the bedrock-mantle endpoint, with IAM controls, project level billing, and AWS handling capacity spikes, the buying decision shifted from best standalone inference stack to whether a company needed performance beyond what its cloud vendor already wrapped into procurement and security workflows.
-
Project Mantle is not just a new endpoint. AWS describes it as a distributed inference engine for large scale model serving, with higher default quotas, unified capacity pools, CloudWatch metrics, and project scoped isolation. That directly overlaps with Wafer's pitch of abstracting away serving operations.
-
The practical wedge for AWS is migration friction. Bedrock supports OpenAI compatible Chat Completions and Responses APIs, and custom model import supports OpenAI compatible schemas for imported models. A team can keep much of its app code the same while moving traffic onto AWS managed infrastructure.
-
Internal comparables show why independent inference vendors still had room before Mantle. A Fireworks customer said Bedrock had a smaller open model catalog in 2025, and chose Fireworks for faster access to new weights like DeepSeek. That suggests Wafer's opening is speed, optimization, and model freshness, not basic hosted availability.
The next phase of this market is a split between bundled baseline inference from hyperscalers and specialist providers that win on faster model onboarding, lower latency, and better unit economics on specific workloads. As AWS keeps expanding Mantle and Bedrock import flows, independent platforms will need to feel materially better in production, not just easier to host.