60% of AI Will Be On-Prem
Tenry Fu, CEO of Spectro Cloud, on why 60% of AI will be on-prem
Frontier model demand can explode even while the model layer itself gets harder to differentiate. Anthropic is growing because coding agents and other high frequency workflows made token usage surge, not because every task forever needs the single best model. As quality gaps narrow, enterprises increasingly split work, using frontier models for planning and hard reasoning, and cheaper open models on local GPUs for repetitive implementation work.
-
This looks like cloud before multi-cloud. Early on, AWS could grow incredibly fast while compute itself still moved toward utility pricing. Spectro Cloud is built around the same next step for AI, where enterprises route each request to the model that fits its cost, latency, security, and reliability needs.
-
The economic trigger is agentic coding. In the interview, monthly token spend at Spectro Cloud rose 10x from January to June 2026. That kind of usage keeps frontier labs growing fast today, while also pushing buyers to replace large chunks of steady state inference with open models they can run on hardware they already control.
-
The defensible layer shifts upward. Spectro Cloud argues the router itself will commoditize, and the harder product is managing the full stack from servers and Kubernetes to model serving and governance across thousands of clusters. That is similar to how Fireworks sells low latency open model infrastructure rather than a proprietary model, and how Anthropic monetizes premium reasoning demand at the top end.
The market is heading toward a barbell. Frontier labs keep capturing the highest value reasoning workloads, while a broad middle of inference becomes programmable, price sensitive, and increasingly local. The winners will be the companies that either stay clearly best at the top, or make multi-model operations simple enough that enterprises can treat models like interchangeable infrastructure.