60% On-Prem Enterprise AI
Tenry Fu, CEO of Spectro Cloud, on why 60% of AI will be on-prem
This mix points to enterprise AI becoming an infrastructure problem, not just a model selection problem. Once inference shifts onto a company’s own GPUs, the hard part becomes keeping thousands of clusters, drivers, models, and policies in sync across data centers, stores, hospitals, factories, and air gapped sites. That is why the control plane gains value. It lets one central team roll out the same approved stack everywhere, patch it remotely, and meter usage across local, cloud, and API paths.
-
The economics push work on premises first. Spectro Cloud saw its own monthly token bill rise 10x from January through June as agentic coding usage expanded. Its product thesis is that open weight models are now good enough for many coding, security, and internal workflow tasks, so repeated inference is cheaper on owned GPUs than through metered APIs.
-
Edge is the next leg of the same shift. PaletteAI is built so a central platform team can remotely install the OS, Kubernetes, networking, GPU drivers, inference engines like vLLM, and approved models across thousands of intermittently connected sites. That matters in retail, defense, healthcare, and manufacturing, where data is created far from a central cloud region.
-
The competitive line is between full stack control planes and narrower tools. DataRobot governs models, prompts, agents, and compliance across cloud and on premises environments. NVIDIA Run:ai focuses on GPU scheduling and utilization. Spectro Cloud sits lower in the stack, managing the fleet itself from bare metal up through model serving, which becomes more important as Kubernetes spreads into inference operations.
The next phase is a VMware like layer for AI infrastructure. As more enterprises standardize on hybrid inference, the winning platforms will be the ones that make self hosted AI feel operationally boring, with repeatable blueprints, remote upgrades, policy controls, and routing across edge, cloud, and frontier APIs from one system.