Control of Runtime Drives Enterprise Capture
General Compute
Taking control of the runtime turns General Compute from a fast chip reseller into a fuller inference platform. Right now the company already owns the API layer, routing, streaming path, and dedicated deployment motion, while the ASIC runtime still sits with a hardware partner. Pulling batching, scheduling, and KV cache management in house lets it tune the part of serving that decides latency, reliability, and how much of each enterprise contract stays with General Compute instead of leaking to a supplier.
-
The runtime is where enterprise serving economics get decided. It controls queueing, retries, admission control, batching, and token streaming, which means it shapes p99 latency, throughput per dollar, and SLA performance, the exact things buyers pay for in dedicated capacity and private model deals.
-
Owning this layer also lowers supplier dependence. General Compute says the runtime on its ASIC is owned by a hardware partner today, with plans to take more of it in house over the next 6 to 8 months. That reduces the risk that roadmap delays, pricing changes, or software bottlenecks at the partner level cap product velocity.
-
The closest analog is Groq, which pairs custom silicon with its own cloud stack, while GPU native platforms like Together AI and Fireworks rely more on standard serving frameworks and then differentiate through developer experience, tuning, and adjacent products. General Compute moving down stack is the path to a more Groq like hardware and software bundle.
If this transition works, General Compute can sell enterprises one serving surface from self serve API to reserved clusters to private weights and co location. That makes the company harder to compare on simple tokens per second benchmarks alone, and better positioned to win larger multi year infrastructure contracts as inference buying shifts from model access to full production systems.