Home  >  Companies  >  Groq
AI inference chip and cloud tool for running large language models efficiently

Revenue

$90.00M

2024

Valuation

$3.50B

2026

Funding

$901.13M

2024

Details
Headquarters
Mountain View, United States
CEO
Jonathan Ross
Website
Milestones
FOUNDING YEAR
2016

Revenue

Sacra estimates that Groq generated $90M in revenue in 2024. The company primarily generates revenue from selling cloud services for companies to run AI on its chips, similar to how companies purchase access to OpenAI models or AI models from Amazon Web Services. Groq has guided investors to more than $500M in revenue for 2025, revised down from an earlier projection of more than $2B; Groq attributed the shortfall to data-center capacity constraints that shifted some revenue into 2026.

Groq's cloud business was operating at a loss in 2025, with investor documents showing more than $40M in cloud revenue against more than $64M in cloud expenses. Groq also sells chip systems and data center operating services to enterprises and telecommunications companies, including Bell Canada. The company reports that nearly 2 million developers and teams have used its services, indicating strong adoption of its GroqCloud platform.

The revenue growth reflects the company's transition from a pure hardware play to a cloud-first business model, capitalizing on the massive demand for high-speed AI inference capabilities across enterprise customers and the developer community.

Valuation & Funding

Groq was valued at $3.5B following its $350M funding round led by Disruptive in August 2026, with Nvidia planning to participate. This was down from $6.9B at its September 2025 round, before Nvidia entered into a $20B licensing transaction with Groq that included hiring founder Jonathan Ross and other senior employees and returning capital to investors.

In June 2026, Groq raised a separate $650M round to fund its transition from a proprietary AI chip company into an AI neocloud operating Nvidia infrastructure. The company plans to expand from 54 MW of capacity to more than 200 MW in 2027.

Groq has raised more than $1.7B in disclosed funding, with notable investors including Disruptive, Nvidia, BlackRock, Neuberger Berman, Cisco Investments, Samsung Catalyst Fund, and D1 Capital Partners.

Product

Groq is a specialized AI inference company built around its proprietary Language Processing Unit (LPU) chip architecture. The LPU uses a statically-scheduled tensor streaming processor design that moves data through a deterministic pipeline, delivering ultra-low latency and up to 10x energy efficiency compared to traditional GPUs.

The core product is GroqCloud, a fully managed inference platform that developers can access through an OpenAI-compatible REST API. Developers can switch to Groq by changing just three lines of code—their API key, base URL, and model name—and immediately start streaming tokens at hundreds of tokens per second with sub-10 millisecond first-token latency. GroqCloud supports curated open-source models like Llama-3, Qwen, and Mixtral, with performance benchmarks showing speeds like 1,345 tokens per second for Llama-3 8B and 662 tokens per second for Qwen-3 32B. The platform also hosts OpenAI's open models gpt-oss-120B and gpt-oss-20B—launched in partnership with HUMAIN on day-zero availability—with 128K context and integrated server-side tools, achieving 500+ tokens/sec and 1,000+ tokens/sec respectively. The platform includes Python and JavaScript SDKs, streaming responses, and integrations with popular frameworks like LangChain, LlamaIndex, and Vercel AI SDK. Groq has also partnered with Meta to power the official Llama API, offering up to 625 tokens/sec throughput to developers building on Meta's open-source models.

Compound, Groq's agentic orchestration product built on GroqCloud, enables multi-step agent workflows that chain model calls, tool use, and retrieval, targeting developers building autonomous applications. It supports server-side tools including web search, code execution, Wolfram Alpha calculations, and parallel browser use through a single API call. After a beta that reached over 100,000 developers and 5 million requests, Compound is now generally available.

For enterprise customers, Groq offers GroqRack compute clusters—on-premises or colocation racks containing 64 to 576+ LPUs per rack. These systems target hyperscalers, sovereign clouds, and regulated industries requiring data residency, with customers like Aramco Digital ordering hundreds of racks for large-scale deployments.

Business Model

Groq operates a three-tier business model that ramps customers from cloud services to dedicated hardware. The company's B2B go-to-market approach targets both individual developers and enterprise customers through different engagement models.

GroqCloud uses a pay-per-token pricing model similar to OpenAI, where customers pay only for tokens generated. The platform offers both self-serve access with no minimum commitments and enterprise plans with annual volume commitments and volume discounts. This usage-based model allows customers to easily estimate costs before engaging with sales teams.

The hardware business involves selling GroqRack systems directly to enterprises, telecommunications companies, and cloud providers. These customers typically require on-premises deployment for data sovereignty, regulatory compliance, or performance requirements that cloud services cannot meet.

Groq's cost structure benefits from its vertically integrated approach, controlling both chip design and the software stack. The company manufactures chips through foundry partnerships, currently using GlobalFoundries' 14nm process with plans to move to Samsung's 4nm process for next-generation chips. This vertical integration allows Groq to optimize the entire inference pipeline while maintaining higher margins than pure software plays.

Groq has added a technology licensing revenue stream through a non-exclusive agreement with Nvidia, enabling Nvidia to integrate Groq's inference technology while Groq continues operating independently under new CEO Simon Edwards.

Founder Jonathan Ross, President Sunny Madra, and other team members joined Nvidia to support scaling the licensed technology, while GroqCloud services continue without interruption. This licensing model represents a fourth revenue stream alongside cloud services, hardware sales, and enterprise contracts.

Competition

GPU incumbents

Nvidia dominates the AI inference market with its H200, B200, and GH200 Grace-Hopper systems, maintaining over 80% of deployed inference GPUs. While Nvidia typically requires 8-16 GPUs to match Groq's token-per-second performance on large language models, the company's CUDA software ecosystem remains the stickiest competitive barrier. Nvidia is pushing its Blackwell B200 and GB200 systems to reset the performance-per-watt curve while bundling networking components to lock in customers.

The competitive dynamic between Nvidia and Groq has shifted from pure rivalry to active productization. A non-exclusive licensing agreement signed in December 2025 led to Nvidia announcing the NVIDIA Groq 3 LPX inference accelerator rack as part of its Vera Rubin platform, featuring 256 LPU processors, 128GB on-chip SRAM, and 640 TB/s scale-up bandwidth, targeted at low-latency large-context agentic inference and expected to ship in the second half of 2026. This validates Groq's LPU architecture commercially but creates a direct channel-conflict risk: Nvidia can now embed Groq-derived inference capability into its own AI factory stack and sell it to customers who might otherwise deploy GroqRack or GroqCloud.

AMD is positioning its MI300X, MI325X, and upcoming MI355X chips as cost-effective alternatives, with some benchmarks showing advantages on specific workloads. However, AMD's ROCm tooling still lags significantly behind CUDA in developer adoption, and the company faces supply chain challenges that have delayed competitive responses to Nvidia's latest generations.

Intel's Gaudi 3 chips claim 50% better performance than H100 at 40% lower total cost of ownership, but software fragmentation across oneAPI and SynapseAI has slowed enterprise adoption. Intel is targeting cloud partnerships by bundling Gaudi 3 with other data center components.

Specialized inference chips

Cerebras and SambaNova are building competing ASIC architectures optimized for AI inference, arguing that GPU economics break down when batch sizes are small for applications like chat, agent loops, and code assistance. These companies are positioning around extreme throughput-per-watt and deterministic latency, similar to Groq's approach.

Hyperscalers are developing internal silicon solutions including Google's TPU v5, Amazon's Inferentia 3, and Microsoft's Cobalt and Maia chips. While these primarily serve internal workloads, they represent potential competition for third-party inference demand as hyperscalers may offer these capabilities to external customers.

Cloud inference platforms

Traditional cloud providers like AWS, Google Cloud, and Microsoft Azure are expanding their AI inference offerings, leveraging existing customer relationships and integrated billing. These platforms can bundle inference with other cloud services, creating switching costs that pure-play inference companies like Groq must overcome through superior performance or pricing.

TAM Expansion

New products and technology

Groq is developing next-generation 4nm LPU chips through its partnership with Samsung Foundry, promising significant performance-per-watt improvements over current 14nm parts. This technology advancement opens opportunities in higher-value systems markets and enables support for larger context windows and real-time multimodal models.

The company's acquisition of Definitive Intelligence bundled data preparation, orchestration, and dashboard capabilities with Groq hardware, allowing enterprises to purchase turnkey inference systems rather than just chips. This vertical software stack expansion moves Groq up the value chain from hardware provider to complete solution vendor.

Customer base expansion

Groq has grown to over 3.5 million developers as of early 2026, up from 2.8 million in mid-2025 and 360,000 eighteen months prior, with 75% of Fortune 100 companies maintaining accounts on the platform. By supporting open-source communities around models like Llama 3, Gemma, and Mixtral, Groq expands beyond Big Tech buyers into small and medium business developer tool budgets.

Groq has established a structured pathway into federal research and national laboratory workloads through multiple government relationships. The company signed an MOU with the U.S. Department of Energy (December 2025) for potential collaboration under the DOE Genesis Mission on low-latency AI inference for scientific discovery, energy-efficient inference, domestic supply chains, and benchmarks, alongside its existing Carahsoft reseller partnership and FedRAMP roadmap.

Bell Canada has selected Groq as the exclusive inference provider for Bell AI Fabric, a planned six-facility, up to 500MW hydro-powered sovereign AI compute network in Canada. The network is rolling out progressively, with an initial 7MW site operational in Kamloops, a second 7MW site in Merritt, and a larger 26MW facility at Thompson Rivers University planned for 2026.

Geographic expansion

Groq has secured a $1.5 billion commitment from Saudi Arabia to fund the largest non-hyperscaler inference cluster with over 19,000 LPUs, positioning Groq as the reference platform for Arabic language models and establishing a major presence in the Middle East market. The company's international footprint now spans multiple regions: a data center in Helsinki through Equinix meeting European data sovereignty requirements while leveraging green hydroelectric power for lower operating costs; a 4.5MW Equinix facility in Sydney representing its first Asia-Pacific footprint, with Canva among its early Australian customers; and a United Kingdom data center through Equinix, further expanding its European coverage.

Risks

Nvidia dependency: Groq's non-exclusive technology licensing agreement with Nvidia means Nvidia can integrate Groq's inference technology across its broader product portfolio, potentially using it to compete more directly with GroqCloud and Groq's hardware business. The departure of founder Jonathan Ross and President Sunny Madra to Nvidia further concentrates key institutional knowledge with a party that now holds a license to Groq's core IP.

Leadership continuity: Successor CEO Simon Edwards left Groq to join Bloom Energy as CFO effective April 13, 2026—less than four months after taking the role—and no authoritative announcement of a replacement has been made. Back-to-back CEO transitions in a short window, combined with the earlier loss of founder Ross and President Madra, leaves Groq's executive bench materially thinned.

Revenue execution: Groq's annual revenue around the time of the Nvidia transaction was approximately $100 million, well below its earlier $2 billion and revised $500 million 2025 targets, while investor documents showed the cloud business operating at a loss. Sustained capacity limitations or continued cloud unit losses would impair Groq's ability to reach profitability before requiring additional capital.

News

DISCLAIMERS

This report is for information purposes only and is not to be used or considered as an offer or the solicitation of an offer to sell or to buy or subscribe for securities or other financial instruments. Nothing in this report constitutes investment, legal, accounting or tax advice or a representation that any investment or strategy is suitable or appropriate to your individual circumstances or otherwise constitutes a personal trade recommendation to you.

This research report has been prepared solely by Sacra and should not be considered a product of any person or entity that makes such report available, if any.

Information and opinions presented in the sections of the report were obtained or derived from sources Sacra believes are reliable, but Sacra makes no representation as to their accuracy or completeness. Past performance should not be taken as an indication or guarantee of future performance, and no representation or warranty, express or implied, is made regarding future performance. Information, opinions and estimates contained in this report reflect a determination at its original date of publication by Sacra and are subject to change without notice.

Sacra accepts no liability for loss arising from the use of the material presented in this report, except that this exclusion of liability does not apply to the extent that liability arises under specific statutes or regulations applicable to Sacra. Sacra may have issued, and may in the future issue, other reports that are inconsistent with, and reach different conclusions from, the information presented in this report. Those reports reflect different assumptions, views and analytical methods of the analysts who prepared them and Sacra is under no obligation to ensure that such other reports are brought to the attention of any recipient of this report.

All rights reserved. All material presented in this report, unless specifically indicated otherwise is under copyright to Sacra. Sacra reserves any and all intellectual property rights in the report. All trademarks, service marks and logos used in this report are trademarks or service marks or registered trademarks or service marks of Sacra. Any modification, copying, displaying, distributing, transmitting, publishing, licensing, creating derivative works from, or selling any report is strictly prohibited. None of the material, nor its content, nor any copy of it, may be altered in any way, transmitted to, copied or distributed to any other party, without the prior express written permission of Sacra. Any unauthorized duplication, redistribution or disclosure of this report will result in prosecution.