Home  >  Companies  >  Etched
Etched
Designs and ships rack-scale AI inference systems including custom chips, liquid-cooled racks, and software for high-throughput, low-latency model inference

Funding

$625.00M

2024

View PDF
Details
Headquarters
San Jose, United States
CEO
Gavin Uberti
Website
Milestones
FOUNDING YEAR
2022

Valuation & Funding

Etched's most recent valuation is $10.3B, set in its Series C round of $300M announced on July 23, 2026, led by Sequoia Capital, with participation from Andreessen Horowitz, SK Hynix, Jane Street, and Diffusion.

Before the Series C, Etched disclosed on June 30, 2026 that it had raised $800M in total across all prior rounds, including a $500M round closed in December 2025 at a $5B post-money valuation. Earlier, Etched raised a $120M Series A in June 2024 and a $5.36M seed round in June 2023.

Across all rounds, Etched has raised approximately $1.1B in total lifetime funding. Other investors in the company's financing history include VentureTech Alliance, Stripes, Peter Thiel, Ribbit Capital, Radical Ventures, Primary VC, Positive Sum, Hudson River Trading, Jump Trading, and Two Sigma.

Product

Etched builds rack-scale inference systems for running frontier AI models in production. Rather than selling a standalone chip, it delivers a co-designed stack of custom silicon, boards, liquid cooling, proprietary interconnect, and software as a complete inference machine. Customers buy a tightly coupled rack designed to operate as one inference system, instead of procuring accelerators, servers, networking, and thermal hardware from separate vendors and integrating them.

The architecture targets the two phases of inference that most GPU clusters handle inefficiently. Prefill, where the system digests a long prompt or context window, and decode, where it generates tokens back to the user one at a time, have different compute and memory profiles. Etched designs the same rack to handle both phases, rather than relying on separate hardware for each.

Two technical components anchor the system. Low Voltage Inference allows Etched's compute blocks to run at under half the voltage of typical AI chips, which lets the rack sustain high utilization without thermal throttling. Cluster Scale Memory creates a shared memory pool across the rack's scale-up domain through a proprietary low-latency interconnect and an HBM/SRAM hybrid design, so the rack behaves more like one large machine than a set of loosely connected accelerators.

The software layer extends beyond the hardware. Etched builds firmware, kernel drivers, virtualization support, power and thermal management, DMA transfer, PCIe-visible memory management, and cluster-scale orchestration software, because customers deploy racks in shared production environments where multi-tenancy, observability, and operational reliability matter alongside peak throughput.

As of summer 2026, Etched's A0 silicon had returned from TSMC on the N4P process node, first racks were shipping to customers, and the system had been validated against production traffic patterns including models such as DeepSeek, Qwen, Mamba, and Llama.

Business Model

Etched sells vertically integrated AI infrastructure to a small number of very large B2B buyers through a high-touch, negotiated sales process. The unit of sale is the rack or cluster system, not a chip or a software license, so Etched captures margin across silicon, boards, thermal design, rack integration, and system software in a single contract.

Monetization is anchored in large system contracts tied to deployed inference capacity. Customers commit to specific rack deployments, and revenue is recognized as systems are manufactured, shipped, and accepted. The model is closer to a capital equipment sale than a SaaS subscription, though software and support are bundled into the deployment.

Vertical integration is the structural differentiator. Owning the chip, rack, interconnect, cooling, and software reduces the margin leakage a fabless chip vendor would face when OEMs, networking vendors, and integration partners each take a cut. It also lets Etched optimize the full system for tokens per watt and end-to-end latency rather than a single component in isolation.

The cost structure is capital-intensive by design. Etched has opened a Taiwan factory, built a San Jose data center and NPI prototyping lab, added an 80,000-square-foot facility in Milpitas, and constructed a 10-megawatt validation lab. That infrastructure is expensive, but it functions as both manufacturing capacity and a customer qualification environment, where Etched can validate racks against real production traffic before shipment.

Competition

Etched is entering a market where the competitive unit has shifted from the chip to the full rack. Incumbents, hyperscalers, and specialist startups are converging on the same system-level framing.

GPU incumbents

Nvidia remains the default baseline for inference buyers, and its relevance has increased as it has moved up the stack. The Vera Rubin NVL72 platform combines 72 GPUs, 36 CPUs, NVLink switching, and networking into a rack-scale system backed by TensorRT-LLM, Dynamo, and years of CUDA ecosystem inertia. Etched is competing not against a chip but against a complete rack architecture with a broad OEM partner base.

AMD is a secondary but increasingly credible incumbent threat. Its Helios rackscale platform combines Instinct GPUs, EPYC CPUs, Pensando networking, and ROCm software, and Microsoft has committed to deploy Helios on Azure for frontier-model inference. AMD's willingness to participate in heterogeneous inference pipelines, explicitly splitting prefill and decode across different compute engines in its Cerebras partnership, makes it a systems integrator rather than just a chip vendor, the same framing Etched is using.

Specialist inference challengers

Groq, Cerebras, and SambaNova are the most direct specialist peers. Groq has moved furthest toward cloud-first distribution, operating across multiple data centers globally and processing trillions of tokens weekly, which gives it go-to-market reach Etched does not yet have. Cerebras completed its IPO in May 2026 and is expanding manufacturing capacity, while its AMD partnership introduces a disaggregated inference model that could reduce demand for Etched's single-vendor integrated approach.

SambaNova competes less on raw throughput and more on turnkey enterprise deployment, particularly in sovereign, regulated, and air-gapped environments. Its Intel partnership, splitting prefill, decode, and agentic tool execution across different components, makes it a heterogeneous systems integrator for the agent era. Fractile and Majestic Labs compete from a memory-topology angle, while d-Matrix has moved from component supplier toward full rack-scale systems through its acquisition of GigaIO's data center business, converging on nearly the same product boundary as Etched.

Hyperscalers and managed inference

AWS Trainium2 already powers the majority of inference on Bedrock as part of a custom silicon stack running at a revenue run rate above $10B. Google's Trillium and Ironwood TPU lines serve Gemini and other workloads at scale. Microsoft's Maia 200, described as built specifically for inference and live in production datacenters, raises the benchmark any external vendor must clear to win Azure-adjacent workloads.

Managed inference platforms like Together AI, Fireworks AI, and Baseten absorb demand that might otherwise justify buying specialized hardware, offering API abstraction, concurrency handling, and time-to-model-access that some buyers value more than raw compute architecture. These platforms are not direct competitors to Etched's rack sales, but they reduce the pool of buyers that need dedicated inference infrastructure at all.

TAM Expansion

Etched's expansion logic is that inference is becoming the dominant AI compute workload, and that buyers with the highest inference demand are still early in building dedicated infrastructure.

New products and workload expansion

Etched's architecture is designed for many-trillion-parameter MoE models, long-context workloads, and agentic tasks where context reuse and memory latency matter disproportionately, and the system targets both prefill and decode within a single rack.

As AI usage shifts from single-turn chat toward longer-horizon agent work, the compute profile of inference changes: tasks run longer, context windows grow, and cost per completed task becomes the relevant metric rather than cost per token. That shift increases the value of Etched's latency and throughput claims because customers are buying end-user experience and task completion economics, not raw FLOPs.

A logical product expansion from here is cluster management, model-serving software, profiling, and scheduling tooling that makes Etched racks easier to operate at scale. Moving up the software stack would let Etched add recurring revenue on top of one-time system sales, similar to Cerebras's move from hardware into pay-per-token inference.

Customer base expansion

Etched's initial customer base is concentrated among frontier AI labs, cloud providers, and hyperscalers. The next layer is AI-native application companies and inference clouds that need dedicated high-performance infrastructure for coding, search, document processing, and enterprise agents but do not have their own chip programs.

Sovereign and regulated deployments are an adjacent segment. Because Etched sells complete rack systems rather than cloud access, it can serve buyers with data residency, security, or latency requirements that make shared public-cloud GPU pools unsuitable, including governments, financial institutions, healthcare networks, and defense-adjacent operators.

As agentic inference becomes a broader enterprise workload rather than a niche developer one, the buyer universe for low-latency, high-throughput inference systems extends beyond the handful of frontier labs that anchor Etched's current backlog.

Geographic and supply-chain expansion

Etched has established manufacturing presence in Taiwan, which shortens the distance to foundry, packaging, and memory partners and gives the company a path to serve Asian cloud providers and sovereign buyers as a second commercial wave after North American hyperscalers.

SK Hynix's strategic investment in Etched matters here: preferred access to HBM and advanced packaging capacity could become a competitive advantage if inference demand stays supply-constrained, because memory availability is one of the binding constraints on how fast any inference rack vendor can scale production.

As data-center power becomes more scarce and expensive, vendors that improve tokens per watt or reduce deployment footprint can expand geographically into more constrained sites. With the IEA projecting U.S. electricity demand growing close to 2% annually through 2030 with data centers as a major driver, power efficiency is also a market-access issue in regions where capacity is limited.

Risks

Architecture concentration: Etched's silicon is optimized for transformer inference patterns, and a material shift in frontier model architectures or serving patterns away from the assumptions embedded in Etched's chip and cluster design could create a mismatch between a fixed hardware roadmap and a changing software frontier.

Buyer concentration: Etched's initial backlog is concentrated among a small number of frontier AI companies, cloud providers, and hyperscalers with long qualification cycles, growing in-house silicon ambitions, and enough bargaining power to use Etched as a benchmarking lever against incumbent suppliers rather than as a durable platform commitment.

Heterogeneous inference: The normalization of multi-engine inference pipelines, where AMD pairs with Cerebras to split prefill and decode, and Intel pairs with SambaNova to split prefill, decode, and agentic tool execution across different components, threatens Etched's single-vendor integrated rack thesis because buyers may prefer best-engine-per-stage modularity over a monolithic architecture from one vendor.

News

DISCLAIMERS

This report is for information purposes only and is not to be used or considered as an offer or the solicitation of an offer to sell or to buy or subscribe for securities or other financial instruments. Nothing in this report constitutes investment, legal, accounting or tax advice or a representation that any investment or strategy is suitable or appropriate to your individual circumstances or otherwise constitutes a personal trade recommendation to you.

This research report has been prepared solely by Sacra and should not be considered a product of any person or entity that makes such report available, if any.

Information and opinions presented in the sections of the report were obtained or derived from sources Sacra believes are reliable, but Sacra makes no representation as to their accuracy or completeness. Past performance should not be taken as an indication or guarantee of future performance, and no representation or warranty, express or implied, is made regarding future performance. Information, opinions and estimates contained in this report reflect a determination at its original date of publication by Sacra and are subject to change without notice.

Sacra accepts no liability for loss arising from the use of the material presented in this report, except that this exclusion of liability does not apply to the extent that liability arises under specific statutes or regulations applicable to Sacra. Sacra may have issued, and may in the future issue, other reports that are inconsistent with, and reach different conclusions from, the information presented in this report. Those reports reflect different assumptions, views and analytical methods of the analysts who prepared them and Sacra is under no obligation to ensure that such other reports are brought to the attention of any recipient of this report.

All rights reserved. All material presented in this report, unless specifically indicated otherwise is under copyright to Sacra. Sacra reserves any and all intellectual property rights in the report. All trademarks, service marks and logos used in this report are trademarks or service marks or registered trademarks or service marks of Sacra. Any modification, copying, displaying, distributing, transmitting, publishing, licensing, creating derivative works from, or selling any report is strictly prohibited. None of the material, nor its content, nor any copy of it, may be altered in any way, transmitted to, copied or distributed to any other party, without the prior express written permission of Sacra. Any unauthorized duplication, redistribution or disclosure of this report will result in prosecution.