Home  >  Companies  >  Design Arena
Design Arena
Crowdsourced benchmark platform that compares AI models by presenting their design outputs side-by-side for human voting

Revenue

$60.00M

2026

Details
Headquarters
San Francisco, United States
CEO
Grace Li
Website
Milestones
FOUNDING YEAR
2025

Revenue

Sacra estimates that Design Arena hit $60M in annualized revenue in July 2026.

Design Arena's rapid revenue growth was driven primarily by enterprise evaluation contracts with frontier AI model laboratories. The company launched its first product in July 2025 with eight models and around 3,000 users, began enterprise outreach in September 2025, and closed its first major lab deal shortly thereafter.

Revenue is concentrated among a small number of high-value enterprise customers rather than millions of consumer users. Frontier labs pay for private model evaluations, preference data, launch-day benchmarking, and diagnostic analysis, while the consumer product remains free. The platform had approximately 5.3 million users in August 2026, but the vast majority of revenue comes from B2B customers, with consumer usage generating data rather than revenue.

Valuation & Funding

Design Arena raised a $7.9M seed round announced on August 3, 2026, led by Index Ventures with participation from Conviction, A*, Valkyrie, and Y Combinator.

The company participated in Y Combinator's Summer 2025 batch. The $7.9M seed represents its total disclosed primary funding to date.

Product

Design Arena is a multi-model creation interface and crowdsourced benchmark for generative AI. A user describes a website, mobile app, game, slide deck, logo, image, video, or voice clip, and the platform sends the same prompt to multiple AI models simultaneously. It then presents the outputs side by side without identifying the model behind each one.

In each blind comparison, four randomly selected models are paired off, and the user ranks them from first to fourth through a series of pairwise choices. Each session generates multiple ranking observations, which feed into a continuously updated public leaderboard.

For images and audio, the platform calls competing APIs with matched prompts. For full-stack web applications and games, it places each model in a standardized sandbox where it can read and write files, run shell commands, search the web, generate media, interact with databases via Supabase, and deploy the finished result to Vercel. Design Arena records each action as a structured tool call, allowing it to reconstruct how an agent produced the final output.

Separate leaderboards cover foundation models, AI builder products such as Cursor and Replit, and autonomous coding agents such as Devin and Claude Code. Rankings use a Bradley-Terry model displayed on an Elo-style scale, with models needing around 200 pairwise comparisons to reach reliability thresholds. By August 2026, Design Arena covered website design, full-stack apps, mobile apps, games, slides, images, image editing, logos, SVGs, 3D assets, video, video editing, text-to-speech, data visualization, and several multimodal transformation tasks.

An API provides programmatic access to leaderboard data across all three evaluation groups. Access requires an application and API key, and public displays must credit and link to Design Arena.

Business Model

Design Arena uses a multi-sided model with free consumer acquisition and paid enterprise products. Free access generates prompts and preference votes, which feed private evaluations, preference datasets, and model-performance analytics sold to AI laboratories.

Benchmarking is embedded in the creation workflow. Users access expensive frontier models through one interface and receive a generated artifact, while Design Arena collects prompts, pairwise preferences, and interaction data. The model resembles LMArena's approach to building its text-model benchmark through free chat access before expanding into paid evaluations, APIs, and model routing.

Model inference accounts for most costs because each prompt may trigger four or more competing generations. Agentic tasks also require sandboxed compute for file systems, shell execution, database access, and deployment infrastructure. These costs rise with engagement, but each session can produce an artifact for the user, several pairwise votes, a benchmark update, reusable prompt data, and provider-performance signals.

Enterprise evaluation contracts, preference-data licensing, and custom benchmarking programs carry the highest margins because they reuse Design Arena's dataset and ranking infrastructure without proportionate increases in free consumer inference. The company has approximately 10 employees and relies on provider APIs for model access, users for subjective evaluation, and cloud services for infrastructure.

Competition

Design Arena competes across AI model evaluation, human-preference data, and creative tooling. The market is split among mass-market arenas, expert evaluation providers, automated benchmarks, and creative-software incumbents that generate implicit preference signals from their user bases.

Mass-market arenas

Arena.ai, the successor to LMArena and Chatbot Arena, is the closest direct competitor. Its WebDev Arena uses the same blind-comparison format for frontend and full-stack web development, with roughly 492,000 votes across 106 models by mid-2026. Arena.ai covers chat, code, image, video, and search, allowing it to funnel a broader general AI audience into creative evaluations. It also has an academic lineage and established relationships with model labs.

Design Arena specializes in creative artifacts and taste across SVG, logos, slides, 3D assets, video editing, text-to-speech, and Android-native development, categories that Arena.ai does not organize as cohesively. Arena.ai could expand further into design categories at relatively low cost using its existing consumer traffic and provider relationships.

Expert and enterprise evaluation

Contra Labs uses vetted professional creatives rather than undifferentiated public voters. Its Human Creativity Benchmark combines pairwise preference with scalar ratings and written explanations across ideation, mockup, and refinement stages. Each judgment contains more diagnostic information than a simple A-versus-B vote, making it more actionable for post-training work.

Scale AI, Surge AI, and Invisible bundle expert red-teaming, preference labeling, and secure enterprise delivery with contracted evaluator workforces and formal audit trails. Artificial Analysis competes in image-generation arenas and pairs quality benchmarks with cost, latency, and provider-level performance data through its Optima enterprise product. These providers can distinguish their offerings from Design Arena's focus on immediate popularity by measuring professional quality, production readiness, or brand suitability.

Creative-tool incumbents

Adobe, Canva, Figma, Framer, Gamma, and Vercel control high-intent creative workflows and can generate proprietary preference and usage data without operating a public leaderboard. They observe which outputs users publish, edit, export, or discard, producing behavioral signals that may be more commercially relevant than explicit votes. Gamma's experience indicates that AI improves activation by eliminating the blank-page problem, while editability and production depth remain product constraints.

TAM Expansion

Design Arena started as a website-design leaderboard and has expanded into a general-purpose evaluation and creation layer for generative AI. Its addressable market now extends beyond public rankings into evaluation infrastructure, new modalities, model routing, and paid creation tools.

Evaluation infrastructure for model developers

The largest near-term opportunity is converting the public benchmark into a full-stack evaluation platform. Private arenas would let model labs test unreleased systems against public and proprietary baselines, run release-to-release regression testing, and receive diagnostic breakdowns by task type, user cohort, and prompt complexity.

This would shift the product from ranking models to identifying why users prefer one output and where a model underperforms. The preference dataset could also be used to train reward models and design judges, particularly for visual and interactive outputs where automated grading remains weak. Prolific, Handshake, and Mercor have generated substantial enterprise spending on human feedback and evaluation infrastructure, creating a budget category that Design Arena could address through its organic, product-embedded data collection.

New modalities and agent evaluation

Design Arena has expanded from web design into images, video, audio, slides, games, 3D assets, mobile apps, and autonomous coding agents. Each modality adds an evaluation budget from model providers and a category of preference data.

The agent evaluation harness records tool calls, file edits, execution, and deployment behavior in standardized sandboxes, allowing Design Arena to benchmark the stack from foundation model to finished application. The potential market extends beyond output aesthetics to coding-agent vendors such as Cursor and Devin, app-building platforms, and enterprise software buyers comparing model-and-harness combinations.

Preference-aware routing and creation platform

With millions of users accessing multiple models through one interface, Design Arena could expand from benchmarking into a paid creation workspace with model routing. Its models directory already combines preference performance with pricing, context windows, and output limits. A routing layer could select models by quality, cost, latency, and use case, capturing value downstream of the benchmark.

Premium features such as higher generation limits, faster queues, iterative editing, project memory, and export to tools such as Figma or Pitch could monetize users who value the creation product independently of the leaderboard. Geographic and demographic segmentation of preference data could also inform localized advertising, ecommerce, and product design decisions across the 190-plus countries where Design Arena already has users.

Risks

Benchmark credibility: Design Arena's business depends on its perceived neutrality, but its primary paying customers are the model labs whose products it ranks, creating a structural tension in which commercial relationships could undermine trust without clear separation between paid integration services and independent vote collection.

Customer concentration: Revenue appears to come from a small number of frontier AI laboratories, so the loss or budget reduction of even one or two major lab customers could materially reduce annualized revenue, while the addressable pool of frontier labs willing to pay millions annually for preference data is limited.

Inference cost exposure: Every user prompt triggers four or more competing model generations and sandboxed compute for agentic tasks, giving the free consumer product that generates preference data variable costs that scale with engagement, particularly as the platform expands into compute-intensive modalities such as video, full-stack applications, and long-running autonomous agents.

DISCLAIMERS

This report is for information purposes only and is not to be used or considered as an offer or the solicitation of an offer to sell or to buy or subscribe for securities or other financial instruments. Nothing in this report constitutes investment, legal, accounting or tax advice or a representation that any investment or strategy is suitable or appropriate to your individual circumstances or otherwise constitutes a personal trade recommendation to you.

This research report has been prepared solely by Sacra and should not be considered a product of any person or entity that makes such report available, if any.

Information and opinions presented in the sections of the report were obtained or derived from sources Sacra believes are reliable, but Sacra makes no representation as to their accuracy or completeness. Past performance should not be taken as an indication or guarantee of future performance, and no representation or warranty, express or implied, is made regarding future performance. Information, opinions and estimates contained in this report reflect a determination at its original date of publication by Sacra and are subject to change without notice.

Sacra accepts no liability for loss arising from the use of the material presented in this report, except that this exclusion of liability does not apply to the extent that liability arises under specific statutes or regulations applicable to Sacra. Sacra may have issued, and may in the future issue, other reports that are inconsistent with, and reach different conclusions from, the information presented in this report. Those reports reflect different assumptions, views and analytical methods of the analysts who prepared them and Sacra is under no obligation to ensure that such other reports are brought to the attention of any recipient of this report.

All rights reserved. All material presented in this report, unless specifically indicated otherwise is under copyright to Sacra. Sacra reserves any and all intellectual property rights in the report. All trademarks, service marks and logos used in this report are trademarks or service marks or registered trademarks or service marks of Sacra. Any modification, copying, displaying, distributing, transmitting, publishing, licensing, creating derivative works from, or selling any report is strictly prohibited. None of the material, nor its content, nor any copy of it, may be altered in any way, transmitted to, copied or distributed to any other party, without the prior express written permission of Sacra. Any unauthorized duplication, redistribution or disclosure of this report will result in prosecution.