Creation Platforms Internalize Evaluation
Design Arena
The real threat is that evaluation can become a built in exhaust stream of the product itself, not a separate destination users visit. Adobe, Canva, Figma, Gamma, Vercel, OpenAI, and Anthropic already run creation surfaces where people generate, edit, save, publish, and discard outputs. Those actions reveal which model output actually helped someone finish work, which is often a stronger signal than a public upvote because it is tied to a real task and a downstream outcome.
-
Creation products see the whole workflow. Canva lets users generate assets, then download, share, or continue editing them in the editor. OpenAI offers image generation and editing in ChatGPT, plus Canvas for iterative writing and coding. Anthropic Artifacts gives Claude users a dedicated place to create, revise, organize, and share outputs. Those steps create implicit labels at industrial scale.
-
Installed distribution matters as much as model quality. Canva was at $4B ARR by the end of 2025, Figma at $1.05B revenue in 2025, and Vercel at $340M annualized revenue by February 2026. Gamma reached about $102M ARR by October 2025. If platforms at that scale add side by side generation or routing across models, they can build their own arenas from existing traffic instead of sending users to a standalone evaluator.
-
This is the same bundling pattern already reshaping adjacent AI creation categories. Gamma defined AI slides, then the category was absorbed into larger suites and labs, while Vercel turned AI app creation into a native part of deployment workflows. When the workflow owner also owns the evaluation loop, it can improve prompts, routing, fine tuning, and default model selection without buying a separate ranking product.
The market is heading toward private, workflow embedded arenas that sit inside design editors, coding products, and chat based creation tools. The winners will be the products that turn every generation, edit, and publish event into training and routing data. That pushes standalone evaluators toward higher trust use cases, cross platform neutrality, and benchmarking that workflow owners cannot easily reproduce on their own.