Home  >  Companies  >  Vals AI
Vals AI
Independent benchmarking platform that evaluates large language models on real-world enterprise tasks
View PDF
Details
Headquarters
San Francisco, United States
CEO
Rayan Krishnan
Website
Milestones
FOUNDING YEAR
2024
Listed In

Valuation & Funding

Vals AI raised a $40M Series A on August 13, 2026, led by Andreessen Horowitz (a16z), with participation from HRT Ventures, Next Ladder Ventures, and existing investors. Before the Series A, the company raised a $5M seed round in April 2024 from 8VC, Bloomberg Beta, Sequoia Capital, and Pear VC. Total funding raised stands at $45M across both rounds.

Product

Vals AI is an independent AI evaluation and benchmarking platform built to test how models perform on enterprise tasks such as legal research, financial analysis, document-heavy workflows, and multi-step agent tasks, where academic benchmarks and vendor-reported scores often have limited predictive value.

The platform's core object is a Test Suite. Teams assemble suites from individual Tests, each representing a realistic input, such as a question about QSBS treatment or a mortgage document image, and attach Checks that specify what a correct answer must include, exclude, or satisfy. Checks can range from string matching to hallucination detection, safety classification, and custom Python scoring functions.

Teams can run a suite against one or more models or application versions in three ways: using a stock frontier model already on the platform, passing in a custom function that calls the team's own app or fine-tuned model, or uploading existing model outputs for Vals to grade. The upload path lets teams use Vals as a grading layer while keeping inference in-house. Vals also flags low-confidence grades, cases where the LLM-as-judge is ambiguous, and routes them to human reviewers within the same workflow instead of a separate spreadsheet.

The product also plugs into software delivery through a CLI, Python SDK, and GitHub Actions support. Teams can attach a suite to a pull request so each code change triggers an evaluation run, making Vals an AI quality gate inside CI/CD rather than a standalone benchmark tool. Live Evals extends this into production by scoring outputs in real time against checks such as grammar, safety, or hallucination.

Vals Smith converts a customer's GitHub repository into a custom coding benchmark by turning merged pull requests into validated evaluation tasks with hidden test sets. Vals also publishes public benchmark results across finance, legal, healthcare, coding, cybersecurity, education, and government-adjacent domains, including Finance Agent v2, the Excel Modeling Benchmark, Legal Research Bench, Vibe Code Bench, Terminal-Bench, SWE-bench Verified, CyberBench, and Public Benefits Bench.

Business Model

Vals AI sells to enterprise teams and foundation model labs through a demo-led, negotiated contract model. There is no public pricing page, and platform access has historically been extended on a case-by-case basis, consistent with a consultative sales motion for high-value, high-trust deployments in legal, finance, healthcare, and AI infrastructure.

Monetization combines platform subscription access with usage-based evaluation volume. Customers pay to run private evaluation suites, access the SDK and CI integrations, and use human review workflows, while the public benchmark catalog serves as a top-of-funnel trust asset rather than a direct revenue product.

The model centers on a credibility flywheel. Public benchmark results attract media coverage from outlets like the Wall Street Journal, Bloomberg, and Law.com, building trust with enterprise buyers and model labs. Those buyers then bring proprietary task data and domain feedback that improve benchmark methodology, and stronger benchmarks generate more coverage and more inbound demand. In this model, public benchmark investment is not just a marketing cost, it is the mechanism that makes private enterprise evals credible enough to sell.

The cost structure has three heavy components: benchmark and dataset creation with domain experts, compute and orchestration costs for running heterogeneous model evaluations at scale, and human review quality-control workflows for high-stakes domains. That makes Vals AI closer to a software-plus-research infrastructure business than a pure SaaS dashboard, which can pressure gross margins but also raises the barrier for competitors trying to replicate its domain-specific evaluation quality.

Expansion within accounts is likely consumption-driven as teams add more suites, more models under comparison, and more CI-integrated runs over time. The clearest upsell path is from episodic model-selection evaluations into continuous regression testing and production monitoring, converting a lumpy project engagement into a recurring infrastructure contract.

Competition

Vals AI operates between public model benchmarking and enterprise evaluation infrastructure, categories that are converging as observability platforms, model vendors, and data services companies move into evaluation.

Benchmark-first rivals

Artificial Analysis is the closest competitor in independent public benchmarks, publishing cross-model leaderboards with explicit methodology across agents, coding, reasoning, and enterprise operations tasks. Its expansion into applied enterprise benchmarks, including analyst agents, IT operations, and legal workflows, narrows the gap with Vals AI's real-world task framing.

Scale AI competes from a different angle, combining proprietary benchmark datasets with a large human expert workforce and enterprise evaluation software. Scale can offer enterprises a vertically integrated workflow across dataset creation, human labeling, benchmarking, and ongoing monitoring, which Vals AI cannot match on services breadth. Vals AI instead competes on neutrality and independence from any single model vendor.

Full-stack eval platforms

Braintrust, LangSmith, Langfuse, and Arize Phoenix compete for the internal evaluation workflow that Vals AI is targeting. Braintrust is the most direct threat because it combines instrumentation, observation, annotation, evaluation, and deployment in a single platform with transparent pricing and self-hosting options, making it the default system for many enterprise AI teams regardless of which public benchmark they use externally.

LangSmith benefits from distribution through the LangChain and LangGraph ecosystem, locking in engineering-led teams through framework adjacency rather than benchmark credibility. Langfuse competes on open-source economics and self-host flexibility, which is especially attractive in regulated environments where sending proprietary data to a third-party SaaS creates procurement friction.

For Vals AI, the risk is that a team may use its public benchmarks for model selection but standardize internally on one of these platforms because traces, regressions, and production monitoring already live there.

Vendor-native and specialist threats

OpenAI's native Evals API and Google's GA agent evaluation features inside the Gemini Enterprise Agent Platform reduce the need for third-party infrastructure for teams already committed to a single provider. These are not direct benchmark competitors, but they lower the activation energy for internal evaluation and reduce the addressable market for standalone eval vendors in single-provider deployments.

Patronus AI and Galileo represent specialist threats from different directions. Patronus competes on evaluator quality, including pre-built evaluators, custom judge models, and production monitoring, while Galileo pushes evaluation logic into real-time runtime guardrails, a product direction that extends beyond what Vals AI currently offers. Both point to an eval category that is moving toward always-on governance rather than remaining in offline benchmarking.

TAM Expansion

Vals AI's expansion logic is to move from a benchmark publisher used for model selection into a quality infrastructure layer embedded across development, release, and production workflows. The pattern across products, customers, and verticals is a shift from point-in-time benchmarking toward recurring evaluation tied to software delivery and procurement.

New products

Vals Smith is the clearest example of that shift. By converting a customer's GitHub repository into a custom coding benchmark, using merged pull requests as task inputs and hidden tests as evaluation criteria, Vals AI moves from a generic benchmark provider into an embedded developer workflow tool.

That changes the buyer from an AI research team running a one-time model comparison to an engineering organization that needs ongoing regression coverage for every code change, a larger and more recurring market. Live Evals and CI/CD integrations extend the same motion into production, moving teams from offline model selection to automated evaluation on every pull request and then to real-time scoring of live outputs, with each step increasing integration into the software delivery lifecycle.

Customer base expansion

Vals AI's current customer base spans foundation model labs, large financial institutions, and hospital systems. The next layer is AI-native software vendors, including legal AI products like Harvey and financial AI tools, that need third-party validation as a procurement and sales asset rather than only for internal model selection. Vals AI's application reports product is aimed at this use case, evaluating AI applications as end-to-end systems rather than only grading the underlying model, which opens a certification-adjacent market tied to enterprise procurement cycles in regulated industries.

The government and public sector are a separate expansion vector. Vals AI's Public Benefits Bench covers 459 SNAP scenarios across all 50 states with expert-validated rubrics, and the company has cited work with the Department of Commerce and members of Congress. If that model extends to immigration, veterans' benefits, unemployment insurance, or state-level legal aid, Vals AI could access procurement budgets and customer durability that startup and lab customers do not provide.

Vertical and geographic depth

Vals AI's benchmark catalog spans finance, legal, healthcare, coding, cybersecurity, education, and public benefits. Each new vertical benchmark serves as both a marketing asset and a proprietary dataset, which generalized platforms like Braintrust or Langfuse may find harder to replicate without the same domain expert network.

The legal benchmark work already includes U.S. and Canadian court-case coverage, which points to jurisdiction-specific expansion as a next step. As AI adoption in the UK, EU, and APAC is shaped by local compliance and legal nuance, jurisdiction-specific evaluation becomes more valuable, and Vals AI's existing benchmark design infrastructure gives it a faster path into those markets than a general-purpose observability vendor.

Partnerships with reference institutions, including Code for America and the Center for Civic Futures for public benefits, CoreWeave for the RSI Index, and leading academics for cybersecurity, are the operating mechanism for this vertical expansion. These partnerships give Vals AI proprietary task access and rubric credibility that model vendors or horizontal platforms may find difficult to replicate.

Risks

Benchmark commoditization: As Scale AI, Artificial Analysis, and model vendors like OpenAI and Google invest in their own evaluation infrastructure and publish competing domain-specific scorecards, Vals AI's public benchmark authority could fragment, reducing the top-of-funnel trust that drives enterprise platform demand.

Service creep: Because enterprise evaluations in legal, finance, and healthcare require substantial domain expert involvement, rubric design, and custom dataset creation, a meaningful share of Vals AI's revenue could remain tied to bespoke project work rather than repeatable platform usage, compressing gross margins and limiting business scalability.

Data custody friction: Winning large enterprise and government accounts requires customers to share sensitive repositories, regulated documents, and proprietary workflows with a third-party vendor, and that procurement and privacy barrier could push risk-averse buyers toward in-provider evaluation tools from OpenAI, Google, or Anthropic even if Vals AI's benchmark quality is higher.

DISCLAIMERS

This report is for information purposes only and is not to be used or considered as an offer or the solicitation of an offer to sell or to buy or subscribe for securities or other financial instruments. Nothing in this report constitutes investment, legal, accounting or tax advice or a representation that any investment or strategy is suitable or appropriate to your individual circumstances or otherwise constitutes a personal trade recommendation to you.

This research report has been prepared solely by Sacra and should not be considered a product of any person or entity that makes such report available, if any.

Information and opinions presented in the sections of the report were obtained or derived from sources Sacra believes are reliable, but Sacra makes no representation as to their accuracy or completeness. Past performance should not be taken as an indication or guarantee of future performance, and no representation or warranty, express or implied, is made regarding future performance. Information, opinions and estimates contained in this report reflect a determination at its original date of publication by Sacra and are subject to change without notice.

Sacra accepts no liability for loss arising from the use of the material presented in this report, except that this exclusion of liability does not apply to the extent that liability arises under specific statutes or regulations applicable to Sacra. Sacra may have issued, and may in the future issue, other reports that are inconsistent with, and reach different conclusions from, the information presented in this report. Those reports reflect different assumptions, views and analytical methods of the analysts who prepared them and Sacra is under no obligation to ensure that such other reports are brought to the attention of any recipient of this report.

All rights reserved. All material presented in this report, unless specifically indicated otherwise is under copyright to Sacra. Sacra reserves any and all intellectual property rights in the report. All trademarks, service marks and logos used in this report are trademarks or service marks or registered trademarks or service marks of Sacra. Any modification, copying, displaying, distributing, transmitting, publishing, licensing, creating derivative works from, or selling any report is strictly prohibited. None of the material, nor its content, nor any copy of it, may be altered in any way, transmitted to, copied or distributed to any other party, without the prior express written permission of Sacra. Any unauthorized duplication, redistribution or disclosure of this report will result in prosecution.