Procurement Barriers Favor In-Provider Evals
Vals AI
The real threat is not that provider built evals are better, it is that they are easier to approve. A bank, hospital, or government team can run OpenAI, Vertex, or Anthropic evaluation workflows inside an existing model contract, cloud account, and security review, instead of sending prompts, documents, traces, and internal workflows to a separate vendor. In practice, that can beat a sharper benchmark if procurement wants one less data sharing exception to sign off.
-
OpenAI and Google already package evals as native product features. OpenAI exposes an Evals API for creating and running evaluations on platform data schemas, while Vertex AI returns row level and summary metrics inside its own evaluation workflow. That means the buyer is adding a feature, not a new vendor.
-
Anthropic also positions evaluation inside its own enterprise stack, and the same pattern shows up across adjacent tooling. LangSmith wins many engineering led teams because evals sit next to tracing and app development, while self host and open source options appeal in regulated settings where third party SaaS review is painful.
-
This is especially important for Vals AI because its core use cases need the hardest data to share, legal files, financial workflows, health documents, and internal repositories. Those are exactly the materials that trigger vendor risk review, data residency questions, and slower public sector procurement.
The market is likely to split in two. Default evals will move into model and cloud platforms for single provider deployments, while independent vendors will need to win where neutrality matters enough to justify separate procurement, especially multi model testing, regulated audits, and benchmarks that customers want kept outside any one model provider.