Vals builds proprietary vertical benchmarks
Vals AI
This is how Vals turns evaluation work into a compounding moat. A new benchmark is not just a leaderboard page, it is a reusable pile of hard to source examples, answer keys, and grading logic for a specific job like reading Canadian case law or long credit agreements. That gives Vals something general eval tools do not naturally accumulate, domain specific test data that can be reused across model launches, customer sales, and future product expansion.
-
Vals publishes vertical benchmarks across legal, finance, healthcare, coding, education, and public sector work, and many are built with industry experts on non public datasets. That makes each benchmark useful as proof of expertise in a sales process and as an internal data asset for the next benchmark version.
-
Braintrust and Langfuse are built to help teams run evals, trace failures, and catch regressions in their own apps. Their core product is workflow infrastructure, not a growing library of proprietary vertical test sets. That makes Vals more like a benchmark publisher with software attached than a pure observability vendor.
-
The legal example shows why this matters internationally. Vals already covers Canadian case law, and adjacent legal software companies are using country specific data assets to enter new markets because local courts, rules, and document formats differ in ways generic benchmarks miss.
If Vals keeps adding verticals and jurisdictions, the benchmark catalog can become the front door for model buyers and the training ground for better enterprise eval products. The winner in this layer is likely the company with the deepest labeled task data in the most valuable workflows, not the company with the most generic tracing features.