Vaakya
Evidence

How we evaluate

Published before any results exist, so you can judge the method independently of the numbers it produces.

Reporting rules

We report measurements only where we have taken them. Where a figure has not been measured on the stated hardware and dataset, the cell shows a dash rather than an estimate. Cost figures are computed from list GPU prices and will differ from yours.

We publish cases where our models lose to a frontier baseline. A benchmark that only contains wins is marketing, and the people we want as customers know that.

Dataset construction

Real call recordings from a design partner, not synthetic dialogue. Held-out test set separated at the conversation level rather than the turn level, so no call contributes to both training and evaluation. Human labels from reviewers who work in the domain, with inter-annotator agreement reported alongside results.

Personally identifying information is removed before any label review. Nothing from a partner's archive is published; only aggregate metrics, and only with written permission.

What we measure

MetricDefinitionWhy it matters
Objection accuracyAgreement with human label on objection classTask correctness
Script fidelityCorrect script per language runDownstream data integrity
Entity preservationNamed entities intact through the turnCorrupted names break records
Compliance failureResponses breaching approved languageThe regulatory risk
ASR word errorTranscription error on code-mixed audioUpstream ceiling on everything
Turn latencyEnd to end, per stageWhether it sounds like a conversation
Cost per callAmortised GPU vs per-tokenThe economic case

Baselines: a current frontier API, and a strong open model served without fine-tuning, so the reported gain is attributable to fine-tuning rather than to self-hosting.

Reproducibility

The evaluation harness is released to design partners with their model, and we intend to open-source the harness itself once the first benchmark publishes. Model versions and prompt templates are pinned and reported with every result.