How we evaluate
Published before any results exist, so you can judge the method independently of the numbers it produces.
Reporting rules
We report measurements only where we have taken them. Where a figure has not been measured on the stated hardware and dataset, the cell shows a dash rather than an estimate. Cost figures are computed from list GPU prices and will differ from yours.
We publish cases where our models lose to a frontier baseline. A benchmark that only contains wins is marketing, and the people we want as customers know that.
Dataset construction
Real call recordings from a design partner, not synthetic dialogue. Held-out test set separated at the conversation level rather than the turn level, so no call contributes to both training and evaluation. Human labels from reviewers who work in the domain, with inter-annotator agreement reported alongside results.
Personally identifying information is removed before any label review. Nothing from a partner's archive is published; only aggregate metrics, and only with written permission.
What we measure
| Metric | Definition | Why it matters |
|---|---|---|
| Objection accuracy | Agreement with human label on objection class | Task correctness |
| Script fidelity | Correct script per language run | Downstream data integrity |
| Entity preservation | Named entities intact through the turn | Corrupted names break records |
| Compliance failure | Responses breaching approved language | The regulatory risk |
| ASR word error | Transcription error on code-mixed audio | Upstream ceiling on everything |
| Turn latency | End to end, per stage | Whether it sounds like a conversation |
| Cost per call | Amortised GPU vs per-token | The economic case |
Baselines: a current frontier API, and a strong open model served without fine-tuning, so the reported gain is attributable to fine-tuning rather than to self-hosting.
Reproducibility
The evaluation harness is released to design partners with their model, and we intend to open-source the harness itself once the first benchmark publishes. Model versions and prompt templates are pinned and reported with every result.