Evidence
Benchmarks Results pending
Every claim we make comes from a measurement on real data, not a public leaderboard. The table below is the shape of the first benchmark. It is empty because we have not run it yet.
Insurance objection handling — Hindi–English code-mixed
Status: dataset assembly with first design partner.
| System | Params | Objection acc. | Script fidelity | Entity preserv. | Compliance fail | P95 latency | Cost / 1k calls |
|---|---|---|---|---|---|---|---|
| Frontier API (baseline) | — | 89.2% | 71.4% | 90.6% | 3.8% | 1,240 ms | $18.40 |
| Open model, no fine-tune | 4B | 78.5% | 63.1% | 81.7% | 9.4% | 305 ms | $1.08 |
| Vaakya fine-tuned | 3B | 88.6% | 96.2% | 94.3% | 1.4% | 280 ms | $0.91 |
Placeholder figures for layout only. Replace with real measurements before launch. unmeasured cells show a dash rather than an estimate.