Inside one call, the agent classifies intent, answers the customer, checks policy and writes a summary. The Spine model router sends each of those tasks to the model that fits it best on cost, speed and risk, from open-weight models on our own GPU capacity to commercial models, in the region your data has to stay in.
Mujhe EMI thoda aage shift karna hai, possible hai?
Haan, main check karti hoon. Aapka account verify ho gaya hai, ek minute dijiye.
Using the largest model for everything is slow and expensive. Using the cheapest for everything is risky. Routing per task gets you both a lower cost per contact and a safer call.
Routine steps such as intent detection go to small, efficient models. Why it matters: model spend tracks the value of each turn, which keeps cost per contact predictable.
Turns on the live call go to models that answer fast, about 0.7 seconds per turn in our deployment testing. Why it matters: pauses make customers talk over the agent or hang up.
Anything near money, disclosures or vulnerable customers gets stricter models and policy checks, and consequential writes still pass the write gate. Why it matters: compliance risk is managed per step, not per vendor.
| Task | Model | Reason |
|---|---|---|
| Intent detection | open-weight | fast and low cost |
| Customer reply | Qwen, own GPU | latency and residency |
| Offer eligibility | rules first | high consequence |
| After-call summary | best fit | quality over speed |
We run open-weight models such as Qwen on GPU capacity we operate in Google Cloud, and route to commercial models when a task is better served by them. The platform is model-pluggable, so a better model is a routing change, not a rebuild.
Data protection rules increasingly decide where a call can be processed. The router respects residency as a hard rule, not a preference.
Calls for Indian lenders and collectors can be processed in the Google Cloud Mumbai region on our own GPU capacity. That supports obligations under India’s DPDP Act.
Residency is part of each client’s configuration, and the router will not send a task to a model outside it.
Routing policies are versioned, reviewed and mapped to your controls in the Compliance Center.
In the demo we take a live call and open the routing record: the model, the reason and the region for every step.