Model router

No single model is right for every turn. So we do not pick just one.

Inside one call, the agent classifies intent, answers the customer, checks policy and writes a summary. The Spine model router sends each of those tasks to the model that fits it best on cost, speed and risk, from open-weight models on our own GPU capacity to commercial models, in the region your data has to stay in.

spine·voice · router
region mumbai · policy v5
Model router · call #20931live
Customer

Mujhe EMI thoda aage shift karna hai, possible hai?

route · intent → open-weight, own GPU · low latency
Agent

Haan, main check karti hoon. Aapka account verify ho gaya hai, ek minute dijiye.

route · reply → open-weight (Qwen) · GCP Mumbai
route · EMI change → policy check, then write gate
data in region3 models usedwrite gated
intent → small model · reply → open-weight · summary → best fit · residency: india ✓ · audit:chained ✓ ·
Hinglish lending call, simulated · rendered by spine·voice
then the router logs every choice →
SPINE · ROUTING · call #20931live
turn 03Reply to customeropen-weight model, own capacityAuto · read
turn 05Reschedule instalmentpolicy check, then specialistHeld · human
✓ model, version and region recorded · audit entry signed
How routing decides

Three questions for every task: what does it cost, how fast must it be and how much can go wrong.

Using the largest model for everything is slow and expensive. Using the cheapest for everything is risky. Routing per task gets you both a lower cost per contact and a safer call.

01 · Cost

Pay for what the task needs

Routine steps such as intent detection go to small, efficient models. Why it matters: model spend tracks the value of each turn, which keeps cost per contact predictable.

02 · Latency

Keep the conversation natural

Turns on the live call go to models that answer fast, about 0.7 seconds per turn in our deployment testing. Why it matters: pauses make customers talk over the agent or hang up.

03 · Risk

Match care to consequence

Anything near money, disclosures or vulnerable customers gets stricter models and policy checks, and consequential writes still pass the write gate. Why it matters: compliance risk is managed per step, not per vendor.

spine·voice // routing policytenant lender_in_
Model router · policy v5sample
TaskModelReason
Intent detectionopen-weightfast and low cost
Customer replyQwen, own GPUlatency and residency
Offer eligibilityrules firsthigh consequence
After-call summarybest fitquality over speed
region: India (Mumbai)Edit policy
Routing policy, sample data · rendered by spine·voice
Open and commercial models

Open-weight models on our own GPUs, commercial models where they earn it.

We run open-weight models such as Qwen on GPU capacity we operate in Google Cloud, and route to commercial models when a task is better served by them. The platform is model-pluggable, so a better model is a routing change, not a rebuild.

  • No lock-in to one model vendorPrice changes or a model retirement upstream do not stall your contact center.
  • Capacity we controlOpen-weight models on our own GPU capacity keep live-call latency and cost under our management, not a shared queue.
  • Every model change is testedA new model goes through simulation and evaluation and a ramp gate like any other release.
  • Every choice is recordedThe model, version and region used for each turn land on the audit chain, so you can explain any answer later.
Data residency

Your customers’ conversations stay where the law says they should.

Data protection rules increasingly decide where a call can be processed. The router respects residency as a hard rule, not a preference.

✓India, from Mumbai

Calls for Indian lenders and collectors can be processed in the Google Cloud Mumbai region on our own GPU capacity. That supports obligations under India’s DPDP Act.

✓Region set per client

Residency is part of each client’s configuration, and the router will not send a task to a model outside it.

✓Governed like everything else

Routing policies are versioned, reviewed and mapped to your controls in the Compliance Center.

Book a demo

See which model answered each turn, and why.

In the demo we take a live call and open the routing record: the model, the reason and the region for every step.