Start with the intent inventory
Before choosing a model, cluster real conversations. In every project we have run, fewer than a dozen intents cover three quarters of contacts. Automate those extremely well and escalate the rest.
Evaluation is the product
Build a graded test set of at least 300 real questions with expected behaviours. Run it on every prompt change. Without this you are shipping vibes.
Track refusal quality too: a confident wrong answer is far more expensive than a graceful handoff.
Design the human handoff first
Decide the escalation SLA before the happy path. Users forgive an assistant that says 'let me get someone'; they do not forgive being trapped.