When to hand off to a human, and how to decide

Handoff triggers should be written rules you can audit, not a judgement made fresh each time.

Sachin Aathreyaa K MCo-founder, CEO and CPO2026-04-291,358 words

Every conversation with an assistant has a moment where continuing costs more than escalating. Finding that moment reliably is the difference between a tool that helps and one that traps people in a loop.

The mistake is treating it as a judgement call. Judgement made fresh each time is inconsistent and impossible to audit. Written rules can be tested and improved.

The five rules worth writing down

1. Two failed attempts

If the assistant has given two answers and the visitor has rephrased twice, escalate. The third attempt rarely succeeds and each one costs patience. This single rule catches most bad experiences.

2. Explicit request

If someone asks for a human, they get one immediately. No qualification step, no “let me try to help first”. Resisting an explicit request is the fastest way to make someone hostile, and it is a pattern people recognise and resent.

3. Detected frustration

Short, repeated or capitalised messages, or profanity, indicate the conversation has gone wrong. Escalate rather than trying to recover it. This does not need a sentiment model to work at a basic level; message length collapsing and repetition are simple signals.

4. Outside the knowledge boundary

If the question concerns something the assistant has no source for, such as account specific data, a legal question, or anything requiring a commitment, escalate instead of approximating. This is a design decision, not a fallback.

5. Commercial signal

A question that indicates a serious evaluation, such as asking about contract terms, security review or team pricing, is worth a person even if the assistant could answer. This is the one rule where the goal is not deflection.

The tension to be honest about

Rule five conflicts with how these tools are usually sold. Deflection rate is the headline metric, and escalating a good lead lowers it. That is a sign the metric is wrong rather than the behaviour.

On a marketing site, a conversation that reaches a salesperson is often the best possible outcome. Measuring success by how many people you avoided talking to gets this exactly backwards.

When nobody is available

Handoff needs a defined behaviour outside working hours, and “the message disappears” is not one. The options are collecting contact details with a stated response time, offering an alternative channel, or being explicit that the team is offline and when they return.

Say the actual expected time. “We will get back to you soon” is worth less than “we reply within one working day”, and if you cannot commit to a number, say that instead of implying one.

Auditing the rules

Once a month, read ten escalated conversations and ten that were not escalated but ran long. The second group is where the rules are failing. If you find conversations that should have escalated and did not, tighten the rule that should have caught them rather than adding a new one.

Writing triggers you can audit

A trigger that says escalate when the assistant seems unsure is not a rule. It is a hope, and it produces behaviour nobody can explain when it goes wrong. The useful triggers are all things you can check in a log afterwards.

Explicit request is the first and it should be unconditional. When a visitor asks for a person, arguing with them, suggesting an article, or asking what the question is about first reads as obstruction. It is the fastest way to make an assistant resented.

Repeated failure is the second. Two consecutive refusals on one thread, or two answers followed by a rephrasing of the same question. Rephrasing is the clearest signal available that an answer missed, and it is easy to detect because the embeddings of the two questions will be close.

Marked topics are the third, and they are configuration rather than inference. Cancellations, complaints, anything contractual, anything involving money already paid. Marking these explicitly is safer than hoping a confidence score catches them, because confidence measures support in the retrieved passages and says nothing about stakes.

The escalation rate that looks like progress

A falling escalation rate is usually reported as the assistant improving. Sometimes it is. Often it means the assistant has become more willing to guess.

The way to tell them apart is to watch refusal rate alongside it. A genuinely improving assistant escalates less because it answers more, so refusals fall too but answers with citations rise to match. An overconfident one shows escalations and refusals both falling with no corresponding rise in cited answers, which means questions that used to reach a person are now being answered by something that does not know it is wrong.

The opposite direction is also misread. Escalations rising after a content change often means the assistant is correctly declining to improvise on a topic your new page half covers. That is the system working, and reacting by loosening the rules undoes it.

What to say at the moment of handoff

The wording matters more than teams expect, because this is the point where a visitor decides whether the assistant was useful or was a delay.

Say that you are handing over, say why in one clause, and say what happens next with a realistic time. Do not apologise at length, do not blame the question, and do not offer one more article as a consolation.

If nobody is available, say so and give the actual response time. An honest queue time beats a promise of immediacy you cannot staff, and it beats a chat window that silently goes nowhere, which is what most implementations do at nine in the evening.

Staffing the promise you make

The failure that damages trust most is not escalating too early or too late. It is escalating into nothing.

Before enabling handoff, decide who receives it, during which hours, and what the visitor is told outside those hours. A chat that promises a person and delivers silence at nine in the evening is worse than an assistant that said it could not help and pointed at a contact page.

If there is no coverage, say so and set a real expectation. Visitors accept a next working day answer. They do not accept an implied immediacy that does not arrive, and they judge the company rather than the widget.

This is a staffing decision wearing a product costume, and it is worth settling before anything is configured.

The rule underneath all of them

Escalate when continuing would require guessing. That single test covers most of the specific triggers above and is the one to fall back on when writing a new one.

Terms in this articleHuman HandoffFallback Model

Questions

An explicit request for a person, repeated failure on one thread, and any topic you have marked always human such as cancellations or anything contractual.

Not on its own. Watch refusal rate alongside it. Both falling with no rise in cited answers means the assistant has become willing to guess.

Say so and give a realistic response time. An honest queue time beats a promise of immediacy you cannot staff.

Repeated failure on the same thread, an explicit request for a person, and topics you marked as always human. Rules you can audit, not a judgement made fresh each time.

Only if frustration is measured as rephrasing rather than guessed from tone. Sentiment detection on short messages is unreliable enough to route people wrongly.

Say so and capture the thread. Silently queueing a visitor who was told a person is coming is worse than not offering.

Signal

Have a question about how this works?

Access is by waitlist, demo request or public-content pilot. Ask directly and we will answer, including where it will not fit.