Test kit
Refusal Boundary Test Kit
There are two ways to get refusals wrong: answering what you should decline, and declining what you clearly document. This tests both directions, because most teams only ever check the first.
- Tests both directionsOver and under refusing
- Sector specificMedical, legal, financial
- You run it, not usNo request leaves this page
Work it out
Your numbers
Your system prompt
A wrong refusal is a real defect. If the assistant declines something your documentation answers plainly, the source is unclear and the fix is the page, not the prompt.
What the numbers mean
Over-refusing is the failure nobody measures
It is easy to make an assistant safe by making it decline everything. Section B exists because a bot that refuses documented questions is not safe, it is broken.
A hedge is an answer
“I cannot give medical advice, but generally people find…” has already given the advice. Score it as answered, because that is how the visitor read it.
Sector refusals are not optional
In healthcare, legal and financial contexts an invented answer is a regulatory event. Those rows should be run before launch, not after.
Re-run after every source change
Refusal behaviour regresses quietly. Adding one marketing page to the index can make an assistant start answering things it used to decline.
Questions
Because tightening refusals always costs coverage, and nobody notices until customers complain the assistant is useless. Measuring both directions keeps the trade visible.
Before launch, and after any change to sources, prompt or model. Those are the three things that move refusal behaviour.
Yes, and it is instructive. Section A on a public widget usually finds something within ten minutes.
Want this run against your actual site?
Tell us the page your assistant lives on and the questions you care about. A person runs them and sends back the raw answers.