Data boundaries, a plain explanation for site owners
What you can and cannot honestly tell your visitors about an AI assistant on your site, without a legal team.
Vishal ChiniwarCo-founder and CTO2026-02-041,396 words
When you put an assistant on your website, you take on a position you may not have thought about explicitly: you are the party your visitors are trusting, regardless of who built the software. They did not choose your vendors. They chose to type something into your site.
The boundary you control
Three decisions are genuinely yours, and they are the ones with the most effect on exposure:
**What you index.** If you upload internal documents to make the assistant more capable, you have moved private content into a pipeline that includes third parties. Sometimes that is fine. It should be deliberate.
**What you ask for.** Every field you collect is data you now hold. A lead form asking for a phone number creates an obligation that no form would not have.
**How long you keep conversations.** Longer retention gives you better analytics and a larger exposure. This is a real tradeoff with no universally correct answer.
The boundary your vendor controls
Model provider terms, retention at the provider, processing location, and subprocessor relationships. You cannot change these but you can ask about them, and you should before you publish anything describing them.
The specific questions worth asking any vendor are in the article on vendor evaluation. The short version: ask for it in writing, and treat a vague answer as an answer.
What to tell visitors
A short, plain disclosure at the point of interaction is better than a long policy nobody opens. Something that covers: that they are talking to an AI assistant, that the conversation is stored, roughly what for, and how to reach a human instead.
Three sentences near the chat input does more real work than three pages of policy, and the two are not alternatives. You need the policy as well. The disclosure is what people actually read.
Things not to say unless you can evidence them
- That conversations are not used for model training. This depends on provider terms you may not have verified.
- That data never leaves a particular region. Verify the processing location first.
- That you are compliant with a named framework. Compliance is a determination, not a description.
- That data is deleted immediately. Deletion propagation through backups and provider systems takes time.
- That the assistant cannot make mistakes. It can.
Each of these is commonly written on websites and each is checkable. The cost of being caught overstating data handling is disproportionate to whatever the claim was worth.
The honest position for an early-stage product
If you are launching something and the details are not settled, say so. “We are finalising our provider agreements and will publish retention terms before anyone is charged” is a sentence that costs nothing and protects everything.
Visitors do not expect an early-stage product to have a mature compliance posture. They do expect it not to lie about having one.
What a data boundary actually is
A data boundary is a plain statement of what data goes where, who processes it, how long it is kept and what leaves your control. It is not a compliance badge and it is not a paragraph in a privacy policy written to be unfalsifiable.
The test of a good one is whether a technically literate customer could draw the diagram from your description. If they cannot, the description is marketing.
Most buyers ask about this before they ask about accuracy, which surprises teams who expected the product conversation to be about answer quality.
The four questions worth answering explicitly
What is stored. Conversation transcripts, retrieved passages, contact details if captured, and the index built from your content. Each of those is a different retention question.
Who processes it. Every model provider in the path, every hosting platform, every analytics or error monitoring service. This is the subprocessor list and it needs to be current, not written once.
How long it is kept. Retention is a product decision with legal consequences, not a setting. Shorter retention costs you the ability to debug and to find content gaps, so the honest answer involves a tradeoff rather than a reassuring number.
What crosses a boundary you do not control. When a question and its retrieved passages go to a model provider, that is a boundary crossing, and whether the provider retains it depends on their terms rather than yours.
Reading a vendor's answer
Vague verbs are the signal. Data is handled securely, information may be processed, we take privacy seriously. None of these is falsifiable and none tells you where anything goes.
Ask for the subprocessor list rather than the assurance. A vendor who cannot produce one either does not know their own chain or does not want to publish it, and both are answers.
Ask what happens to a transcript containing something a visitor should not have typed, because they will. Whether redaction happens at write time or read time is the difference between the value never being stored and it being stored and hidden.
And ask which claims are contractual and which are current practice. Practice can change without notice. That question is uncomfortable to ask and uncomfortable to answer, which is why it separates vendors.
Being on the answering side
If you are the vendor, the honest early stage position is to describe design intent and to name what is not settled, separately and clearly.
Saying that a signed data processing agreement is planned before paid rollout is credible. Pointing at a template and implying it is in force is not, and a buyer's legal team will find the difference.
Naming a compliance framework you have not been audited against is the fastest way to lose a technical buyer, because it tells them your other claims need checking too.
What visitors should be told, and where
Most of this discussion is about what you tell a buyer. There is a separate and smaller question about what you tell the person typing into the widget.
The answer is less than teams fear and more than most provide. A short line at the start of the conversation saying this is an AI assistant, that the conversation is stored, and linking to the privacy page. One sentence, not a consent wall.
The reason to be brief is that a long disclosure at the point of a question is read by nobody and reduces the chance the question gets asked at all. The reason to include it is that visitors paste things you did not ask for, and the disclosure is what makes storing it defensible.
Where a conversation captures contact details, the disclosure has to be at the point of capture rather than at the start, because that is where the visitor is actually deciding.
Terms in this articleIndex
Related reading
Retention decisions are product decisions
How long you keep conversation data determines what analysis is possible. Copying someone else's number decides your product by accident.Vishal Chiniwar2026-01-21Where your website content goes when you train an assistant
Draw the data flow before you promise anything about it. Here is the map, and the questions it forces you to answer.Vishal Chiniwar2026-02-11
Questions
Signal
Have a question about how this works?
Access is by waitlist, demo request or public-content pilot. Ask directly and we will answer, including where it will not fit.