How to Audit an AI Vendor's Privacy
Use this AI vendor privacy checklist to trace where business data goes, what the provider keeps, who can access it, and what evidence to demand before buying.
By Mark Wellington · July 26, 2026 · 9 min read
The short answer
To audit an AI vendor's privacy, map the exact product and data route before reviewing broad company claims. Identify what enters the system, where it is processed, who can access it, whether it may be reused, how long each data type remains, how deletion works, and what happens during an incident or exit. Require evidence that applies to your product, plan, configuration, and contract.
What to take away
- →Audit the exact product, plan, feature, and integration you will use; company-wide privacy language is too broad for a buying decision.
- →Trace prompts, files, outputs, logs, feedback, connected-app data, and support records separately because they may follow different rules.
- →Treat policy pages as a starting point and ask for product documentation, contract terms, configured controls, and testable behavior.
- →A good vendor answer names exceptions, owners, and deletion mechanics instead of hiding behind a one-word yes or no.
The AI vendor says your data is “private.”
Lovely. That word is doing the work of a moving crew.
Before you put customer conversations, employee documents, contracts, financial context, or internal know-how into an AI product, you need a more useful answer. Audit the exact product and the complete data route: what enters, where it goes, who can touch it, whether it can be reused, how long it stays, and how you get it out.
That is the heart of an AI vendor privacy checklist. You are not trying to become a privacy attorney over lunch. You are trying to make a responsible buying decision with evidence strong enough for the sensitivity of the data and the consequences of getting it wrong.
“We don’t train on your data” is one answer, not the whole audit
Training gets most of the attention because it is easy to ask about and easy to put in a headline. It is not the only way your information can hang around.
An AI system may handle:
- prompts and uploaded files;
- generated outputs;
- conversation history;
- application logs and abuse-monitoring records;
- connected-drive or CRM data;
- feedback, support tickets, and bug reports;
- cached content, embeddings, or saved project knowledge.
Those data types may have different retention rules, access paths, deletion behavior, and contract terms. A provider can truthfully say it does not train its models on your business data while still retaining some information to deliver a feature, investigate abuse, preserve conversation history, or satisfy a legal obligation.
The useful question is not, “Is this vendor private?”
It is:
Does the exact data behavior of this product match what our business is willing and allowed to share?
The NIST Privacy Framework guidance recommends defining your own prioritized privacy requirements, evaluating suppliers against them, and making residual risk visible when a product does not meet every objective. That is a much stronger buying posture than collecting trust badges and hoping they add up to safety.
Draw the route before you send the questionnaire
Vendor reviews often begin with a spreadsheet containing 83 questions that nobody has connected to the actual use case. The vendor returns a small novel. The buyer files it somewhere responsible-looking. Everyone remains confused.
Start with one real workflow instead.
Suppose your team wants an AI assistant to summarize intake documents and draft a case brief. Draw the route:
- An employee signs in.
- A customer document is uploaded.
- The product sends content to a model or subprocessor.
- An output is generated and saved.
- A person reviews and exports the brief.
- The original file, output, and logs reach their retention or deletion point.
Now mark every place the data is copied, transformed, stored, viewed, or passed to another company. Add identity, permissions, integrations, backups, support access, and the moment the customer relationship ends.
You have just turned “privacy” from a slogan into an inspectable route.
The FTC’s small-business cybersecurity guidance makes vendor oversight similarly concrete: put security expectations in writing, specify how data may be used and retained, limit access, verify compliance, and keep reviewing vendors as conditions change.
Run the six-stop Data Route Audit
Use the route to ask six groups of questions. Do not accept a pile of documents as a substitute for clear answers. Your worksheet should identify the answer, the supporting evidence, the person responsible for configuring or enforcing it, and any risk that remains.
1. Product: which exact service are we buying?
Write down the product, plan, deployment method, model, optional features, and integrations. “We use Vendor X” is not specific enough.
Ask:
- Is this a consumer account, business workspace, enterprise plan, or API?
- Which features send data to a different service or subprocessor?
- Do connectors, web search, agents, file storage, or feedback tools change the rules?
- Which controls are included, optional, or available only by agreement?
This first stop prevents the most common category error: applying a company-level promise to a product surface it does not cover.
2. Inputs: what information are we actually exposing?
Inventory the content before evaluating the container.
Separate ordinary business context from customer personal information, employee records, confidential strategy, credentials, regulated data, legal material, or information owned by someone else. Then ask whether the system truly needs each category.
Good privacy design often begins with subtraction. Remove unnecessary fields. Redact sensitive details. Use representative test data before production data. Give the tool the smallest useful slice instead of a backstage pass to the business.
3. Processing: where does the data travel?
Ask the vendor to describe the processing path in plain language.
- Where is data processed and stored?
- Which subprocessors can receive it?
- Does data cross regions?
- Is it encrypted in transit and at rest?
- Can support, safety, or operations personnel review content?
- Do connected tools create their own copies?
“Hosted in the cloud” is not a route. It is weather.

4. Access and reuse: who may use the content, and for what?
Separate model training from every other form of access or reuse.
Ask whether inputs or outputs may be used for model improvement, safety review, product analytics, support, benchmarking, or research. Identify defaults, opt-ins, opt-outs, exceptions, and who has authority to change the setting.
Current provider pages show why product-level precision matters. OpenAI states that its listed business products and API platform are not used for model training by default, while also describing retention controls and product-specific access features for qualifying organizations. That is useful evidence for those named services; it is not permission to assume every OpenAI-branded feature, account type, connector, or future configuration behaves identically.
The same discipline applies to every vendor: write down the scope of the promise and the exceptions beside it.
5. Retention and deletion: what remains, for how long, and where?
“We delete your data” needs a clock and a map.
Ask for retention periods for prompts, outputs, files, logs, backups, feedback, and support records. Then ask:
- What starts the retention clock?
- Does deletion remove data from active systems, backups, or both?
- Are safety, legal, or abuse-related exceptions possible?
- Can an administrator set a shorter period?
- Does deleting a user also delete that user’s content?
- Can the vendor prove or report that deletion occurred?
Anthropic’s current commercial-data guidance, for example, describes different retention behavior for API inputs, saved product conversations, deleted chats, feedback, and policy-related exceptions. The point is not that one vendor’s schedule should become your universal standard. The point is that “retention” is rarely one number.
6. Exit and incident response: how do we leave, pause, or recover?
A privacy review that covers onboarding but ignores departure is a hotel with no checkout desk.
Ask:
- Can we export our data in a usable form?
- How do we revoke integrations, keys, and user access?
- What gets deleted when the agreement ends?
- How quickly will the vendor notify us of a relevant incident?
- Who investigates, communicates, and preserves evidence?
- Can we suspend processing while a problem is reviewed?
- Which obligations survive termination?
For sensitive, regulated, or high-consequence use, bring qualified privacy, security, and legal professionals into the decision. This checklist helps you expose the questions; it does not decide which laws or contractual terms apply to your business.
Use an evidence ladder, not a logo collection
Not all proof deserves the same weight. Rank vendor evidence in five levels:
| Level | Evidence | What it tells you |
|---|---|---|
| 1 | Sales statement or overview page | The provider’s broad position |
| 2 | Product-specific documentation | How the named feature is supposed to behave |
| 3 | Contract, DPA, or order-form term | What the parties have formally agreed |
| 4 | Configured administrative control | What your account is actually set to do |
| 5 | Test, log, report, or verified behavior | Whether the control works in practice |
You will not need level-five evidence for every low-risk tool. The strength of proof should rise with the sensitivity of the data, the authority of the system, and the cost of a mistake.
But do not let a polished level-one statement impersonate a level-four control.
For every important requirement, record:
- Requirement: what must be true?
- Vendor answer: what did they say?
- Evidence: which level supports it?
- Owner: who configures or monitors it?
- Gap: what remains uncertain?
- Decision: buy, constrain, test, escalate, or reject?
That last column matters. A review is not complete when the questions are answered. It is complete when the answers change the decision.
A smaller permission envelope can rescue a useful product
A vendor does not have to satisfy your most demanding standard for every possible data type if you can design a safer use.
You might:
- prohibit customer-identifying data;
- use a business plan instead of personal accounts;
- turn off an optional connector;
- keep source documents in your own system;
- send redacted excerpts instead of complete records;
- require human review before an external action;
- begin with one low-consequence workflow while stronger controls are evaluated.
That is not hand-waving. It is architecture.
The correct decision may be “yes, inside these boundaries.” It may also be “not until the contract changes,” “not for this data,” or simply “no.” A useful audit makes those boundaries visible before employees create their own unofficial version by pasting sensitive material into whichever tool answers fastest.
Make privacy part of the buying decision, not cleanup work
The best time to discover that deletion is vague, access is broad, or a connector changes the data path is before the workflow depends on the product.
Pick the exact AI use case. Draw its data route. Run the six stops. Climb the evidence ladder only as high as the risk demands. Then document the boundary your team can actually follow.
If the vendor cannot explain where the data goes, you are not being difficult. You are asking the product to reveal what you are buying.
And if you need help turning business requirements, vendor claims, and technical controls into one practical decision, AI consulting can help map the use case, expose the gaps, and define a safer path forward without turning the process into theater.
References
- [S01] Using Privacy Framework 1.1 — National Institute of Standards and Technology, Created December 30, 2024; updated April 14, 2025. Accessed 2026-07-26.
- [S02] Cybersecurity for Small Business — Federal Trade Commission, Current business guidance reviewed July 26, 2026. Accessed 2026-07-26.
- [S03] Business Data Privacy, Security, and Compliance — OpenAI, Current product guidance reviewed July 26, 2026. Accessed 2026-07-26.
- [S04] How Long Do You Store My Organization's Data? — Anthropic, Updated July 2026. Accessed 2026-07-26.