NobleCloak
Playbooks

The AI vendor due-diligence questions that actually matter

Author

Audrey

NobleCloak's AI Correspondent · AI-drafted, fact-checked against our sourced evidence before publishing.

Date Published

There is no shortage of AI vendor questionnaires. Ncontracts publishes a good free one; Vanta shipped an AI Security Assessment template in June 2026; every GRC platform has a bank of AI questions. The problem was never finding questions. The problem is that most of them get answered in marketing language — "we take security seriously," "your data is encrypted," "we're SOC 2 compliant" — that reads like an answer and tells you nothing you could defend at an exam.

This playbook is the short list. Twelve questions, chosen because a real answer to each changes your risk assessment, paired with what a genuine answer looks like versus the marketing version that fails. If a vendor can't answer these — or answers in brochure — that itself is a finding. Proportionality applies throughout: NCUA's 07-CU-13 framework expects diligence scaled to the vendor's risk, and Reg S-P asks for "reasonable" oversight. A grammar checker gets three of these; the AI transcription tool reading eight people's calendars gets all twelve.

The twelve

1. What data does your product access, and at what scope? Real answer: a specific list of data types and OAuth scopes ("read-only calendar and Drive metadata; we do not read file contents unless a user attaches a file"). Marketing answer: "only the data needed to provide the service." That's a non-answer. Cross-check it against what the OAuth grant actually shows in your console — see Reading OAuth grants — the shadow-AI map in your Workspace and Entra. If the console says Drive read and the vendor says "we don't touch your files," someone's wrong, and you need to know which.

2. Do you use our data — or our customers' data — to train your models? The single most important question. Real answer: an unambiguous yes or no, with a contractual commitment and a way to opt out. "No, customer data is never used for training; this is in our DPA at section X." Marketing answer: "we may use aggregated, anonymized data to improve our services." That's a yes wearing a disguise — "improve our services" is training, and "anonymized" is doing a lot of undefined work.

3. Where is our data stored, for how long, and how is it deleted? Real answer: region, retention period, and a deletion mechanism with a timeframe. Marketing answer: "in secure, industry-leading data centers." For the Otter.ai case — recordings retained — this is the question that turns "retained" from a vague worry into a documented fact you can act on.

4. Which sub-processors and third-party AI models do you rely on? Most AI vendors are wrappers around someone else's model (OpenAI, Anthropic, Google) plus a stack of sub-processors. Real answer: a current sub-processor list and the underlying model providers, with a commitment to notify you of changes. Marketing answer: silence, or "we use best-in-class partners." You are inheriting every one of those relationships; you're entitled to the list.

5. Can you provide a SOC 2 Type II, ISO 27001, or equivalent attestation? Real answer: a current report under NDA, or an honest "not yet, here's our timeline and what we do instead." Marketing answer: "we are SOC 2 compliant" with nothing to show. "Compliant" is not a report. Ask for the report and the date of the audit period; a SOC 2 is a snapshot, and a two-year-old one is nearly worthless.

6. What is your breach-notification commitment? For RIAs this isn't optional — Reg S-P now requires service providers to notify you of a breach within 72 hours. Real answer: a specific SLA in the contract. Marketing answer: "we'll notify you promptly." "Promptly" doesn't survive an exam; the number does. See The Reg S-P service-provider oversight file step by step.

7. What AI features are enabled by default, and what's the update policy? Real answer: a list of AI features and a commitment that new ones are opt-in, not auto-enabled. Marketing answer: "we're always innovating." That's the problem, not the reassurance — vendors adding AI silently on update is how your surface expands without your knowledge.

8. Can a human at your company access our data, and under what controls? Real answer: named access controls, logging, and the conditions (support tickets, break-glass). Marketing answer: "access is strictly limited." Ask who, when, and whether it's logged.

9. How do you handle our data in your AI processing — is it isolated per customer? Real answer: a description of tenant isolation and whether prompts/outputs are logged or retained. Marketing answer: "your data is kept private and secure." Private from whom, and retained where?

10. What happens to our data when we terminate? Real answer: a deletion commitment with a timeframe and, ideally, a certificate of destruction. Marketing answer: "you can export your data anytime." Export is not deletion. You want to know what they keep after you leave.

11. Do you carry cyber-liability insurance, and will you indemnify for an AI-caused data incident? Real answer: coverage limits and contractual indemnity language. Marketing answer: "we maintain appropriate insurance." A number and a clause, or it doesn't count.

12. Can you point to how your AI governance maps to a recognized framework? Not "are you certified" — there's no AI certification that means much yet — but whether they can speak to NIST AI RMF functions: how they Govern AI risk, Map their systems' impacts, Measure performance and harms, and Manage incidents (NIST). Real answer: a vendor who can describe their own governance in those terms. Marketing answer: "we follow NIST" as a logo on a slide. The framework is a vocabulary for a real conversation, not a badge.

How to tell a real answer from a marketing one

A pattern runs through all twelve. Real answers are specific, falsifiable, and often contractual — a number, a scope, a section reference, a document you can hold. Marketing answers are adjectival and unverifiable — "secure," "industry-leading," "best-in-class," "compliant," "appropriate." Three tests:

  • Can I check it independently? A claimed OAuth scope you can verify in your console. A claimed SOC 2 you can read. A claimed retention period you can find in the DPA. If the only evidence is the vendor's own adjective, it's not evidence.
  • Is it in the contract? A promise in an email is a hope; a promise in the DPA is a right. Push the answers that matter — training, breach notice, deletion — into the contract.
  • Does it name a mechanism, or a feeling? "Encrypted at rest with customer-managed keys" is a mechanism. "We take your privacy seriously" is a feeling. Score mechanisms; discard feelings.

When a vendor won't answer, the non-answer is your finding. You don't need them to confess to a risk; a refusal to specify data-use or breach-notice terms is, itself, a documented diligence result. Write it down: "Vendor declined to confirm training use in writing as of [date]" is a defensible, dated observation an examiner respects far more than a checkbox marked "yes."

The honest part

Twelve good questions raise your diligence from a checkbox to a conversation — but the answers are still the vendor's own words. A questionnaire is a self-attestation; it's only as true as the vendor is honest, and only current as of the day it's answered. It doesn't replace a SOC 2, and a SOC 2 doesn't replace ongoing monitoring, because a vendor who answered clean in January can flip on a training feature in March. This is one layer — the ask — sitting on top of the observe layer (what the OAuth grants show) and under the monitor layer (re-running both over time). No single layer is the whole program. But asking the right twelve, and writing down the answers and the refusals, is diligence an examiner can see.


Chasing twelve real answers out of every AI vendor — and independently verifying the checkable ones against what the tool actually does — is precisely the work the research library behind every Discover report exists to carry. We come to the vendor conversation already knowing what they've told other institutions, what their DPA says, and where the marketing answer hides the real one.

Request your Discover scan — we run it with you. Then file the answers into What an examiner-ready AI evidence binder contains.