What an examiner-ready AI evidence binder contains
Author
Audrey
NobleCloak's AI Correspondent · AI-drafted, fact-checked against our sourced evidence before publishing.
Date Published
The compliance officer's real 2026 problem isn't that they've done nothing about AI. It's that they've done things — a policy here, a vendor check there — and have no way to show it as a coherent whole when someone asks. The emotional truth, named plainly: "I'll be asked about AI, and I have nothing to hand over." The fix is not a bigger program. It's an evidence binder: one organized place where everything you've done is dated, sourced, and legible to an outsider in fifteen minutes.
This playbook is the table of contents. Build these seven sections and you can answer "what are you doing about AI?" by sliding a binder across the table instead of giving a speech. It works whether "binder" means a literal tabbed folder, a shared drive, or a PDF — the structure is what matters. Everything here is assembled from the other playbooks in this series; this is where their outputs land.
Why a binder beats a better answer
Examiners don't grade eloquence; they grade evidence. An institution that speaks confidently about its "AI governance approach" but produces nothing scores worse than one that hands over a modest, dated, honest file. The binder does three jobs at once: it proves the work happened (dated artifacts, not assertions), it shows the process is repeatable (a monitoring trail, not a one-time scramble), and it demonstrates candor (you documented what you don't cover, which is the single most credible thing a program can do). Ncontracts' 2026 State of TPRM survey (n=173) found AI risk now tied with cybersecurity as institutions' top third-party concern — the examiner walking in has read the same headlines. The binder is how you meet that with substance instead of a shrug.
The table of contents
Tab 1 — Governance summary (the one page)
The cover artifact: a single page that says how many AI tools you found, how many were shadow (nobody approved them), how many touch sensitive data, your top risks, and what you did about each. Name the owner (usually you), the review cadence, and the date. This is the page a board sees and an examiner reads first. If you build nothing else, build this — but it should sit on top of the tabs that back it up.
Tab 2 — The AI inventory
Your dated list of every AI tool in use: tool, vendor, users, data touched, how it got in (approved / feature-flip / shadow), and a risk flag. This is the output of How to run an AI data call, and it's the factual spine of the whole binder — every other tab refers back to a row in this one. Include the date it was run and the date it's next due, because an inventory with no re-run date reads as a one-time event, and examiners can tell the difference.
Tab 3 — Shadow-AI findings (the OAuth evidence)
The raw observations from your identity provider: the third-party app grants, their scopes, their breadth, and what you did with each. This is where the canonical find lives:
> Otter.ai holds Calendar + Drive read access for 8 users since March; recordings retained; no SOC 2 on file — revoked April 2, grant removed, users notified.
Notice the second half. The finding plus the dated resolution is what makes this a strength rather than a liability. A documented, remediated shadow-AI grant shows the program works; the same grant found by the examiner instead of by you is a write-up. This tab is the output of Reading OAuth grants — the shadow-AI map in your Workspace and Entra.
Tab 4 — Vendor due-diligence records
One dated record per AI vendor that matters: the questionnaire answers, the reviewed attestation (or a documented note that none exists and how you compensated), the sub-processor list, and a risk rating. Built from The AI vendor due-diligence questions that actually matter. Crucially, include the refusals — "vendor declined to confirm training-data use in writing, [date]" is a finding, not a gap, and its presence proves you asked the hard questions rather than the easy ones.
Tab 5 — Policies and procedures
The written documents: your AI-use policy, your service-provider oversight policy, and — for RIAs and BDs — your Reg S-P incident-response program with the 72-hour vendor clock and 30-day customer clock. These come from The Reg S-P service-provider oversight file step by step. Each should be dated and show evidence of approval (board minute reference, principal sign-off). A policy with no approval date is a draft, and examiners read it as one.
Tab 6 — The monitoring trail
The section that separates a real program from a one-time cleanup. Dated log entries showing the inventory was re-run, grants were re-reviewed, vendors were re-checked, and renewals triggered fresh diligence. It doesn't need to be heavy — a few timestamped lines per quarter — but it needs to exist and be current this year. "Ongoing oversight" with no dated trail is indistinguishable from oversight that stopped after onboarding, and that's exactly the gap Reg S-P's monitoring requirement and NCUA's proportional-diligence expectation are written to catch.
Tab 7 — Framework crosswalk (optional but powerful)
A short table mapping your artifacts to a recognized vocabulary — most usefully the NIST AI RMF functions: Govern (your policies and owner), Map (your inventory and data-flow notes), Measure (your risk ratings and diligence findings), Manage (your remediation and monitoring trail) (NIST). This isn't required by any rule — no US regulator mandates AI-specific controls today — but it translates your work into terms an examiner recognizes and shows you're thinking in a real framework rather than improvising. A CU-flavored 07-CU-13 version of this crosswalk works the same way for the proportional-diligence lens.
Assembling it without a compliance department
The binder looks like a lot until you notice it's mostly outputs you already generated running the other playbooks. The order of operations for a 1–2 person program:
- Run the data call → Tab 2.
- Read the OAuth grants → Tab 3.
- Diligence the vendors the first two surfaced as high-risk → Tab 4.
- Write (or dust off) the three policies → Tab 5.
- Start the monitoring log the day you finish, so it has a first entry → Tab 6.
- Add the crosswalk and the one-page summary last → Tabs 7 and 1.
A focused week gets you a defensible first edition. The binder is never "done" — it's a living file you diff each quarter — but a dated first edition beats an eloquent nothing every single time.
The honest part
An evidence binder proves you ran a reasonable, proportional, documented process. It does not prove your vendors are safe, that no AI tool will ever leak data, or that you've caught every shadow tool — and the binder should say so. The most credible tab you can include is a short, honest statement of coverage limits: this program inventories AI reachable through your identity provider and your vendor list; it is point-in-time evidence re-run on a stated cadence; it does not monitor employee use of personal devices or catch AI a vendor runs invisibly on their own backend. Naming what you don't cover is not a weakness in the binder. In a trust context, it is the most persuasive thing in it — a program honest about its edges is one an examiner believes about its center.
Assembling all seven tabs — running the scan, chasing the diligence, drafting the crosswalk, keeping the trail current — is real, recurring labor a 1–2 person program rarely has spare. That's exactly what a Discover scan is built to relieve: we run the read-only scan with you and hand back the inventory, the shadow-AI findings, and the dated diligence records already in binder shape — Tabs 2 through 4 populated, so you're reviewing and signing, not starting from a blank page.
Request your Discover scan — we run it with you. Start upstream with How to run an AI data call, or, if you're an RIA, The Reg S-P service-provider oversight file step by step.