What a security review asks
The questions every enterprise AI platform gets asked, what a good answer looks like, and which ones are genuinely hard.
The questions are the same everywhere, and they arrive in roughly this order. Knowing them before the meeting is the difference between a two-week review and a two-month one. Answers that hold up cite a configuration or a contract clause; answers that do not cite a vendor blog post.
This is the companion to estimating a bill: together they are the two things you actually have to defend when picking a platform.
1. Where does the data go
The first question, always, and the one that decides whether the rest of the meeting happens.
| Ask | A good answer |
|---|---|
| Does the request leave our tenant? | names the boundary: same account, same region, same VPC, or a named third party |
| Which region does inference run in? | a specific region, pinned by configuration, not "usually" |
| Is there a cross-region fallback? | says whether routing can silently move the request, and how to disable it |
| Who is the model provider? | names them, because a first-party platform reselling a third-party model is still a disclosure |
The trap is a platform that runs "inside your cloud account" while the model itself is a managed endpoint the provider operates. Both can be true at once. Ask where the inference runs, not where the service is billed.
2. Is our data used for training
The answer you need is contractual, not architectural, and it should be in writing.
Distinguish three things that get conflated:
- Training the base model on your inputs. The answer must be no, and it should be a contract term.
- Retention for abuse monitoring. Common, often 30 days, sometimes waivable on enterprise terms. Ask the number and who can read it.
- Your own fine-tuning. Your data by definition, but ask where the resulting weights live and who else can invoke them.
3. What is logged, and for how long
Prompts and completions are the most sensitive artifacts in the system, because they contain whatever the user pasted in.
Cover each of these separately: the platform's own logs, your application's logs, the observability tooling in between, and any prompt-caching layer. A team that carefully disables provider logging and then writes full prompts to its own log aggregator has moved the problem, not solved it.
4. Who can call it
This is ordinary access control and it is usually the easiest section, provided the platform integrates with the identity system you already run.
| Ask | Why it matters |
|---|---|
| Does it use our existing identity provider? | a separate user directory is a separate offboarding problem |
| Can we scope access per model? | frontier models are a cost and a risk decision |
| Is there an audit trail of who invoked what? | required for almost every framework |
| Can a service account act on a user's behalf? | this is where row-level security quietly breaks |
That last row is the subtle one. If a retrieval layer runs as a service account with broad read access, every user of the application inherits that access through the answers, regardless of what their own permissions say.
5. How does retrieval respect existing permissions
The hardest question on the list, and the one most likely to be answered badly.
An index built from documents the indexing job could read will happily answer questions about documents the asking user cannot. There are three honest answers:
- Filter at query time using the user's identity, with the index carrying the access metadata. Correct, and the most work.
- Partition the index per permission boundary. Simple, and multiplies your index cost.
- Index only what everyone may see. Perfectly legitimate, and it constrains the product.
"The model will not reveal it" is not one of the three.
6. What happens when the model changes
Models are deprecated on the provider's schedule, not yours, and outputs shift between versions even when the interface does not.
Ask whether you can pin a version, what the deprecation notice period is, and what the migration path looks like. Then ask the harder internal question: if the model silently got better or worse, would you find out? A prompt tuned against one version is an undocumented dependency on it.
7. How do you know it is still working
Evaluation is a governance question, not just an engineering one, because "the AI was wrong" is an incident someone has to own.
The minimum defensible position is a held-out set of real cases with known-good answers, run on a schedule, with a threshold that alerts. Without it you cannot distinguish a model regression from a prompt change from a data change.
8. What is the failure mode
Ask what the system does when the model is unavailable, slow, or returns something malformed. The answer should be a defined behavior rather than an exception trace reaching a user, and for anything customer-facing it should degrade to something honest rather than guessing.
The artifacts worth having ready
A review goes faster when these exist before it starts:
- A one-page data-flow diagram showing where the request goes and what crosses a boundary
- The retention numbers, for the platform and for your own logs
- The access model, including what the retrieval layer runs as
- The evaluation set and its current results
- The list of models in use and their pinned versions
Gotchas
- "It runs in your account" and "the provider sees your prompts" can both be true. Ask about inference, not billing.
- Prompt caching is a retention surface. A cached prefix is your data sitting somewhere, for some duration, by definition.
- A pilot exemption becomes production. Whatever a security review conditionally approves for a proof of concept is what will still be running a year later.
- The retrieval layer's identity is the whole permission story. Most access-control failures in these systems are one over-privileged service account.
- Deprecation notices are not addressed to you. They arrive on a provider status page, and somebody has to be watching it.
See also
- Estimating an LLM bill for the other half of the decision
- AWS Bedrock AI Strategy for that platform's governance surface
- Snowflake Cortex: A Working Engineer's Guide for its access model
Related pages
- Estimating an LLM billThe token arithmetic every platform bills on — where the money actually goes, and which levers move it.
- AWS Bedrock AI StrategyConsolidated overview of Bedrock capabilities, use cases, architecture patterns, governance, and a sequenced project portfolio for a wealth management firm.
- Snowflake Cortex AI: A Working Engineer's GuideImplementation-level Cortex reference — product map, setup, AISQL functions, Cortex Search, the token cost model, monitoring queries, governance, and gotchas.
- AWS Bedrock Use Cases for RIAsA catalog of 112 concrete Bedrock applications for a wealth management firm, grouped by function — document intelligence, data engineering, CRM, compliance, analytics, backfills, and agents.
- Snowflake Cortex AI: The Big PictureA non-technical orientation to Cortex AI — what it is, why it matters strategically, where the real value is, what it costs, and what to be skeptical about.