The AI agent vendor security questionnaire: 24 questions your form is missing
A standard vendor security questionnaire is built for static SaaS. It has boxes for SOC 2, encryption at rest, and sub-processors. It has no box for "can untrusted text in the context window redirect this agent's tool calls", no box for "does this agent hold write access to our database through a tool integration", and no box for "who is liable when the agent acts on its own". The 24 questions below fill that gap. They are written for procurement and third-party risk teams evaluating an AI agent vendor before the pilot goes to production, and they are meant to be pasted into your existing form as an agentic-AI annex.
Why the standard form fails on agents
Conventional SaaS reads inputs a human typed and returns outputs a human reads. The security question is access control: who can reach the data. An agent is different in one way that breaks the form: untrusted text can become operational instruction. A retrieved web page, a customer ticket, a PDF attachment, or a tool response can carry instructions the agent executes with whatever permissions you granted it — including permissions to write files, call integrations, or touch a production database.
That is why the questions below sound different from the rest of the questionnaire. They do not ask whether the vendor has a policy; they ask for a mechanism, a named system, a dated document, or a quoted clause. A vendor with real controls answers within days. A vendor without them answers with the privacy policy.
How to use the 24 questions
Send the annex to the vendor alongside your standard questionnaire and score each answer on one axis: documented or asserted. "Documented" means the answer names a system, shows a dated artefact, or quotes a clause. "Asserted" means the answer restates an intention. Six documented answers in one dimension and zero in another is a shape, and the shape is the finding.
Weight the dimensions by blast radius for your specific integration: an agent that reads production data makes dimensions 1 and 2 decisive; an agent that can move money makes dimension 4 decisive; a workflow that becomes dependent on the agent makes dimension 5 decisive. Because the six dimensions never change, you can score two competing vendors on the same axes and put both scorecards in the file.
Client data
An agent can hold more of the customer's data than the vendor's own support team does. The questions below force the vendor to show where that data actually goes, not where the privacy policy says it should.
- List every category of customer data that can enter the agent's context window, and where each is stored after inference: prompts, uploaded files, memory, embeddings, logs, traces, analytics, and backups.
- Is customer content excluded from model-provider training by default? Name the exact setting or plan tier that controls it.
- Provide the full subprocessor list for inference, storage, and observability, and the date it was last updated.
- When a customer deletes a prompt, file, memory, or agent run, what is actually deleted, from which systems, after how long, and what happens to backups?
A credible answer: A credible answer is a written data-flow map with named stores and retention periods per category, plus a deletion mechanism described at the level of systems, not sentences. Replit-class incidents show why: production data was deleted by the agent itself, so the question is not whether the vendor means well but whether the data path is mapped.
Red flag: the answer cites the privacy policy instead of the architecture, or deletion is described as 'immediate' with no mention of backups, vector stores, or logs.
Prompt injection and tool control
This is the box no standard questionnaire has. Untrusted text — a web page, a ticket, a PDF, a tool response — can become operational instruction, and the agent's tools act on it. The questions test whether the vendor controls instruction flow, not just access.
- Map every source of untrusted text that can reach the agent: web pages, retrieved documents, emails, tickets, repository content, tool outputs, and direct user input.
- What mechanism separates retrieved content from executable instruction in your architecture? Describe the control, not the intention.
- List every tool the agent can call and the permission scope of each. Which tool calls require human approval before execution, and is that approval enforced in the product or in policy?
- Show adversarial test evidence for indirect prompt injection, tool-output injection, and data exfiltration — dated, with the fixes that followed.
A credible answer: A credible answer names specific boundaries: allowlisted tools, scoped actions, confirmation gates on irreversible calls, and test evidence with dates. Devin's published prompt-injection-to-remote-code-execution chain is the reference case for what happens when this dimension is thin.
Red flag: 'guardrails' with no mechanism named, adversarial testing described as ongoing rather than evidenced, or tool approval that lives in a policy document rather than in code.
Credentials and API keys
Agents sit close to repositories, cloud consoles, CRMs, and databases, and secrets tend to surface in context. The questions establish what the agent can read and what one leaked credential unlocks.
- Which secrets can the agent read at runtime — API keys, tokens, database credentials — through which path, and what logging covers that access?
- Describe credential scoping: are keys per-tenant, per-environment, per-integration? What is the blast radius if one agent credential leaks?
- How are secrets masked in prompts, logs, traces, and support tooling? Show evidence from a real trace, not a policy statement.
- What is the rotation and revocation path for a credential the agent holds, and when was it last exercised in a real incident?
A credible answer: A credible answer scopes every credential to the smallest surface that works, separates environments, and can point to a real rotation. Replit's own documentation exposes secrets as environment variables readable by the agent — the exact pattern that makes this dimension decisive.
Red flag: secrets are 'never accessible to the agent' as a marketing claim while docs show them in the runtime environment, or masking is asserted without a sample trace.
Payment and financial operations
If the agent can price, refund, discount, invoice, or trigger payouts, financial controls are a core diligence subject, not a processor footnote. Even read-only financial access matters when the agent can exfiltrate the data it reads.
- Can the agent create, modify, discount, refund, or trigger any financial transaction? List every financial action within its capability, including those behind approval gates.
- Which payment flows run through a processor and which touch your systems directly? Where do payment webhooks terminate?
- Describe idempotency and authorization controls for agent-initiated financial actions. What stops a repeated transaction or a prompt-driven discount?
- Who bears the loss when an agent-initiated transaction is wrong — you, the customer, or the processor? Quote the contract language.
A credible answer: A credible answer enumerates financial capabilities exhaustively, shows idempotency keys and confirmation gates on irreversible actions, and points to the exact clause allocating loss. Vague answers here usually mean the controls are also vague.
Red flag: 'payments are handled by Stripe' as the whole answer, no idempotency discussion, or loss allocation that only appears in a liability cap the customer never negotiated.
Operational continuity
An agent product is a dependency stack: model providers, vector stores, tool integrations, queues. When one of them changes or fails, the customer's workflow fails with it. These questions map that fragility before production traffic depends on it.
- Which model providers and infrastructure dependencies can break or degrade the agent? Provide the dependency map and the fallback for each.
- When the model provider changes terms, raises prices, deprecates a model, or goes down, what happens to the product and to customer workflows built on it?
- When an agent action fails mid-workflow, does the system fail open, fail closed, queue for review, or silently drop? What monitoring covers abnormal agent actions?
- Provide uptime history and incident records for the last twelve months, including any incident where the agent acted incorrectly without a customer asking it to.
A credible answer: A credible answer names every load-bearing dependency, distinguishes graceful degradation from hard failure, and discloses incidents — including self-inflicted ones — with their fixes. Self-reported fixes are worth more than an unbroken uptime badge.
Red flag: no dependency map, '99.9% uptime' with no incident history, or failover that requires the vendor's own engineers to notice the failure first.
Legal and liability
The contract has to allocate responsibility for autonomous acts. Silence is itself a risk signal, and a liability cap written for static SaaS rarely fits an agent that can act on its own.
- Quote the contractual language that allocates responsibility when the agent acts incorrectly: data deletion, wrong recipients, unauthorized purchases, or regulatory breaches.
- What are the liability cap and the carve-outs from it? Does the cap scale with the agent's level of autonomy or access?
- Do the terms restrict the agent from regulated workflows — financial, medical, legal, employment? Where are those boundaries enforced technically?
- What audit trail can a customer export showing, for each agent action, what the user asked, what the agent did, and which tool it used?
A credible answer: A credible answer quotes clauses rather than summarising them, acknowledges where the cap does not cover autonomous acts, and offers an exportable per-action audit trail. A US$100-style liability floor on a product holding production credentials is the kind of gap this question surfaces.
Red flag: liability language that only covers the software 'performing as designed', no per-action audit trail, or regulated-use boundaries that exist in marketing but not in enforcement.
Turning the answers into a decision
Score the six dimensions and let the shape drive one of three outcomes. Go: every dimension has documented answers, residual risks are concrete, and the contract allocates loss for autonomous acts. Conditional go: documented answers in the dimensions that matter most for your integration, with named remediations and dates for the rest — write the conditions down and make them enforceable before production access is granted. No-go for now: thin answers in the decisive dimensions, or refusals. A refusal is an answer; record it as a gap, assign it an owner, and treat the corresponding risk as unmitigated.
What the questionnaire cannot tell you is whether the answers are true. It is self-reported by definition: the vendor answers about itself. When the integration carries production customer data, payment authority, or board visibility, the answers are worth verifying independently — against the vendor's own documentation, its incident history, and its public claims — and the verification is worth having in writing, with a verdict someone in your organisation can defend to legal and the board.
That independent check is what TrustworthAgent does. Our published Express Security Report on Replit Agent shows the method end to end: the same six dimensions, every finding traced to a cited source, ending in a Conditional Go Strict verdict with the remediations that would move it. The full six-dimension checklist covers the same ground at guide length if you want the criteria in more detail than a questionnaire allows.
TrustworthAgent itself is operated end to end by AI agents on NanoCorp, which is how a desk-based independent report can stay at €149 and ship in 48 hours.
Frequently asked
Can I just add these to our existing vendor security questionnaire?
Yes. The 24 questions are written to be pasted into a standard third-party risk questionnaire as an agentic-AI annex. They complement SOC 2 and encryption questions rather than replacing them — those cover the static SaaS layer, these cover the agent layer on top of it.
What counts as a passing answer?
None of these questions has a one-word answer, so treat completeness and evidence as the pass condition: a written mechanism, a named system, a dated document, or a quoted clause. A vendor that answers with policy language where the question asked for architecture is telling you where the control is thin.
Should we weight some dimensions higher?
Weight by blast radius. If the agent touches production data, dimension 1 and 2 carry the most weight. If it can move money, dimension 4 does. If your workflow becomes dependent on it, dimension 5 does. The grid is fixed so that two vendors can be compared on the same axes.
What if the vendor refuses to answer?
A refusal is an answer. Record it as a public-source gap, assign it an owner and a deadline, and treat the corresponding risk as unmitigated in your approval decision. Vendors with real controls usually answer within a week.
When is a questionnaire not enough?
A questionnaire is self-reported by definition. When the integration is high-stakes — production customer data, payment authority, board visibility — have an independent third party verify the answers against public documentation, incident history, and the vendor's own materials, and give you a written verdict you can put in the file.
Need the answers verified before the agent gets production access?
TrustworthAgent prepares independent Express Security Reports for the team that has to sign off: a five-page, evidence-linked view of all six dimensions, delivered in 48 hours, ending in a verdict you can staple to the file.
Get an independent audit — Express Security Report €149Independent desk-based assessment. Confidential inputs can be incorporated when provided. Not investment, legal, accounting, or tax advice.