AI Agent Due Diligence Checklist: 6 Security Dimensions
A practical framework for investors, acquirers, and enterprise risk teams asking how to audit an AI agent company security posture before a deal, integration, or material dependency is approved.
What this checklist is designed to decide
Traditional software diligence asks whether a vendor stores data safely and maintains ordinary SaaS controls. AI agent diligence must go further: it must determine what the agent can do, what it can see, what can influence its decisions, which credentials it can touch, and who is accountable when an automated action causes harm. The standard is not perfection. The standard is whether the buyer has enough evidence to approve, condition, delay, or reject the deal.
Need a buyer-ready security view on a specific AI agent company?
TrustworthAgent prepares independent Express Security Reports for investors, acquirers, and enterprise buyers. Reviews are evidence-based, confidential, and clear about public-source limits versus items requiring private diligence.
Customer Data Security
The first diligence question is not whether the model is capable. It is what customer material the agent can see, store, transform, send to vendors, or expose through logs and support workflows.
What to verify
- A data-flow map covering prompts, files, connectors, memory, embeddings, logs, traces, analytics, backups, support tools, and model-provider calls.
- Tenant isolation, data residency, retention periods, deletion mechanics, backup expiration, and whether customer content is excluded from training by default.
- Subprocessor list, DPA availability, support-access controls, audit logs for internal access, and customer-configurable controls for sensitive data.
Red flags
- No written data map, or only a generic privacy policy that does not describe agent memory, retrieval, tool output, or observability stores.
- Customer content can be retained indefinitely in prompts, debug traces, vector stores, or third-party model logs without an enterprise opt-out.
- The company cannot explain which employees, vendors, or automated systems can inspect customer prompts and outputs during support or debugging.
Questions to ask
- Which customer data categories enter the agent context window, and where are they stored after inference?
- Can a customer permanently delete prompts, uploaded files, memories, embeddings, and logs, including backups?
- Which model providers and subprocessors process customer data, under what retention and training terms?
Prompt Injection and Tool-Control Exposure
An AI agent converts language into action. Due diligence should treat every page, file, ticket, email, repository, database row, and tool response as a potential instruction source unless the company proves otherwise.
What to verify
- A boundary map showing all untrusted content sources, retrieval paths, context assembly logic, system prompts, tool schemas, and approval checkpoints.
- Adversarial testing evidence for direct prompt injection, indirect prompt injection, tool-output injection, instruction hierarchy bypass, and data-exfiltration attempts.
- Tool permission design: allowlists, scoped actions, confirmation gates, rate limits, human-in-the-loop requirements, and post-action audit trails.
Red flags
- Security claims rely on prompt wording alone, with no policy engine, parser separation, tool sandboxing, or repeatable adversarial test harness.
- The agent can browse untrusted content and call privileged tools in the same workflow without a separate authorization step.
- The company cannot produce examples of failed attack prompts, mitigations, regression tests, or incident-response handling for prompt-injection findings.
Questions to ask
- What prevents untrusted retrieved content from overriding the system instruction or tool-use policy?
- Which actions require human approval, and can a prompt-injection payload suppress or alter that approval request?
- How are prompt-injection regressions tested before new tools, connectors, models, or retrieval sources are released?
Credentials, Secrets, and API-Key Security
Agent companies frequently operate near repositories, cloud consoles, CRMs, payment processors, internal databases, and customer workspaces. A serious review assumes secrets will appear in context unless isolation controls are evidenced.
What to verify
- Secrets architecture: storage location, encryption, key-management service, rotation process, revocation path, environment separation, and least-privilege scopes.
- Controls that prevent credentials from being copied into prompts, model requests, trace logs, browser automation output, support tickets, or user-visible responses.
- Connector authorization model, OAuth scopes, service-account ownership, offboarding mechanics, and break-glass access approvals.
Red flags
- Shared API keys, long-lived tokens, broad OAuth scopes, manual credential handling, or no customer-visible revocation mechanism.
- Prompt, trace, or observability vendors can receive raw secrets or bearer tokens as part of normal agent operation.
- Production and test credentials are not separated, or the company cannot show when a compromised integration can be contained.
Questions to ask
- Where do customer API keys live, who can access them, and how quickly can they be rotated or revoked?
- Can the agent read secrets from a repository, ticket, document, browser session, or terminal and transmit them to a model provider?
- What happens if an integration token is compromised through an agent workflow?
Payments and Financial-Operation Security
If an agent can quote prices, approve refunds, generate invoices, change billing records, reconcile revenue, or trigger payouts, financial controls become a security diligence topic, not merely a commerce integration detail.
What to verify
- Payment architecture: hosted checkout versus direct card handling, webhook endpoint ownership, idempotency, refund permissions, chargeback handling, and ledger reconciliation.
- Authorization for financial actions, including role-based controls, manual approvals, dual-control requirements, and alerting for high-risk transactions.
- Agent-specific abuse cases such as prompt-driven discounting, invoice tampering, unauthorized refunds, customer-account takeover, and webhook replay.
Red flags
- The agent can modify prices, coupons, refunds, invoices, or subscription states without durable audit logs and human approval thresholds.
- Webhook verification, idempotency, and replay protection are undocumented or delegated to informal application logic.
- The company cannot distinguish payment metadata needed for operations from regulated cardholder data that should never enter the agent layer.
Questions to ask
- Which payment actions can the agent initiate, recommend, approve, or execute without a human operator?
- How are webhook events authenticated, replay-protected, reconciled, and investigated after a disputed transaction?
- Can a malicious prompt, support ticket, or customer message cause pricing, billing, refund, or payout changes?
Operational Continuity and Dependency Risk
Autonomous AI products often become hidden workflow dependencies. Buyers should assess whether the company can continue operating when a model, queue, vector store, browser, API, or cloud region fails.
What to verify
- Dependency map for model providers, orchestration services, vector databases, queues, browser automation, authentication, storage, observability, and payment systems.
- Uptime history, incident register, status-page process, recovery objectives, backup and restore evidence, customer export paths, and escalation ownership.
- Fail-safe behavior: whether failed actions are blocked, queued, retried, silently skipped, or executed later with stale context.
Red flags
- No incident register, no public or customer-facing status process, and no evidence of recovery testing for critical dependencies.
- The agent fails open during outages, keeps acting on stale context, or retries privileged actions without idempotency and human review.
- Customers cannot export workflows, memory, logs, policies, or configuration if the vendor terminates service or suffers a prolonged outage.
Questions to ask
- What is the operational plan when the primary model provider, vector database, or tool API is degraded or unavailable?
- Which workflows stop safely, and which continue with reduced capability?
- Can customers recover state, audit history, and configuration without the vendor being operational?
Legal, Contractual, and Liability Security
The legal file must match the product's operational reality. When an AI agent can act, decide, write, transact, or advise, buyers need contractual clarity about responsibility, remedies, and residual risk.
What to verify
- Terms, privacy policy, DPA, subprocessors, acceptable-use policy, security commitments, service levels, indemnities, liability caps, audit rights, and data-processing instructions.
- Alignment between marketing claims and contractual disclaimers for autonomy, accuracy, compliance, security, regulated use cases, and human oversight.
- Responsibility allocation for incorrect actions, data leakage, IP disputes, privacy requests, model-provider changes, and regulated-workflow failures.
Red flags
- Marketing invites regulated or mission-critical reliance while the contract disclaims responsibility for outputs, actions, availability, and data accuracy.
- No DPA, subprocessor list, security terms, status commitments, or named legal entity for the party handling customer data.
- The company cannot describe who bears loss if the agent deletes data, sends wrong instructions, makes unauthorized changes, or leaks confidential material.
Questions to ask
- Who is contractually responsible when the agent takes a wrong or unauthorized action?
- Which use cases are prohibited, unsupported, or require additional written terms before deployment?
- Do the public claims, sales materials, terms, privacy policy, and security commitments describe the same product risk?
How to use the checklist in a deal room
See the checklist applied to public AI-agent companies
These public reports show how TrustworthAgent applies a documented, independent, public-source methodology. They are not private audits, endorsements, investment recommendations, or claims of affiliation.
Applied example for prompt injection, credential exposure, MCP/plugin blast radius, customer data, and provider concentration.
Read applied report →Applied example for permission scopes, audit trails, intervention handlers, legal-surface gaps, and operational claims.
Read applied report →Applied example for metric provenance, API-key purchase flow, benchmark-methodology risk, and missing legal/status pages.
Read applied report →Additional report drafts are linked only after their public URLs return a successful indexable response; the guide avoids linking to unavailable pages.
Methodology boundaries
TrustworthAgent reviews are independent and are prepared for buyer-side decision support. A public-source report does not imply private access, cooperation by the target company, penetration testing, endorsement, or a complete statement of all facts.
Confidential materials supplied by a client should be handled under written scope, used only for the agreed diligence purpose, and separated from public examples. Findings should be framed as risk judgments based on evidence, limitations, and date-bounded observations; they are not legal, investment, or technical certification advice.
Need a buyer-ready security view on a specific AI agent company?
TrustworthAgent prepares independent Express Security Reports for investors, acquirers, and enterprise buyers. Reviews are evidence-based, confidential, and clear about public-source limits versus items requiring private diligence.