How to vet an AI agent company before you invest, acquire, or integrate
Autonomous AI companies can look commercially mature while still carrying unresolved security, operational, and legal risks. This guide gives investors, acquirers, and enterprise buyers a practical diligence structure before capital, customer data, payment authority, or business workflows are placed in the product.
What makes AI-agent diligence different
Standard diligence asks whether software works, scales, and sells. Agent diligence must also ask whether the software can act. Once a product can interpret ambiguous instructions, call tools, write records, move money, or change files, a buyer is no longer reviewing a passive workflow application.
The review should therefore combine technical architecture, evidence quality, commercial traction, operational continuity, legal allocation of responsibility, and a security model tailored to autonomous behavior. A polished demo is useful evidence, but it is never enough.
Legal and Operational Identity
Start with the entity, not the demo. Buyers need to know who contracts, who controls the product, where the company operates, and whether public claims match legal reality.
- Confirm legal entity name, jurisdiction, operating address, trade names, founders, and beneficial ownership where available.
- Compare website, checkout, privacy policy, terms, security page, status page, and public repositories for inconsistent naming or missing ownership signals.
- Document the observation date for every fact. AI-agent startups can change positioning, product scope, and pricing quickly.
Decision test: If the counterparty cannot establish who owns the product and who accepts contractual responsibility, defer the investment, acquisition, or integration until identity is resolved.
Technical Architecture and Agent Boundary
A conventional SaaS architecture review is insufficient. The central question is what the agent can perceive, decide, remember, and execute without human review.
- Draw the agent boundary: inputs, context assembly, model calls, tool permissions, memory, logs, human approvals, and external actions.
- Separate deterministic software from probabilistic behavior. Record where the model can influence code paths, permissions, payments, communications, or records.
- Ask whether the product supports third-party tools, browser automation, file-system writes, repository access, email sending, or customer-database actions.
- Verify that architecture diagrams are current, specific, and reconciled with public docs, product behavior, and security statements.
Decision test: If the company cannot describe the agent boundary precisely, any valuation, integration plan, or acquisition thesis should carry an explicit technical uncertainty discount.
Security Analysis: Six Dimensions Investors Should Not Skip
Security is the diligence center of gravity for autonomous AI companies. The review should allocate substantial time to six dimensions that determine whether the agent is safe enough to trust with customer systems.
Client data security
Identify what the agent can read, retain, transform, transmit, or use for model improvement. The review should separate user prompts, uploaded files, memory stores, application logs, embeddings, tool outputs, and administrative metadata.
- Ask for a data-flow map from input to model call, storage, observability, support access, and deletion.
- Confirm whether customer data is excluded from training by default or only behind an enterprise setting.
- Check retention periods, tenant isolation, deletion mechanics, backup treatment, and subprocessors.
Prompt injection and tool-control exposure
Agent products fail differently from conventional SaaS because untrusted text can become operational instruction. Treat every external page, document, email, ticket, repository, database row, and tool response as a possible instruction source.
- Map all places where retrieved or user-supplied content enters the agent context window.
- Inspect guardrails that distinguish data from instructions before tool calls are executed.
- Require evidence of adversarial testing against indirect prompt injection, data exfiltration, and unsafe tool chaining.
Credentials and API-key security
AI agents often sit close to repositories, cloud consoles, CRMs, ticketing systems, payment tools, and internal databases. A diligence review should assume secrets will appear in context unless controls prove otherwise.
- Review secret-scanning, masking, least-privilege scopes, key rotation, and revocation paths.
- Check whether the agent can copy secrets into prompts, logs, traces, support tickets, or external tools.
- Require separation between development, staging, production, and customer-specific credentials.
Payment and financial-operation security
If the agent can price, quote, refund, subscribe, invoice, reconcile, or trigger payouts, financial controls become a core diligence subject rather than a payment-processor footnote.
- Confirm which payment data is handled directly, which is delegated to a processor, and where webhooks terminate.
- Test authorization, idempotency, refund controls, audit trails, and manual override paths.
- Inspect controls that prevent prompt-driven discounting, invoice tampering, or unauthorized financial actions.
Operational continuity
An autonomous product can become a hidden dependency in a buyer's workflow. Reliability diligence must cover model providers, vector databases, queues, monitoring, fallbacks, incident response, and graceful degradation.
- Request uptime history, incident records, dependency maps, recovery objectives, and status-page evidence.
- Identify whether failures are fail-open, fail-closed, queued for review, or silently dropped.
- Check whether customers can export state, logs, memories, workflows, and configuration before termination.
Legal and liability security
The contract should allocate responsibility for autonomous acts, training-data disputes, regulated outputs, privacy requests, and third-party model failures. Silence is itself a risk signal.
- Review terms, DPA, subprocessors, acceptable-use policy, indemnities, liability caps, and IP provisions.
- Compare marketing promises against contractual disclaimers and documented product limits.
- Confirm who bears loss when the agent deletes data, sends wrong instructions, or performs an unauthorized action.
Decision test: If any high-impact dimension is undocumented, untested, or delegated to vague policy language, treat the gap as a deal risk rather than a post-close implementation detail.
Performance, Revenue, and Traction Evidence
AI-agent companies often demonstrate capability before durable adoption. Diligence should separate product excitement from recurring usage, paid demand, and observable proof.
- Request cohort-level usage, retention, active accounts, expansion, churn, support volume, and paid conversion by customer segment.
- Reconcile public metrics with billing data, product telemetry, customer references, and contract terms.
- Treat screenshots, demo counters, benchmark claims, waitlists, and unverified testimonials as directional rather than dispositive evidence.
Decision test: If traction depends on a narrow demo workflow or unsourced performance claims, write the investment case around verified usage only.
Operational Risk Assessment
Autonomous systems create operating risks that appear after deployment: silent failures, cascading tool calls, brittle dependencies, excessive support burden, and unclear incident ownership.
- Review monitoring coverage for model errors, tool failures, abnormal actions, cost spikes, latency, and customer-visible incidents.
- Inspect support procedures for urgent customer interruptions, mistaken agent actions, and rollback requests.
- Confirm there is a documented escalation path from customer report to engineering response to post-incident communication.
Decision test: If operational response depends on founders manually watching logs, price the company as an early operational system, not as mature infrastructure.
Reputational and Market Signal Review
Reputation checks should be sober. The goal is not to reward visibility, but to identify claims that create reliance, customer expectations, or downside if contradicted later.
- Catalogue public promises about autonomy, accuracy, security, compliance, uptime, cost savings, and human oversight.
- Compare launch posts, docs, pricing pages, changelogs, and customer quotes for unsupported escalation of claims.
- Record relevant press, community discussion, public incidents, takedown requests, and unresolved criticism without overstating weak signals.
Decision test: If marketing language invites regulated or mission-critical reliance before controls are evidenced, the diligence memo should flag reliance risk.
Legal, Compliance, and Liability Review
Legal review should be tied to the product's actual autonomy. A low-risk content assistant and a high-permission operational agent should not receive the same legal treatment.
- Check whether the company publishes terms, privacy policy, DPA, subprocessors, security commitments, export-control language, and acceptable-use boundaries.
- Identify regulated workflows: financial advice, medical decisions, employment actions, legal drafting, consumer credit, insurance, safety-critical operations, or personal-data processing.
- Require contractual clarity on model providers, customer-data rights, generated-output ownership, audit logs, indemnities, and post-incident remedies.
Decision test: If the contract contradicts the product's operational role, escalate to counsel before close or enterprise rollout.
Decision Record and Go / No-Go Recommendation
The final output should be a defensible decision record, not a narrative summary. It should preserve evidence, uncertainty, residual risk, and the specific conditions required before funds or systems are exposed.
- Assign each material risk an owner, severity, evidence basis, mitigation, and date by which it must be resolved.
- Separate confirmed facts from assumptions, founder assertions, public-source gaps, and items requiring private diligence.
- Use clear outcomes: Go, Conditional Go, Delay, or No-Go. Conditions should be measurable enough for a board, investment committee, or procurement team to enforce.
Decision test: If the decision cannot be defended from evidence, do not convert the guide into confidence. Mark the limits and require additional diligence.
Evidence hierarchy
Confidentiality handling
Method limits
Need a decision record before capital, acquisition, or integration?
TrustworthAgent prepares independent Express Security Reports for buyer teams that need a concise, evidence-linked risk view before exposing customer data, credentials, payment flows, or operating workflows to an autonomous AI company.
Get an independent audit — Express Security Report €149Independent desk-based assessment. Confidential inputs can be incorporated when provided. Not investment, legal, accounting, or tax advice.
Need due diligence on a specific autonomous business?
Besoin d'une DD sur une entreprise autonome précise ?
Or order directly: Express Security Report — €149