An independent audit on your own agent, or on a vendor you are about to trust.

Express Security Report — €149 · 48 hours
Free Public Sample — Methodology Demonstration — Based on Publicly Available Information Only
TrustworthAgent · Express Security Report · ESR-2026-004

Devin AI Agent Security Audit — Due Diligence Report

Cognition AI, Inc. · Autonomous Software Engineering Agent · Devin Cloud, Devin Desktop, Devin CLI

Report Type
Express Security Report
Publication
2026-09-28
Observation cutoff
2026-09-28, 08:30 UTC
Classification
Public Sample
Methodology
TrustworthAgent v1.0 (Desk-Based)
Subject Entity
Cognition AI, Inc.
Subject Product
Devin (Cloud, Desktop, CLI)
Access Granted
None — public sources only
Buyer Intent · On-Demand Audit

Need an AI agent security audit before a deal or deployment?

TrustworthAgent produces this type of independent audit on demand for investors, acquirers, and enterprise risk teams evaluating an autonomous AI company. The Express Security Report maps the agent security threat model across six dimensions: client data, prompt injection and tool control, credentials and API keys, payment and financial operations, operational continuity, and legal and liability risk.

Order Express Security Report — €149 · 48 hours
Recommendation
CONDITIONAL GO STRICT

Cognition is a substantial, well-capitalised company with real compliance artefacts and a permission model for its CLI that is among the most carefully documented we have assessed. The strict condition attaches to a specific gap: a working, publicly documented indirect prompt injection that achieved remote code execution and reached credentials on the cloud agent's own machine, reported to the vendor in April 2025, acknowledged, and published after 120+ days without an answer on fix timelines — with no public advisory or fix note found since. Set against a liability cap of six months of fees or US$100 and a set of security artefacts that cannot be read without signing an NDA, that is a file a risk lead can approve only with the five remediations in §08 answered in writing first.

§ 01

Legal & Operational Identity

Legal Entity
Cognition AI, Inc.
Address of record
550 Third Street, San Francisco, CA 94107
Founded
November 2023
Founders
Scott Wu, Steven Hao, Walden Yan
Headcount
~200 (2026, secondary source)
Capital raised
$400M+ at $10.2B post-money, led by Founders Fund
Reported valuation
$10B (Sep 2025) → $26B (May 2026)
Key acquisition
Windsurf, July 2025 — terms undisclosed

Cognition AI, Inc. is a US corporation with an address of record in San Francisco, given as the notice address for arbitration in its platform terms. [S9][S11] It was founded in November 2023 by Scott Wu, Steven Hao and Walden Yan, launched Devin in March 2024, and reported approximately 200 employees in 2026. [S2]

Funding is disclosed by the company: over $400M raised at a $10.2B post-money valuation led by Founders Fund, with Lux Capital and 8VC as joint leads of the prior round, alongside Neo, Elad Gil, Definition Capital, Swish VC, Bain Capital Ventures, Hanabi Capital and D1 Capital. [S1] Secondary sources report a valuation trajectory of roughly $10B in September 2025 and $26B in May 2026, with Bloomberg reporting in August 2026 early discussions for new financing at $40B or more — discussions only, and not confirmed by the company. [S2]

In July 2025 Cognition signed a definitive agreement to acquire Windsurf, days after Google's $2.4B acquihire of Windsurf's CEO and senior staff. Deal terms were not publicly disclosed. Windsurf was rebranded Devin Desktop in June 2026. [S2]

Materiality for a risk file. Counterparty solvency is not the exposure here. A company at this capitalisation is not going to disappear mid-contract, which removes the vendor-viability question that usually dominates an early-stage AI vendor review and leaves the technical and contractual questions to carry the whole decision. One point does deserve a note: the platform has absorbed an acquired IDE codebase and rebranded it inside twelve months, and the terms of service carry a linked prior version from April 2025, so the surface being assessed is a moving one. [S2][S11]

§ 02

Technical Architecture

Devin is an autonomous software engineering agent. The published product surface is Devin Cloud (cloud agents), Devin Desktop (the former Windsurf IDE), Tab (inline completion), Devin CLI, plus DeepWiki, the Devin API and Ask Devin. [S8][S13]

The cloud agent runs in a Cognition-hosted virtual machine that the vendor and researchers both call a DevBox. Published research confirms the agent has, inside that machine, a shell, a terminal it can open more than once, and a browser it uses to navigate to arbitrary URLs. [S7] This matters more than any single policy statement in this report: the unit of blast radius is a machine with a shell and network egress, not a text completion.

2.1 — Integration and permission surface

The architectural finding of this report is a scope mismatch, and it runs through every dimension below. Cognition publishes an unusually thorough permission model — five modes, a per-tool auto/prompt matrix, deny/ask/allow precedence, path and command and URL scopes, and organisation-level rules that project and user config cannot override. That documentation describes the CLI. [S5] The publicly demonstrated compromise was executed against the cloud DevBox and reproduced from Slack. [S7] We found no equivalent per-tool permission matrix for the cloud agent on any public page. A buyer who reads the CLI permissions documentation and concludes the cloud product carries the same enforceable gates is making an inference the public documentation does not support.

§ 03

Security Assessment — Six Dimensions

Framework

Assessed across TrustworthAgent's six security dimensions — D1 client data, D2 prompt injection and tool control, D3 credentials and API keys, D4 payment and financial operations, D5 operational continuity, D6 legal and liability — using the OWASP Top 10 for LLM Applications as the primary vulnerability reference framework. The grid is fixed and identical on every report.

Security Dimension 01 of 06

3.1 — D1 · Client Data

Threat model: an engineering organisation connects Devin to its repositories. Proprietary source code, infrastructure configuration, authentication logic and ticket content are processed by Cognition's backend and by whichever model provider serves the request.

Cognition publishes its privacy policy and platform terms ungated, covering the website, Devin and Windsurf under one entity. Infrastructure is AWS, with production access granted on a need-to-know basis to a limited number of employees or contractors, MFA required on primary work applications, and encryption in transit and at rest. Annual security training is asserted. [S3][S9] These are real institutional investments, and they are not where the finding is.

Finding — Two vendor sources disagree on training

The enterprise documentation states: "By default, Cognition does nottrain its models on customer data or code." The privacy policy's processing table lists a purpose of "using User Content to train, fine tune and improve the models that power our Services", with User Content among the data categories and Legitimate Interestsas the lawful basis, qualified by "depending on the terms that apply to your use of the Services". [S3][S9]

The two statements are reconcilable — a paid or enterprise agreement can carve out what a general privacy policy permits, and the "depending on the terms" qualifier is doing exactly that work. The finding is not that the vendor is being deceptive. The finding is that the published documents do not say which tiers fall on which side of the line, and a procurement lead cannot close that question by reading. It has to be answered in the contract, naming the tier. This is the single most common way an AI vendor review produces a false assurance: the documentation page says one thing, the binding instrument says another, and the questionnaire captures the documentation page.

Retentionis stated two ways. The documentation says data processed through Devin is retained "only for the duration of the customer relationship unless specified otherwise", while "Feedback & User Interaction Data may be retained as needed, as determined by Cognition". [S3] The privacy policy sets no fixed period, framing retention by necessity and legal obligation. [S9] There is no published number of days for anything.

The licence over customer data is broad.§3.2 of the platform terms grants Cognition a non-exclusive, worldwide, royalty-free, fully paid, sublicensable and transferable licence to "reproduce, distribute, modify, and otherwise use, display, and perform all acts with respect to the Customer Data as may be necessary for Cognition to provide the Services". The necessity-to-provide qualifier is the real boundary and it is a meaningful one, but the verbs and the transferability are worth a legal read before signature. [S11]

Sub-processors are not named publicly.The privacy policy discloses categories — "Third Party Service Providers", "Other Third Parties" — and we found no named sub-processor list on any public page. Model providers appear only generically on the pricing page. Transfers are worldwide including EU, UK and US, relying on Standard Contractual Clauses and the UK IDTA where applicable. For an EU deployment under GDPR Article 28, a named sub-processor list is not a preference, and it is not obtainable without asking. [S8][S9] The policy also flags shared Devin conversation links as a customer-side disclosure path, which is a fair and useful disclosure. [S9]

OWASP relevance: LLM06 (Sensitive Information Disclosure). A repository-scale context window transmitted to a third-party inference pipeline whose sub-processors are undisclosed, under a retention commitment stated without a period, is a direct LLM06 exposure pattern.

Risk: MEDIUM-HIGHFor any deployment without a signed tier-specific training and sub-processor schedule. MEDIUM for a dedicated or on-prem enterprise tenant.
Security Dimension 02 of 06

3.2 — D2 · Prompt Injection & Tool Control

Threat model: Devin is asked to look into a GitHub issue. The issue was filed by someone outside the organisation — a user, a contractor, a bug reporter — and its text contains instructions aimed at the agent rather than at a human reader.

Primary Finding — Highest Severity in this Report

This is not a theoretical attack surface or a proof-of-concept in a lab. A full exploit chain against Devin is published, with screenshots, and it ends in remote code execution and credential access on the agent's own machine. [S7]

Security researcher Johann Rehberger, publishing as part of the "Month of AI Bugs" series, planted instructions in a GitHub issue that pointed Devin at an attacker-controlled website. Devin navigated off-domain, followed the instructions it found there, and downloaded a Sliver command-and-control binary. The binary hit a permission-denied error — and Devin opened a second terminal, added the execute bit, and ran it again. The researcher got a remote shell on the DevBox, and from that shell reached AWS key material present on the machine. [S7]

Three details make this materially worse than a single successful injection. First, the researcher reproduced the attack from Slack, which he characterises as "an entirely unsupervised interaction" — no human was watching a screen when the agent was compromised. Second, persistence was maintained with a session-connected reaction that fired commands the instant the callback landed, so secrets could be exfiltrated in milliseconds even when Devin later cancelled the command. An operator watching carefully and hitting stop does not prevent the exfiltration. Third, the researcher's general observation across agents was that they "like clicking links", and that once an agent is off-domain on an attacker's page, payloads "often work at first try". [S7]

The disclosure history is itself a finding. The vulnerability was reported to Cognition on 6 April 2025 and acknowledged within days. Follow-up queries on fix timelines, status and coordinated disclosure went unanswered for more than 120 days, after which the researcher published. [S7] We searched for a vendor advisory, a fix note or any public response and found none. For a risk team, an unanswered disclosure is a signal about process, not just about one bug: it speaks to how the next report will be handled once the agent is inside your perimeter.

What Cognition does publish, and it is substantial

It would be unfair and inaccurate to present this vendor as having no tool-control story. The CLI permission model is one of the better-documented we have assessed, and a risk team should read it directly. [S5]

The vendor also publishes its own limitations plainly — Devin "may still" hallucinate, introduce bugs and suggest insecure coding practices — and recommends code review before deployment, branch protections and standard engineering review. [S3] That candour is worth crediting.

Why the rating is still HIGH.The controls are documented for the CLI. The compromise was demonstrated on the cloud agent and from Slack, where we found no equivalent published matrix. The researcher's central recommendation to the vendor was to stop relying on model behaviour or in-chat confirmations and build an out-of-band validation step for sensitive operations, along with blocking outbound connections by default, warning users not to join Devin to an enterprise or VPN network, and noting the absence of endpoint protection on the hosts. [S7]Smart mode's model-judged approval is a mitigation of exactly the class the researcher warned against relying on — a model deciding whether another model's action is safe. It raises the bar; it is not an out-of-band gate.

OWASP relevance: LLM01 (Prompt Injection) and LLM02 (Insecure Output Handling), compounded by LLM08 (Excessive Agency). A write-capable agent with a shell, a browser and network egress, consuming untrusted third-party text as context, is the canonical LLM01-into-LLM08 escalation, and here it is demonstrated rather than inferred.

Risk: HIGHFor cloud or Slack-triggered sessions on repositories that ingest third-party content. MEDIUM for CLI use under organisation-level deny rules with Fetch restricted to named domains.
Security Dimension 03 of 06

3.3 — D3 · Credentials & API Keys

Threat model: Devin needs credentials to do useful work — a database URL to run migrations, a cloud key to deploy, a token to call an internal API. Those credentials are present on the machine the agent controls, at the moment untrusted text enters its context.

A first-party Secrets feature on the Settings page is the documented mechanism for API keys, passwords and cookies. [S3] The dedicated secrets documentation page returned 404 at the time of access, so the feature's internals are not publicly evidenced: we cannot state how secrets are encrypted at rest, whether they are scoped per-session or per-repository, whether values are masked from the agent's own context window, or whether rotation is supported. Those are four questions a risk team must ask directly.

Repository access is reasonably designed. The administrator installing the GitHub integration selects which repositories Devin can reach and can adjust permissions at any time from GitHub App Settings — a per-repository grant under the customer's own control, which is better than a personal access token. Slack processing is scoped to explicit invocation: an @Devin tag, a direct prompt, or information shared in an active thread. [S3]

Against that sits the demonstrated exposure.The published exploit reached AWS key material from an attacker-controlled shell on the DevBox, and the researcher's explicit advice to users is that "any secret or private code Devin has access to can be leaked at will to third-party systems by the AI, or an attacker via indirect prompt injection". [S7] The CLI permission model does deny reads and writes of dotenv files and key material in Smart mode, which is a direct and well-aimed mitigation [S5] — but the exploit was not on the CLI.

Identity controls are tier-gated. SAML/OIDC SSO is Enterprise-only, and no SCIM provisioning is named anywhere on the published plan comparison. [S8]For an organisation of any size this is a joiner-mover-leaver problem: without SCIM, deprovisioning a departing engineer's access to an agent that holds repository write access is a manual step someone has to remember.

OWASP relevance:LLM06 (Sensitive Information Disclosure) and LLM08 (Excessive Agency). Credentials co-resident with an injectable agent on a machine with network egress are exfiltratable by design unless egress is restricted or the secrets are held outside the agent's reach.

Risk: HIGHWhere production credentials are made available to a cloud session. MEDIUM-HIGH with sandboxed CLI use, restricted Fetch scopes and short-lived scoped credentials only.
Security Dimension 04 of 06

3.4 — D4 · Payment & Financial Operations

Threat model: the agent's financial exposure. Two distinct questions — can it move money, and can it spend more than budgeted.

On the first question the answer is clean. We found no published capabilityfor Devin to hold spend authority, move funds, or transact on a customer's behalf. There are no payment rails in scope, and a risk team can close that line. The exposure is procurement spend.

Free
$0 — light quota, limited models
Pro
$20 / month
Max
$200 / month
Teams
$80 / month + $40 / month per full dev seat, up to 200 users
Enterprise
By quote — VPC, SSO, admin controls, teamspace isolation
Overage
Extra usage purchased, consumed at API pricing

The finding is that cost per unit of work is not deterministic, and no published spend cap exists.Each paid plan carries an allowance refreshing daily and weekly; beyond it, extra usage is purchased and "consumed at API pricing". The vendor states plainly that cost per message varies with the model used, task size, complexity and reasoning required. [S8]A budget owner therefore cannot derive a worst-case monthly figure from the pricing page, and the overage meter is denominated in a third party's API pricing rather than the vendor's own rate card. Concurrent-session limits — up to 10 on Free, Pro and Max, unlimited on Team and Enterprise — bound parallelism but not spend. [S8]

Centralised billing and an admin dashboard with analytics arrive at the Team tier and above, which is where visibility becomes possible at all. [S8] One live commercial term is time-boxed and worth diarising: SWE-2 is free in Devin Desktop and the CLI only through 10 October 2026. [S8][S14] Any cost model built on current pricing should be re-run after that date.

A second-order risk worth naming. An injected agent that can run arbitrary shell commands on metered infrastructure is a cost-amplification path as well as a security one. Unlimited concurrent sessions on Team and Enterprise tiers removes the natural ceiling.

Risk: MEDIUMNo payment authority, so no funds-movement exposure. The rating reflects non-deterministic metered spend with no published cap and a time-boxed free-model term.
Security Dimension 05 of 06

3.5 — D5 · Operational Continuity

Threat model: Devin sits on a delivery path with a deadline. The question is what happens when it is unavailable, and what the vendor has committed to.

Cognition operates a public Atlassian status page with per-component 90-day uptime and subscription by email, SMS, Slack, Teams, webhook and RSS. Transparency here is above the market norm for AI vendors, and the numbers are specific enough to act on. [S12]

Cloud Agent
99.77% — 90 days
Cloud Agent (Enterprise)
99.9% — 90 days
Cloud Web Client
99.87% — 90 days
Cloud Web Client (Enterprise)
99.99% — 90 days
Enterprise
99.97% — 90 days
Desktop Agent / Tab / Integrations
100% — 90 days

99.77% on the non-enterprise Cloud Agent is roughly five hours of unavailability per month. Enterprise components are measurably better than their general counterparts across the board, which is a real and quantified reason to price the Enterprise tier rather than a sales claim. Two incidents inside the eleven days before our observation cutoff were publicly written up with timestamps and resolved: 24 September 2026, "Devin webapp down", investigating 19:59 UTC to resolved 22:03 UTC, about two hours; and 17 September 2026, "Multiple services experiencing degraded performance", 13:40 to 15:20 UTC, about an hour forty. [S12] Both were handled and communicated properly. Incidents happen; the honest write-up is the good sign.

Finding — The terms disclaim the availability the status page demonstrates

§2.5 of the platform terms states the Services are subject to modification and change at Cognition's sole discretion and that "There are no guarantees made with respect to the quality, stability, availability, or reliability of the Services." We found no SLA with an uptime commitment or service credits on any public page. [S11][S12]

The gap between observed operational behaviour and contractual commitment is the finding. Good uptime that is disclaimed in the binding instrument is not a commitment a risk lead can put in a file. The trust centre does declare a Recovery Time Objective of 48 hours, a Data Access Level of "Internal" and an Impact Level of "Substantial", and asserts a business continuity plan with disaster-recovery testing whose frequency sits behind the NDA. [S4] A 48-hour RTO is a number worth knowing before Devin goes on a critical path.

Upstream dependency is distributed but not disclosed in detail. The platform routes to OpenAI, Anthropic, Google and open-source models alongside in-house SWE-2, which means no single provider outage takes the product down and gives Cognition genuine substitution power — a real resilience advantage over a single-provider agent. [S8] But we found no published dependency map, fallback commitment, or statement of which provider serves which workload, so the customer cannot verify the resilience or know whose terms govern their data on a given request. Monitoring is asserted: continuous logging, error handling, real-time dashboards and alerting on unusual application states. [S3]

Risk: MEDIUMObserved availability is good and honestly reported. The rating reflects the absence of any contractual commitment, a 48-hour RTO, and an undisclosed provider dependency map.
Security Dimension 06 of 06

3.6 — D6 · Legal & Liability

Threat model: the agent causes harm — it leaks code, corrupts a repository, or acts on injected instructions. The question is who carries the loss.

Finding — The liability cap is the greater of six months of fees, or US$100

§10.2 caps each party's aggregate liability at the greater of amounts paid to Cognition in the six months preceding the claim, or one hundred US dollars. §10.1 excludes consequential, incidental and punitive damages, expressly including loss of data and "breach of data or system security". [S11]

Put that against a Teams deployment. At $80/month plus $40 per seat, a twenty-seat team pays roughly $5,300 over six months — that is the recoverable ceiling for any claim, against an agent holding repository write access and credentials. The excluded categories are precisely the ones that matter here: a prompt-injection incident that exfiltrates source code is a loss of data and a breach of system security, both named exclusions. This is not unusual drafting for SaaS at this price point; it is material because the product's blast radius is not usual for SaaS at this price point.

Data-breach responsibility is narrowed further.§4.1 provides that Cognition "will not be responsible for any breach in security except to the extent the breach is due to Cognition's gross negligence" — a markedly higher bar than ordinary negligence. The same section puts routine backup on the customer, disclaims liability for loss, alteration, destruction or corruption of Customer Data, makes the customer responsible for notifying its own employees and customers of a breach and for regulator filings, and has the customer indemnify Cognition against claims from authorised users or data protection authorities on those obligations. [S11]

Finding — The output disclaimer names security assessments specifically

§8.2 provides that where output includes "assessment, review, analysis, evaluation, or examination of code, configurations, security posture, vulnerabilities, defects", Cognition makes no representation that it identifies all relevant issues, and such output is "a non-exhaustive aid only and is not a substitute for your own review, testing, audit, or other independent verification". [S11]

Anyone planning to put Devin into a security review, a code-audit or a compliance workflow should read that clause before designing the control. The vendor has disclaimed, in the binding instrument, exactly the reliance such a workflow would place on it. Everything is "AS IS" with all implied warranties disclaimed, including any warranty that the Services will "be secure". [S11]

What is genuinely favourable in these terms

One clause deserves a specific warning to risk teams.§3.5 assigns all Feedback to Cognition — including "any error, problem or defect reports" — with all right, title and interest assigned, no attribution or compensation, and the Feedback deemed Cognition's Confidential Information. [S11] Read literally, a detailed security or defect report filed through ordinary support channels is assigned to the vendor and classified as their confidential information. Route security findings through security@cognition.ai under separately agreed disclosure terms rather than through support.

Dispute resolution removes the court and the class.Binding arbitration before the AAA's International Centre for Dispute Resolution under Expedited Commercial Rules, a single arbitrator, in English, in New York, New York, with a jury-trial waiver and a class-action waiver; each party bears its own legal and expert fees regardless of outcome. Governing law is New York. [S11] For a European buyer, that is a US forum, US law, and self-funded costs on any dispute. A DPA exists but sits behind the trust-centre NDA, not on a public URL. [S4] Customer-side export-control and sanctions representations are imposed, including non-military-intelligence end-use. [S11]

Risk: HIGHThe cap, the gross-negligence standard for breach, and the security-assessment output disclaimer together leave the customer carrying the material residual. Negotiable at Enterprise tier; non-negotiable below it.
§ 04

Performance & Traction

Cognition is one of the most heavily capitalised companies in the agentic category: over $400M raised at a $10.2B post-money valuation, with secondary reporting of $26B by May 2026 and Bloomberg describing early discussions at $40B or more in August 2026. [S1][S2] The Windsurf acquisition brought an established IDE user base, rebranded to Devin Desktop in June 2026. [S2]

We deliberately publish no customer-count, revenue or retention figure here, because we found none in public sources we would sign our name to. A logo strip appears on the pricing page; we do not treat unlabelled logos as evidence of a reference-able deployment. [S8]

Two traction signals are worth reading, and they point in opposite directions. In its own funding retrospective the company describes the Devin of March 2024 as "still a very junior engineer" — unusual candour, and a marker of how fast the product has had to move. [S1] Against that, the current in-house model SWE-2 is being given away free in Devin Desktop and the CLI through 10 October 2026. [S8][S14] A time-boxed free tier on the flagship model is a distribution play in a category competing hard on developer default; it is not a sign of pricing power, and a buyer should assume the economics after that date are not yet knowable.

Relevance to the verdict: traction is a vendor-viability question, and on this file it is answered. Cognition will be here for the duration of a contract. That removes the concern that usually drives caution on an AI vendor review, and it is precisely why the verdict rests entirely on the technical and contractual findings rather than on company risk.

§ 05

Operational Risks

§ 06

Reputational Risks

The compliance artefacts are real, and they are all behind an NDA. The trust centre lists SOC 2 Type 2, ISO/IEC 27001:2022 and CCPA, with a SOC 2 report, a SOC 2 Type 2 bridge letter, a pentest report, a network diagram, a DPA and a certificate of insurance available on request. The documentation dates SOC 2 Type II to September 2024, covering data security, privacy, processing integrity, confidentiality and availability. [S3][S4] This is a genuine institutional investment and it distinguishes Cognition from most of the category.

The reputational finding is what cannotbe checked without entering a commercial relationship. Access requires a request, an emailed invite and a signed NDA. Until then a buyer cannot read the audit period, the scope boundary, the auditor's identity, whether there were exceptions, the ISO certificate number and issuing body, or the pentest date and scope. [S4] The certifications are therefore asserted-and-gated rather than verified for the purposes of this report, and we rate them as such.

Several trust-centre sections are assertion without artefact. App Security reads "we are putting together a program to monitor internal apps" — candid, and an acknowledged gap. Data Security, Access Control, Infrastructure, Endpoint Security, Incident Response, Continuous Monitoring and Data Privacy each assert industry best practice and offer "more details ... upon request". [S4]

Vulnerability disclosure exists as a channel but not as a demonstrated process. A named address, security@cognition.ai, is published, and Cognition states it will notify Enterprise customers of incidents affecting their environments per contractual obligations. [S3] Against that, the one substantive published finding went 120+ days without an answer on fix timelines. [S7] A bug bounty is not publicly confirmed: the trust-centre FAQ lists the question "Is there a bug bounty program in place?" and the answer is behind the NDA. [S4]

A capture of Devin's system prompt dated 10 April 2025 is publicly archived on GitHub, which is ordinary for this product class and not a finding in itself. [S7] We found no CVE naming Devin or Windsurf and no public litigation involving Cognition in the sources consulted.

Risk: MEDIUM-HIGHReal certifications, fully gated. The rating reflects unverifiable scope plus one unanswered disclosure, not an absence of security investment.
§ 07

Legal Risks

The legal exposure on this file is not drafting exotica; it is the ordinary SaaS risk allocation applied to a product whose failure mode is not ordinary SaaS. Four clauses carry it, and all four are quoted in §3.6: the liability capat the greater of six months' fees or US$100 with loss of data and breach of system security excluded from recoverable damages; the gross-negligence standard for data-breach responsibility, with backup and breach-notification duties pushed to the customer; the output disclaimer that names security assessments specifically; and the Feedback assignment that captures defect and error reports. [S11]

For a European buyer three further points are procurement-relevant. The DPA is behind the NDA rather than public, so Article 28 terms cannot be reviewed before engaging. [S4] No named sub-processor list is published, which is an Article 28(2) question that cannot be answered by reading. [S9] And dispute resolution is binding arbitration in New York under New York law, with a class-action waiver and each party bearing its own costs. [S11]

The favourable counterweights are real and should be recorded in the same file: customer ownership of output IP, a paid-tier IP indemnity, versioned terms with a 30-day change mechanism and a linked prior version, SCCs and the UK IDTA for transfers, and a published set of data-subject rights. [S3][S9][S11]

Risk: HIGHDriven by the cap and the exclusions set against the product's blast radius, not by unusual drafting.
§ 08

Recommendation

CONDITIONAL GO STRICT

Cognition AI, Inc. is a substantial and well-capitalised company with genuine compliance investment — SOC 2 Type 2, ISO/IEC 27001:2022, a public status page with honest incident write-ups, customer ownership of output IP, and a CLI permission model that is among the most carefully specified we have assessed. Vendor viability is not in question on this file.

The verdict is CONDITIONAL GO STRICT rather than CONDITIONAL GO because of the distance between three things that should not be far apart. A working indirect prompt injection achieving remote code execution and credential access on the cloud agent is published, was reported to the vendor in April 2025, was acknowledged, and was released after 120+ days without an answer on fix timelines — and we found no public advisory or fix note for it since. The controls that would bound that attack are documented for the CLI, not for the cloud or Slack paths where it was demonstrated. And the contractual recovery for the resulting harm is capped at the greater of six months of fees or US$100, with loss of data and breach of system security excluded from recoverable damages outright.

None of that makes Devin unusable. It makes the five questions below procurement-blocking rather than best-practice, each to be answered in writing before production access is granted:

1
Require an out-of-band approval gate for the cloud agent, in the contract

The permission model Cognition documents in depth — five modes, deny/ask/allow precedence, Read/Write/Exec/Fetch scopes, and organisation-level rules that project and user config cannot override — is documented for the CLI. The published exploit was executed against the cloud DevBox and reproduced from Slack. Before production access, require in writing: which of those controls apply to the cloud agent and to Slack-triggered sessions, who can change them, and whether an organisation-level deny is enforceable on a cloud session. If the answer is that cloud sessions do not carry an equivalent enforceable gate, restrict Devin to the CLI with organisation-level deny rules and treat cloud and Slack invocation as out of scope for approval.

2
Settle the training question against your own tier, in writing

The enterprise documentation states Cognition does not train on customer data or code by default. The privacy policy lists User Content as processed to train, fine tune and improve the models, under legitimate interests, qualified by 'depending on the terms that apply to your use of the Services'. Both can be true at once, and the published documents do not say which tiers sit on which side. Get a signed statement naming your tier, the retention period for code and session data, and the named sub-processor list including which model provider handles which workload. No public sub-processor list exists, so this cannot be resolved by reading.

3
Price the liability cap against the blast radius, and do it before signing

Aggregate liability is capped at the greater of six months of fees or US$100. Loss of data and 'breach of data or system security' are excluded from recoverable damages outright. Data-breach responsibility is limited to Cognition's gross negligence, backup duty sits with the customer, and the customer carries its own breach-notification obligations. For a Teams deployment at roughly a few thousand euros per six months, the recoverable ceiling is that figure against an agent holding repository write access and credentials. Either negotiate the cap and the gross-negligence standard, or scope Devin's access so the uncapped residual is one your organisation accepts.

4
Treat the NDA-gated trust centre as unverified until you have read the documents

SOC 2 Type 2, ISO/IEC 27001:2022 and a pentest report are all listed, and the documentation dates SOC 2 Type II to September 2024. None of it is readable without requesting access and signing an NDA: not the audit period, not the scope boundary, not the auditor, not the pentest date. Request all of them and check one thing specifically — whether the indirect prompt injection class demonstrated publicly in April 2025 was in the pentest scope, and whether it has been remediated. No public advisory or fix note for it was found.

5
Do not rely on availability language that the terms disclaim

Ninety-day uptime on the non-enterprise Cloud Agent reads 99.77%, roughly five hours of unavailability a month, against 99.9% on the Enterprise equivalent. Two incidents in the eleven days before the observation cutoff were publicly written up and resolved in about two hours and one hour forty. The status page and the 48-hour trust-centre recovery time objective are real operational signals, but §2.5 of the platform terms states there are no guarantees as to quality, stability, availability or reliability, and no SLA with service credits was found on any public page. If Devin sits on a delivery path with a deadline, the commitment has to be negotiated into the contract.

What would move this verdict to GO. A written confirmation that organisation-level deny rules are enforceable on cloud and Slack-triggered sessions with Fetch restricted to an allow-list; a signed tier-specific schedule on training, retention and named sub-processors; a negotiated liability position that does not exclude breach of system security; and a reviewed pentest report showing the April 2025 injection class in scope and remediated. Four of the five are answerable in a single procurement exchange.

What would move it to NO-GO. Granting a cloud or Slack-triggered Devin session standing access to production credentials, or to repositories that ingest third-party content such as public issue trackers, without an out-of-band approval gate and restricted egress. On the published evidence that configuration is the one demonstrated to fail, and the contract does not carry the loss.

§ A

Appendix — Sources, Methodology & Limitations

A.1 — Sources consulted

All sources are public and were accessed on 2026-09-28. Findings above are tagged with the source identifiers below.

A.2 — Methodology

TrustworthAgent Express Security Reports are desk-based assessments conducted using exclusively publicly available information. The six security dimensions assessed in §3.1–§3.6 constitute the TrustworthAgent Security Framework and are identical on every report we publish, so that two vendors can be compared on the same axes and a report written a year from now remains readable against this one. Each dimension is assessed against documented architectural choices, published policy commitments, known vulnerability classes from the OWASP LLM Top 10, and comparable market standards at equivalent company stage and product category.

Risk ratings are drawn from the fixed set LOW / MEDIUM / MEDIUM-HIGH / HIGH and reflect assessed materiality relative to the enterprise deployment scenario described in each threat model, not absolute severity in isolation. The recommendation is drawn from the fixed set GO / CONDITIONAL GO / CONDITIONAL GO STRICT / NO-GO.

A.3 — Limitations

No testing of any kind was performed against Cognition or Devin for this report.No penetration testing, no code review, no internal interviews, no non-public data access, and no purchase of the product. We did not reproduce the exploit described in §3.2; we report it as published third-party research by a named researcher, with its disclosure timeline as stated by that researcher, and we could not independently verify the vendor's response or the current remediation status.

This report reflects publicly available information as of the observation cutoff, 2026-09-28, 08:30 UTC. Cognition's security posture, documentation and policy commitments may have changed since. The following were material to our assessment and could not be verified from public sources: the SOC 2 audit period, scope and auditor; the ISO 27001 certificate number and issuing body; the pentest date, scope and whether the injection class in §3.2 was covered; a named sub-processor list; whether the §3.2 vulnerability has been remediated; any per-tool permission matrix for the cloud agent and Slack paths; the Secrets feature's encryption, scoping, masking and rotation behaviour; the existence of any SLA or bug bounty; SCIM availability; the Windsurf acquisition terms; and insurance limits. Each of these sits behind an NDA, a 404, or is simply unpublished. Where a fact was unavailable we have said so rather than inferred it.

Risk ratings for compliance artefacts reflect verifiability, not quality. Cognition may well hold excellent, broadly scoped certifications; we cannot read them, so we cannot rate them as verified, and we decline to credit what we have not seen.

A.4 — Non-affiliation and right of reply

This is a free, public, unsolicited report published by TrustworthAgent on 2026-09-28 (v1.0) as a demonstration of the Express Security Report methodology. TrustworthAgent has no commercial, financial, shareholding, contractual, partnership or mandate relationship with Cognition AI, Inc., its officers, its investors or its representatives, and holds no position in the company. We were not asked to write this report and were not paid to write it. We do not invest, we do not rate consumers, and we do not sell signals about the companies we audit.

Observations, risk ratings and recommendations are opinions of analysis founded on the methodology described and the sources cited. They are not exhaustive or definitive statements of fact, and they are not investment, legal, tax, accounting, cybersecurity or regulatory advice. Any reader relying on this for a procurement or investment decision should commission a full mandate, which is scoped in writing and can extend to proprietary architecture documentation and authorized testing where the subject grants access.

Cognition AI, Inc. has a right of reply. To request correction of a factual error or publication of a response, write to hello@trustworthagent.com citing this URL, the passages concerned, the supporting facts or sources, and the name and role of the person replying. We undertake to examine any good-faith request and, where it is founded, to correct the error, publish the response or add a dated update note within 5 business days of a complete request.

Intelligence Briefing

Get the next TrustworthAgent security due diligence report in your inbox. One report per publication. No noise.

Due Diligence Request

Need due diligence on a specific autonomous business?
Name the subject and the decision it supports. We reply in writing within one business day.

Or order directly: Express Security Report — €149 · 48 hours

Independent audit · TrustworthAgent

An independent audit on your own agent, or on a vendor.

Express Security Report — 5 pages, all six dimensions, delivered in 48 hours. For a fast go or no-go before an integration, a partnership or an investment.

Express Security Report — €149

Full audit mandate, by quote: request a written scope

Questions: hello@trustworthagent.com

← TrustworthAgent© 2026 TrustworthAgent · Free Public Sample · ESR-2026-004