Legal & Operational Identity
Cognition AI, Inc. is a US corporation with an address of record in San Francisco, given as the notice address for arbitration in its platform terms. [S9][S11] It was founded in November 2023 by Scott Wu, Steven Hao and Walden Yan, launched Devin in March 2024, and reported approximately 200 employees in 2026. [S2]
Funding is disclosed by the company: over $400M raised at a $10.2B post-money valuation led by Founders Fund, with Lux Capital and 8VC as joint leads of the prior round, alongside Neo, Elad Gil, Definition Capital, Swish VC, Bain Capital Ventures, Hanabi Capital and D1 Capital. [S1] Secondary sources report a valuation trajectory of roughly $10B in September 2025 and $26B in May 2026, with Bloomberg reporting in August 2026 early discussions for new financing at $40B or more — discussions only, and not confirmed by the company. [S2]
In July 2025 Cognition signed a definitive agreement to acquire Windsurf, days after Google's $2.4B acquihire of Windsurf's CEO and senior staff. Deal terms were not publicly disclosed. Windsurf was rebranded Devin Desktop in June 2026. [S2]
Materiality for a risk file. Counterparty solvency is not the exposure here. A company at this capitalisation is not going to disappear mid-contract, which removes the vendor-viability question that usually dominates an early-stage AI vendor review and leaves the technical and contractual questions to carry the whole decision. One point does deserve a note: the platform has absorbed an acquired IDE codebase and rebranded it inside twelve months, and the terms of service carry a linked prior version from April 2025, so the surface being assessed is a moving one. [S2][S11]
Technical Architecture
Devin is an autonomous software engineering agent. The published product surface is Devin Cloud (cloud agents), Devin Desktop (the former Windsurf IDE), Tab (inline completion), Devin CLI, plus DeepWiki, the Devin API and Ask Devin. [S8][S13]
The cloud agent runs in a Cognition-hosted virtual machine that the vendor and researchers both call a DevBox. Published research confirms the agent has, inside that machine, a shell, a terminal it can open more than once, and a browser it uses to navigate to arbitrary URLs. [S7] This matters more than any single policy statement in this report: the unit of blast radius is a machine with a shell and network egress, not a text completion.
2.1 — Integration and permission surface
- Source control: GitHub, GitLab and Bitbucket, plus custom git provider support. GitHub Enterprise Server and GitHub Enterprise Cloud with Data Residency are both supported for enterprise connections. [S6][S8]
- Messaging and tickets: Slack and Microsoft Teams; Linear and Jira. [S8] The Slack path is material: it lets a session be started and run without anyone watching a screen.
- MCP: supported, and permission-governed at three granularities — one tool (
mcp__server__tool), a whole server (mcp__server__*), or every MCP tool (mcp__*). [S5] - Models: in-house SWE-2 alongside third-party frontier models from OpenAI, Anthropic and Google plus open-source models. [S8][S14]
- Enterprise-only controls:VPC deployment, SAML/OIDC SSO, centralised admin controls and teamspace isolation are listed on the Enterprise tier only. For dedicated or on-prem deployments the vendor states all customer data stays within the customer's tenant. [S3][S8]
The architectural finding of this report is a scope mismatch, and it runs through every dimension below. Cognition publishes an unusually thorough permission model — five modes, a per-tool auto/prompt matrix, deny/ask/allow precedence, path and command and URL scopes, and organisation-level rules that project and user config cannot override. That documentation describes the CLI. [S5] The publicly demonstrated compromise was executed against the cloud DevBox and reproduced from Slack. [S7] We found no equivalent per-tool permission matrix for the cloud agent on any public page. A buyer who reads the CLI permissions documentation and concludes the cloud product carries the same enforceable gates is making an inference the public documentation does not support.
Security Assessment — Six Dimensions
Assessed across TrustworthAgent's six security dimensions — D1 client data, D2 prompt injection and tool control, D3 credentials and API keys, D4 payment and financial operations, D5 operational continuity, D6 legal and liability — using the OWASP Top 10 for LLM Applications as the primary vulnerability reference framework. The grid is fixed and identical on every report.
3.1 — D1 · Client Data
Threat model: an engineering organisation connects Devin to its repositories. Proprietary source code, infrastructure configuration, authentication logic and ticket content are processed by Cognition's backend and by whichever model provider serves the request.
Cognition publishes its privacy policy and platform terms ungated, covering the website, Devin and Windsurf under one entity. Infrastructure is AWS, with production access granted on a need-to-know basis to a limited number of employees or contractors, MFA required on primary work applications, and encryption in transit and at rest. Annual security training is asserted. [S3][S9] These are real institutional investments, and they are not where the finding is.
The enterprise documentation states: "By default, Cognition does nottrain its models on customer data or code." The privacy policy's processing table lists a purpose of "using User Content to train, fine tune and improve the models that power our Services", with User Content among the data categories and Legitimate Interestsas the lawful basis, qualified by "depending on the terms that apply to your use of the Services". [S3][S9]
The two statements are reconcilable — a paid or enterprise agreement can carve out what a general privacy policy permits, and the "depending on the terms" qualifier is doing exactly that work. The finding is not that the vendor is being deceptive. The finding is that the published documents do not say which tiers fall on which side of the line, and a procurement lead cannot close that question by reading. It has to be answered in the contract, naming the tier. This is the single most common way an AI vendor review produces a false assurance: the documentation page says one thing, the binding instrument says another, and the questionnaire captures the documentation page.
Retentionis stated two ways. The documentation says data processed through Devin is retained "only for the duration of the customer relationship unless specified otherwise", while "Feedback & User Interaction Data may be retained as needed, as determined by Cognition". [S3] The privacy policy sets no fixed period, framing retention by necessity and legal obligation. [S9] There is no published number of days for anything.
The licence over customer data is broad.§3.2 of the platform terms grants Cognition a non-exclusive, worldwide, royalty-free, fully paid, sublicensable and transferable licence to "reproduce, distribute, modify, and otherwise use, display, and perform all acts with respect to the Customer Data as may be necessary for Cognition to provide the Services". The necessity-to-provide qualifier is the real boundary and it is a meaningful one, but the verbs and the transferability are worth a legal read before signature. [S11]
Sub-processors are not named publicly.The privacy policy discloses categories — "Third Party Service Providers", "Other Third Parties" — and we found no named sub-processor list on any public page. Model providers appear only generically on the pricing page. Transfers are worldwide including EU, UK and US, relying on Standard Contractual Clauses and the UK IDTA where applicable. For an EU deployment under GDPR Article 28, a named sub-processor list is not a preference, and it is not obtainable without asking. [S8][S9] The policy also flags shared Devin conversation links as a customer-side disclosure path, which is a fair and useful disclosure. [S9]
OWASP relevance: LLM06 (Sensitive Information Disclosure). A repository-scale context window transmitted to a third-party inference pipeline whose sub-processors are undisclosed, under a retention commitment stated without a period, is a direct LLM06 exposure pattern.
3.2 — D2 · Prompt Injection & Tool Control
Threat model: Devin is asked to look into a GitHub issue. The issue was filed by someone outside the organisation — a user, a contractor, a bug reporter — and its text contains instructions aimed at the agent rather than at a human reader.
This is not a theoretical attack surface or a proof-of-concept in a lab. A full exploit chain against Devin is published, with screenshots, and it ends in remote code execution and credential access on the agent's own machine. [S7]
Security researcher Johann Rehberger, publishing as part of the "Month of AI Bugs" series, planted instructions in a GitHub issue that pointed Devin at an attacker-controlled website. Devin navigated off-domain, followed the instructions it found there, and downloaded a Sliver command-and-control binary. The binary hit a permission-denied error — and Devin opened a second terminal, added the execute bit, and ran it again. The researcher got a remote shell on the DevBox, and from that shell reached AWS key material present on the machine. [S7]
Three details make this materially worse than a single successful injection. First, the researcher reproduced the attack from Slack, which he characterises as "an entirely unsupervised interaction" — no human was watching a screen when the agent was compromised. Second, persistence was maintained with a session-connected reaction that fired commands the instant the callback landed, so secrets could be exfiltrated in milliseconds even when Devin later cancelled the command. An operator watching carefully and hitting stop does not prevent the exfiltration. Third, the researcher's general observation across agents was that they "like clicking links", and that once an agent is off-domain on an attacker's page, payloads "often work at first try". [S7]
The disclosure history is itself a finding. The vulnerability was reported to Cognition on 6 April 2025 and acknowledged within days. Follow-up queries on fix timelines, status and coordinated disclosure went unanswered for more than 120 days, after which the researcher published. [S7] We searched for a vendor advisory, a fix note or any public response and found none. For a risk team, an unanswered disclosure is a signal about process, not just about one bug: it speaks to how the next report will be handled once the agent is inside your perimeter.
What Cognition does publish, and it is substantial
It would be unfair and inaccurate to present this vendor as having no tool-control story. The CLI permission model is one of the better-documented we have assessed, and a risk team should read it directly. [S5]
- Five modes with a published per-tool matrix. Normal (reads auto; writes and shell prompt), Accept Edits, Smart (a fast model judges each non-edit action), Bypass (everything auto), and Autonomous, which pairs with an OS-level sandbox that bounds filesystem and network reach.
- Deterministic rule resolution. Deny → ask → allow → default-prompt, and a deny always wins; an allow never overrides a matching deny however specific it is.
- Real scopes, including egress.
Read(glob),Write(glob),Exec(prefix)andFetch(url-pattern). The Fetch scope is the one that directly addresses the attack above: egress can be restricted to named domains, and the exploit depended on navigating to an attacker-controlled host. - Categories never auto-approved in Smart mode. Package installs, mutating
git,rmandsudo, destructive cloud-CLI andkubectl delete, and anything that reads or writes dotenv files, key material, git config or the agent's own configuration. - Organisation-level precedence, and this is the control to lean on.Admin-enforced deny and ask rules configured in Team Settings cannot be overridden by project or user config, and remain active regardless of the user's permission mode — including Bypass. An enterprise can therefore impose a floor a developer cannot lift.
The vendor also publishes its own limitations plainly — Devin "may still" hallucinate, introduce bugs and suggest insecure coding practices — and recommends code review before deployment, branch protections and standard engineering review. [S3] That candour is worth crediting.
Why the rating is still HIGH.The controls are documented for the CLI. The compromise was demonstrated on the cloud agent and from Slack, where we found no equivalent published matrix. The researcher's central recommendation to the vendor was to stop relying on model behaviour or in-chat confirmations and build an out-of-band validation step for sensitive operations, along with blocking outbound connections by default, warning users not to join Devin to an enterprise or VPN network, and noting the absence of endpoint protection on the hosts. [S7]Smart mode's model-judged approval is a mitigation of exactly the class the researcher warned against relying on — a model deciding whether another model's action is safe. It raises the bar; it is not an out-of-band gate.
OWASP relevance: LLM01 (Prompt Injection) and LLM02 (Insecure Output Handling), compounded by LLM08 (Excessive Agency). A write-capable agent with a shell, a browser and network egress, consuming untrusted third-party text as context, is the canonical LLM01-into-LLM08 escalation, and here it is demonstrated rather than inferred.
3.3 — D3 · Credentials & API Keys
Threat model: Devin needs credentials to do useful work — a database URL to run migrations, a cloud key to deploy, a token to call an internal API. Those credentials are present on the machine the agent controls, at the moment untrusted text enters its context.
A first-party Secrets feature on the Settings page is the documented mechanism for API keys, passwords and cookies. [S3] The dedicated secrets documentation page returned 404 at the time of access, so the feature's internals are not publicly evidenced: we cannot state how secrets are encrypted at rest, whether they are scoped per-session or per-repository, whether values are masked from the agent's own context window, or whether rotation is supported. Those are four questions a risk team must ask directly.
Repository access is reasonably designed. The administrator installing the GitHub integration selects which repositories Devin can reach and can adjust permissions at any time from GitHub App Settings — a per-repository grant under the customer's own control, which is better than a personal access token. Slack processing is scoped to explicit invocation: an @Devin tag, a direct prompt, or information shared in an active thread. [S3]
Against that sits the demonstrated exposure.The published exploit reached AWS key material from an attacker-controlled shell on the DevBox, and the researcher's explicit advice to users is that "any secret or private code Devin has access to can be leaked at will to third-party systems by the AI, or an attacker via indirect prompt injection". [S7] The CLI permission model does deny reads and writes of dotenv files and key material in Smart mode, which is a direct and well-aimed mitigation [S5] — but the exploit was not on the CLI.
Identity controls are tier-gated. SAML/OIDC SSO is Enterprise-only, and no SCIM provisioning is named anywhere on the published plan comparison. [S8]For an organisation of any size this is a joiner-mover-leaver problem: without SCIM, deprovisioning a departing engineer's access to an agent that holds repository write access is a manual step someone has to remember.
OWASP relevance:LLM06 (Sensitive Information Disclosure) and LLM08 (Excessive Agency). Credentials co-resident with an injectable agent on a machine with network egress are exfiltratable by design unless egress is restricted or the secrets are held outside the agent's reach.
3.4 — D4 · Payment & Financial Operations
Threat model: the agent's financial exposure. Two distinct questions — can it move money, and can it spend more than budgeted.
On the first question the answer is clean. We found no published capabilityfor Devin to hold spend authority, move funds, or transact on a customer's behalf. There are no payment rails in scope, and a risk team can close that line. The exposure is procurement spend.
The finding is that cost per unit of work is not deterministic, and no published spend cap exists.Each paid plan carries an allowance refreshing daily and weekly; beyond it, extra usage is purchased and "consumed at API pricing". The vendor states plainly that cost per message varies with the model used, task size, complexity and reasoning required. [S8]A budget owner therefore cannot derive a worst-case monthly figure from the pricing page, and the overage meter is denominated in a third party's API pricing rather than the vendor's own rate card. Concurrent-session limits — up to 10 on Free, Pro and Max, unlimited on Team and Enterprise — bound parallelism but not spend. [S8]
Centralised billing and an admin dashboard with analytics arrive at the Team tier and above, which is where visibility becomes possible at all. [S8] One live commercial term is time-boxed and worth diarising: SWE-2 is free in Devin Desktop and the CLI only through 10 October 2026. [S8][S14] Any cost model built on current pricing should be re-run after that date.
A second-order risk worth naming. An injected agent that can run arbitrary shell commands on metered infrastructure is a cost-amplification path as well as a security one. Unlimited concurrent sessions on Team and Enterprise tiers removes the natural ceiling.
3.5 — D5 · Operational Continuity
Threat model: Devin sits on a delivery path with a deadline. The question is what happens when it is unavailable, and what the vendor has committed to.
Cognition operates a public Atlassian status page with per-component 90-day uptime and subscription by email, SMS, Slack, Teams, webhook and RSS. Transparency here is above the market norm for AI vendors, and the numbers are specific enough to act on. [S12]
99.77% on the non-enterprise Cloud Agent is roughly five hours of unavailability per month. Enterprise components are measurably better than their general counterparts across the board, which is a real and quantified reason to price the Enterprise tier rather than a sales claim. Two incidents inside the eleven days before our observation cutoff were publicly written up with timestamps and resolved: 24 September 2026, "Devin webapp down", investigating 19:59 UTC to resolved 22:03 UTC, about two hours; and 17 September 2026, "Multiple services experiencing degraded performance", 13:40 to 15:20 UTC, about an hour forty. [S12] Both were handled and communicated properly. Incidents happen; the honest write-up is the good sign.
§2.5 of the platform terms states the Services are subject to modification and change at Cognition's sole discretion and that "There are no guarantees made with respect to the quality, stability, availability, or reliability of the Services." We found no SLA with an uptime commitment or service credits on any public page. [S11][S12]
The gap between observed operational behaviour and contractual commitment is the finding. Good uptime that is disclaimed in the binding instrument is not a commitment a risk lead can put in a file. The trust centre does declare a Recovery Time Objective of 48 hours, a Data Access Level of "Internal" and an Impact Level of "Substantial", and asserts a business continuity plan with disaster-recovery testing whose frequency sits behind the NDA. [S4] A 48-hour RTO is a number worth knowing before Devin goes on a critical path.
Upstream dependency is distributed but not disclosed in detail. The platform routes to OpenAI, Anthropic, Google and open-source models alongside in-house SWE-2, which means no single provider outage takes the product down and gives Cognition genuine substitution power — a real resilience advantage over a single-provider agent. [S8] But we found no published dependency map, fallback commitment, or statement of which provider serves which workload, so the customer cannot verify the resilience or know whose terms govern their data on a given request. Monitoring is asserted: continuous logging, error handling, real-time dashboards and alerting on unusual application states. [S3]
3.6 — D6 · Legal & Liability
Threat model: the agent causes harm — it leaks code, corrupts a repository, or acts on injected instructions. The question is who carries the loss.
§10.2 caps each party's aggregate liability at the greater of amounts paid to Cognition in the six months preceding the claim, or one hundred US dollars. §10.1 excludes consequential, incidental and punitive damages, expressly including loss of data and "breach of data or system security". [S11]
Put that against a Teams deployment. At $80/month plus $40 per seat, a twenty-seat team pays roughly $5,300 over six months — that is the recoverable ceiling for any claim, against an agent holding repository write access and credentials. The excluded categories are precisely the ones that matter here: a prompt-injection incident that exfiltrates source code is a loss of data and a breach of system security, both named exclusions. This is not unusual drafting for SaaS at this price point; it is material because the product's blast radius is not usual for SaaS at this price point.
Data-breach responsibility is narrowed further.§4.1 provides that Cognition "will not be responsible for any breach in security except to the extent the breach is due to Cognition's gross negligence" — a markedly higher bar than ordinary negligence. The same section puts routine backup on the customer, disclaims liability for loss, alteration, destruction or corruption of Customer Data, makes the customer responsible for notifying its own employees and customers of a breach and for regulator filings, and has the customer indemnify Cognition against claims from authorised users or data protection authorities on those obligations. [S11]
§8.2 provides that where output includes "assessment, review, analysis, evaluation, or examination of code, configurations, security posture, vulnerabilities, defects", Cognition makes no representation that it identifies all relevant issues, and such output is "a non-exhaustive aid only and is not a substitute for your own review, testing, audit, or other independent verification". [S11]
Anyone planning to put Devin into a security review, a code-audit or a compliance workflow should read that clause before designing the control. The vendor has disclaimed, in the binding instrument, exactly the reliance such a workflow would place on it. Everything is "AS IS" with all implied warranties disclaimed, including any warranty that the Services will "be secure". [S11]
What is genuinely favourable in these terms
- Output IP belongs to the customer.Generated code and work product are the customer's intellectual property and may be used commercially, with the only restriction being that output may not be used to train a model intended to reverse-engineer or build a competing product. That is a clean and commercially reasonable position. [S3]
- An IP indemnity exists — for paid tiers. §9.1 has Cognition defend third-party patent, copyright and trade-secret claims against the Services used in accordance with the agreement, subject to the standard carve-outs (customer specifications, combination with non-Cognition software, modification, failure to follow documentation, unauthorised use, stale version, Customer Data). §9.3 makes it the sole remedy for IP claims. The Free tier is expressly excluded, which matters for anyone piloting on free seats. [S11]
- Terms are versioned and observable. A 30-day change mechanism and a linked prior version from April 2025 mean a buyer can diff the terms over time — better practice than most. [S11]
One clause deserves a specific warning to risk teams.§3.5 assigns all Feedback to Cognition — including "any error, problem or defect reports" — with all right, title and interest assigned, no attribution or compensation, and the Feedback deemed Cognition's Confidential Information. [S11] Read literally, a detailed security or defect report filed through ordinary support channels is assigned to the vendor and classified as their confidential information. Route security findings through security@cognition.ai under separately agreed disclosure terms rather than through support.
Dispute resolution removes the court and the class.Binding arbitration before the AAA's International Centre for Dispute Resolution under Expedited Commercial Rules, a single arbitrator, in English, in New York, New York, with a jury-trial waiver and a class-action waiver; each party bears its own legal and expert fees regardless of outcome. Governing law is New York. [S11] For a European buyer, that is a US forum, US law, and self-funded costs on any dispute. A DPA exists but sits behind the trust-centre NDA, not on a public URL. [S4] Customer-side export-control and sanctions representations are imposed, including non-military-intelligence end-use. [S11]
Performance & Traction
Cognition is one of the most heavily capitalised companies in the agentic category: over $400M raised at a $10.2B post-money valuation, with secondary reporting of $26B by May 2026 and Bloomberg describing early discussions at $40B or more in August 2026. [S1][S2] The Windsurf acquisition brought an established IDE user base, rebranded to Devin Desktop in June 2026. [S2]
We deliberately publish no customer-count, revenue or retention figure here, because we found none in public sources we would sign our name to. A logo strip appears on the pricing page; we do not treat unlabelled logos as evidence of a reference-able deployment. [S8]
Two traction signals are worth reading, and they point in opposite directions. In its own funding retrospective the company describes the Devin of March 2024 as "still a very junior engineer" — unusual candour, and a marker of how fast the product has had to move. [S1] Against that, the current in-house model SWE-2 is being given away free in Devin Desktop and the CLI through 10 October 2026. [S8][S14] A time-boxed free tier on the flagship model is a distribution play in a category competing hard on developer default; it is not a sign of pricing power, and a buyer should assume the economics after that date are not yet knowable.
Relevance to the verdict: traction is a vendor-viability question, and on this file it is answered. Cognition will be here for the duration of a contract. That removes the concern that usually drives caution on an AI vendor review, and it is precisely why the verdict rests entirely on the technical and contractual findings rather than on company risk.
Operational Risks
- Documented controls describe a different product than the demonstrated compromise. The permission matrix, scopes and organisation-level precedence are published for the CLI; the exploit ran on the cloud DevBox and from Slack, where no equivalent matrix was found. A control inventory built by reading the CLI documentation will overstate what governs a cloud session. [S5][S7]
- Unsupervised invocation paths. Slack and Teams integration means sessions start without anyone watching a screen, and the published exploit was reproduced through exactly that path. Any approval model that assumes a human is present is not describing this deployment. [S7][S8]
- No EDR or antivirus on the agent hosts, per the researcher's findings and recommendations, meaning a binary executed on a DevBox meets no endpoint control. [S7]
- Secrets handling is unverifiable from outside. The Secrets feature is referenced but its dedicated documentation page returned 404; encryption, scoping, masking and rotation are all unevidenced publicly. [S3]
- No SCIM, and SSO only at Enterprise.Deprovisioning an engineer's access to an agent holding repository write access is a manual step below Enterprise tier. [S8]
- Metered spend without a published cap, with overage denominated in a third party's API pricing and unlimited concurrent sessions at Team tier and above. [S8]
- A 48-hour recovery time objective against 99.77% observed 90-day uptime on the non-enterprise Cloud Agent, with no contractual availability commitment. [S4][S11][S12]
Reputational Risks
The compliance artefacts are real, and they are all behind an NDA. The trust centre lists SOC 2 Type 2, ISO/IEC 27001:2022 and CCPA, with a SOC 2 report, a SOC 2 Type 2 bridge letter, a pentest report, a network diagram, a DPA and a certificate of insurance available on request. The documentation dates SOC 2 Type II to September 2024, covering data security, privacy, processing integrity, confidentiality and availability. [S3][S4] This is a genuine institutional investment and it distinguishes Cognition from most of the category.
The reputational finding is what cannotbe checked without entering a commercial relationship. Access requires a request, an emailed invite and a signed NDA. Until then a buyer cannot read the audit period, the scope boundary, the auditor's identity, whether there were exceptions, the ISO certificate number and issuing body, or the pentest date and scope. [S4] The certifications are therefore asserted-and-gated rather than verified for the purposes of this report, and we rate them as such.
Several trust-centre sections are assertion without artefact. App Security reads "we are putting together a program to monitor internal apps" — candid, and an acknowledged gap. Data Security, Access Control, Infrastructure, Endpoint Security, Incident Response, Continuous Monitoring and Data Privacy each assert industry best practice and offer "more details ... upon request". [S4]
Vulnerability disclosure exists as a channel but not as a demonstrated process. A named address, security@cognition.ai, is published, and Cognition states it will notify Enterprise customers of incidents affecting their environments per contractual obligations. [S3] Against that, the one substantive published finding went 120+ days without an answer on fix timelines. [S7] A bug bounty is not publicly confirmed: the trust-centre FAQ lists the question "Is there a bug bounty program in place?" and the answer is behind the NDA. [S4]
A capture of Devin's system prompt dated 10 April 2025 is publicly archived on GitHub, which is ordinary for this product class and not a finding in itself. [S7] We found no CVE naming Devin or Windsurf and no public litigation involving Cognition in the sources consulted.
Legal Risks
The legal exposure on this file is not drafting exotica; it is the ordinary SaaS risk allocation applied to a product whose failure mode is not ordinary SaaS. Four clauses carry it, and all four are quoted in §3.6: the liability capat the greater of six months' fees or US$100 with loss of data and breach of system security excluded from recoverable damages; the gross-negligence standard for data-breach responsibility, with backup and breach-notification duties pushed to the customer; the output disclaimer that names security assessments specifically; and the Feedback assignment that captures defect and error reports. [S11]
For a European buyer three further points are procurement-relevant. The DPA is behind the NDA rather than public, so Article 28 terms cannot be reviewed before engaging. [S4] No named sub-processor list is published, which is an Article 28(2) question that cannot be answered by reading. [S9] And dispute resolution is binding arbitration in New York under New York law, with a class-action waiver and each party bearing its own costs. [S11]
The favourable counterweights are real and should be recorded in the same file: customer ownership of output IP, a paid-tier IP indemnity, versioned terms with a 30-day change mechanism and a linked prior version, SCCs and the UK IDTA for transfers, and a published set of data-subject rights. [S3][S9][S11]
Recommendation
Cognition AI, Inc. is a substantial and well-capitalised company with genuine compliance investment — SOC 2 Type 2, ISO/IEC 27001:2022, a public status page with honest incident write-ups, customer ownership of output IP, and a CLI permission model that is among the most carefully specified we have assessed. Vendor viability is not in question on this file.
The verdict is CONDITIONAL GO STRICT rather than CONDITIONAL GO because of the distance between three things that should not be far apart. A working indirect prompt injection achieving remote code execution and credential access on the cloud agent is published, was reported to the vendor in April 2025, was acknowledged, and was released after 120+ days without an answer on fix timelines — and we found no public advisory or fix note for it since. The controls that would bound that attack are documented for the CLI, not for the cloud or Slack paths where it was demonstrated. And the contractual recovery for the resulting harm is capped at the greater of six months of fees or US$100, with loss of data and breach of system security excluded from recoverable damages outright.
None of that makes Devin unusable. It makes the five questions below procurement-blocking rather than best-practice, each to be answered in writing before production access is granted:
The permission model Cognition documents in depth — five modes, deny/ask/allow precedence, Read/Write/Exec/Fetch scopes, and organisation-level rules that project and user config cannot override — is documented for the CLI. The published exploit was executed against the cloud DevBox and reproduced from Slack. Before production access, require in writing: which of those controls apply to the cloud agent and to Slack-triggered sessions, who can change them, and whether an organisation-level deny is enforceable on a cloud session. If the answer is that cloud sessions do not carry an equivalent enforceable gate, restrict Devin to the CLI with organisation-level deny rules and treat cloud and Slack invocation as out of scope for approval.
The enterprise documentation states Cognition does not train on customer data or code by default. The privacy policy lists User Content as processed to train, fine tune and improve the models, under legitimate interests, qualified by 'depending on the terms that apply to your use of the Services'. Both can be true at once, and the published documents do not say which tiers sit on which side. Get a signed statement naming your tier, the retention period for code and session data, and the named sub-processor list including which model provider handles which workload. No public sub-processor list exists, so this cannot be resolved by reading.
Aggregate liability is capped at the greater of six months of fees or US$100. Loss of data and 'breach of data or system security' are excluded from recoverable damages outright. Data-breach responsibility is limited to Cognition's gross negligence, backup duty sits with the customer, and the customer carries its own breach-notification obligations. For a Teams deployment at roughly a few thousand euros per six months, the recoverable ceiling is that figure against an agent holding repository write access and credentials. Either negotiate the cap and the gross-negligence standard, or scope Devin's access so the uncapped residual is one your organisation accepts.
SOC 2 Type 2, ISO/IEC 27001:2022 and a pentest report are all listed, and the documentation dates SOC 2 Type II to September 2024. None of it is readable without requesting access and signing an NDA: not the audit period, not the scope boundary, not the auditor, not the pentest date. Request all of them and check one thing specifically — whether the indirect prompt injection class demonstrated publicly in April 2025 was in the pentest scope, and whether it has been remediated. No public advisory or fix note for it was found.
Ninety-day uptime on the non-enterprise Cloud Agent reads 99.77%, roughly five hours of unavailability a month, against 99.9% on the Enterprise equivalent. Two incidents in the eleven days before the observation cutoff were publicly written up and resolved in about two hours and one hour forty. The status page and the 48-hour trust-centre recovery time objective are real operational signals, but §2.5 of the platform terms states there are no guarantees as to quality, stability, availability or reliability, and no SLA with service credits was found on any public page. If Devin sits on a delivery path with a deadline, the commitment has to be negotiated into the contract.
What would move this verdict to GO. A written confirmation that organisation-level deny rules are enforceable on cloud and Slack-triggered sessions with Fetch restricted to an allow-list; a signed tier-specific schedule on training, retention and named sub-processors; a negotiated liability position that does not exclude breach of system security; and a reviewed pentest report showing the April 2025 injection class in scope and remediated. Four of the five are answerable in a single procurement exchange.
What would move it to NO-GO. Granting a cloud or Slack-triggered Devin session standing access to production credentials, or to repositories that ingest third-party content such as public issue trackers, without an out-of-band approval gate and restricted egress. On the published evidence that configuration is the one demonstrated to fail, and the contract does not carry the loss.
Appendix — Sources, Methodology & Limitations
A.1 — Sources consulted
All sources are public and were accessed on 2026-09-28. Findings above are tagged with the source identifiers below.
- S1 —
cognition.com/blog/series-c· vendor funding announcement · $400M+ at $10.2B post-money, investor list, launch retrospective. - S2 —
en.wikipedia.org/wiki/Cognition_AI· encyclopedia entry citing CNBC and Bloomberg · founding, founders, headcount, Windsurf acquisition, valuation trajectory, Devin Desktop rebrand. Treated as secondary throughout. - S3 —
docs.devin.ai/enterprise/overview· vendor documentation · encryption, AWS, MFA, SOC 2 Type II date, disclosure address, retention, training statement, output IP, GitHub/Slack permission model, stated product limitations, Secrets feature. - S4 —
trust.cognition.ai· vendor trust centre · CCPA, SOC 2 Type 2, ISO/IEC 27001:2022, document inventory, risk profile including the 48-hour RTO, NDA gate. - S5 —
docs.devin.ai/cli/reference/permissions· vendor documentation · the full CLI permission model: modes, rule precedence, Read/Write/Exec/Fetch scopes, MCP permissions, organisation-level precedence, never-auto-approved categories. - S6 —
docs.devin.ai/enterprise/integrations/github-enterprise-server· vendor documentation · GHES and GHEC-with-data-residency connection methods. - S7 —
embracethered.com, "Month of AI Bugs" episode 6 · independent security research by Johann Rehberger · indirect prompt injection to remote code execution on the DevBox, credential access, Slack reproduction, persistence, disclosure timeline, vendor recommendations. - S8 —
devin.ai/pricing· vendor pricing page · all published tiers and prices, quota and overage model, integration matrix, tier-gated security features, model providers. - S9 —
cognition.com/legal/privacy-policy· vendor legal · entity name, User Content training purpose under legitimate interests, retention framing, transfers and SCCs, third-party categories, data-subject rights. - S10 —
cognition.com/legal/terms-of-service· vendor legal · website terms, pointing to the separate platform terms for the service. - S11 —
cognition.com/legal/platform-terms-of-service· vendor legal · §2.5 availability, §3.2 customer-data licence, §3.5 Feedback assignment, §4.1 breach standard, §8 AS IS and output disclaimer, §9 indemnification, §10 liability cap, §11 arbitration, §12.6 governing law, §12.10 export. - S12 —
devinstatus.com· vendor status page · per-component 90-day uptime and the two September 2026 incidents with timestamps. - S13 —
devin.ai· vendor marketing site · product surface inventory. - S14 —
cognition.com/blog/swe-2· vendor blog · SWE-2 as current in-house model and the free-through-2026-10-10 term. - Reference framework: OWASP Top 10 for LLM Applications.
- Pages fetched that returned 404 and yielded no evidence:
docs.devin.ai/product-guides/devin-secretsanddocs.devin.ai/collaborate-with-devin/devin-api. The Secrets feature and the Devin API are therefore evidenced only indirectly, through S3 and S8.
A.2 — Methodology
TrustworthAgent Express Security Reports are desk-based assessments conducted using exclusively publicly available information. The six security dimensions assessed in §3.1–§3.6 constitute the TrustworthAgent Security Framework and are identical on every report we publish, so that two vendors can be compared on the same axes and a report written a year from now remains readable against this one. Each dimension is assessed against documented architectural choices, published policy commitments, known vulnerability classes from the OWASP LLM Top 10, and comparable market standards at equivalent company stage and product category.
Risk ratings are drawn from the fixed set LOW / MEDIUM / MEDIUM-HIGH / HIGH and reflect assessed materiality relative to the enterprise deployment scenario described in each threat model, not absolute severity in isolation. The recommendation is drawn from the fixed set GO / CONDITIONAL GO / CONDITIONAL GO STRICT / NO-GO.
A.3 — Limitations
No testing of any kind was performed against Cognition or Devin for this report.No penetration testing, no code review, no internal interviews, no non-public data access, and no purchase of the product. We did not reproduce the exploit described in §3.2; we report it as published third-party research by a named researcher, with its disclosure timeline as stated by that researcher, and we could not independently verify the vendor's response or the current remediation status.
This report reflects publicly available information as of the observation cutoff, 2026-09-28, 08:30 UTC. Cognition's security posture, documentation and policy commitments may have changed since. The following were material to our assessment and could not be verified from public sources: the SOC 2 audit period, scope and auditor; the ISO 27001 certificate number and issuing body; the pentest date, scope and whether the injection class in §3.2 was covered; a named sub-processor list; whether the §3.2 vulnerability has been remediated; any per-tool permission matrix for the cloud agent and Slack paths; the Secrets feature's encryption, scoping, masking and rotation behaviour; the existence of any SLA or bug bounty; SCIM availability; the Windsurf acquisition terms; and insurance limits. Each of these sits behind an NDA, a 404, or is simply unpublished. Where a fact was unavailable we have said so rather than inferred it.
Risk ratings for compliance artefacts reflect verifiability, not quality. Cognition may well hold excellent, broadly scoped certifications; we cannot read them, so we cannot rate them as verified, and we decline to credit what we have not seen.
A.4 — Non-affiliation and right of reply
This is a free, public, unsolicited report published by TrustworthAgent on 2026-09-28 (v1.0) as a demonstration of the Express Security Report methodology. TrustworthAgent has no commercial, financial, shareholding, contractual, partnership or mandate relationship with Cognition AI, Inc., its officers, its investors or its representatives, and holds no position in the company. We were not asked to write this report and were not paid to write it. We do not invest, we do not rate consumers, and we do not sell signals about the companies we audit.
Observations, risk ratings and recommendations are opinions of analysis founded on the methodology described and the sources cited. They are not exhaustive or definitive statements of fact, and they are not investment, legal, tax, accounting, cybersecurity or regulatory advice. Any reader relying on this for a procurement or investment decision should commission a full mandate, which is scoped in writing and can extend to proprietary architecture documentation and authorized testing where the subject grants access.
Cognition AI, Inc. has a right of reply. To request correction of a factual error or publication of a response, write to hello@trustworthagent.com citing this URL, the passages concerned, the supporting facts or sources, and the name and role of the person replying. We undertake to examine any good-faith request and, where it is founded, to correct the error, publish the response or add a dated update note within 5 business days of a complete request.
Get the next TrustworthAgent security due diligence report in your inbox. One report per publication. No noise.
Need due diligence on a specific autonomous business?
Name the subject and the decision it supports. We reply in writing within one business day.
Or order directly: Express Security Report — €149 · 48 hours
Independent audit · TrustworthAgent
An independent audit on your own agent, or on a vendor.
Express Security Report — 5 pages, all six dimensions, delivered in 48 hours. For a fast go or no-go before an integration, a partnership or an investment.
Full audit mandate, by quote: request a written scope
Questions: hello@trustworthagent.com