Cognition's Devin uses GPT-6 Astra to test its own code

OpenAI has announced that Cognition's coding agent Devin now uses GPT-6 Astra to test the software it writes and to demonstrate that the work is correct. The stated purpose is to cut down how much code human engineers need to review by hand, so that more of it can ship without that manual check. OpenAI's announcement, posted to its own site, does not go beyond that single claim: it gives no release date, no explanation of how GPT-6 Astra actually performs or verifies the testing, no benchmark numbers or adoption figures, and no comparison to how Devin tested its work before this change. No engineer, executive or spokesperson is named or quoted on either side. The underlying announcement page returned a Cloudflare block on a follow-up fetch, so this account rests on the title and text OpenAI surfaced through its index rather than the full write-up.

Key facts

  • GPT-6 Astra, an OpenAI model, now improves Devin's ability to test software it has written.
  • Devin is Cognition's coding agent; the integration lets it show that its own work is correct.
  • OpenAI frames the goal as letting engineers review less code while shipping more of it.
  • No mechanism, benchmark, date or named spokesperson accompanies the announcement.

Why it matters

A coding agent that can test and demonstrate the correctness of its own output addresses the main bottleneck in agentic coding: a human still has to review everything the agent produces before it ships. If GPT-6 Astra genuinely strengthens that self-testing step, the practical effect OpenAI describes is engineers reviewing less code by hand while more of it goes out the door.

Who it affects

The direct parties are Cognition, which builds and ships Devin, and OpenAI, whose model now sits inside that testing step. Engineering teams already using Devin are the ones who would see less manual review; the announcement does not name any specific customer or team beyond Cognition itself.

How to use it

The source gives no pricing, licensing, availability date or rollout details for this capability, so there is nothing concrete to report here beyond the fact that the integration exists between Devin and GPT-6 Astra.

How solid is it

The claim comes directly from OpenAI's own announcement, which is a first-party source rather than a third-party report, but the announcement itself is thin: one sentence describing the capability and its goal, with no supporting data, methodology or named individual behind it. A follow-up attempt to read the full announcement page was blocked, so nothing beyond that sentence could be independently confirmed.

Risks and caveats

Without a description of how GPT-6 Astra tests or verifies Devin's work, or any figures on how often that testing catches problems, it is not possible to judge how much manual review this actually replaces. The claim that engineers can review less code rests entirely on OpenAI's own framing, with no independent measurement offered.