Cloud agents get more productive as they scale: Benchmarks from the Global 1000
Cognition operates the largest install base of production enterprise cloud agents in the world. We took a hard look at what our usage data says. Here’s what we found.
The prevailing narrative in enterprise AI is that spending is out of control. Token bills blow through budgets with little to show for it.
That narrative hasn’t matched what we’ve seen on the ground.
We’ve spent the last two years deploying cloud agents into the largest enterprises on the planet, building one of the biggest observable populations of software engineering agents in production.
These aren’t pilots. Every company operates under real regulatory, security, and competitive pressure.
So we went to the data: aggregated, anonymized usage from March through July 2026 across Global 1000 companies in every major vertical and region.
Our analysis found that, deployed well, cloud agents show real economies of scale: cost per output falls while usage and the value delivered keep growing.
Three things stand out in the data:
- Cloud agents get more efficient as they scale
- The gains reach every stage of the SDLC
- Spend focuses on the work with the highest output
Today we’re publishing the data as a benchmark for engineering leaders evaluating their own cloud agent strategy. These are the same metrics our product and applied AI teams use with customers.
A note on measurement: a session is any human interaction with the Devin cloud agent, from a research question to writing code. Productive hours estimate the engineering time saved in each session, using a framework Cognition built for this purpose. ACUs (agent compute units) measure the compute consumed.
Finding 1: Cloud agents become more efficient with scale
A persistent assumption is that AI costs scale faster than results: pilots look efficient because they’re small, then the economics fall apart in production. Our data shows the opposite. Dollar cost per session fell 21% below the March baseline while session volume grew 147% (Figure 1).
Dollar cost per merged PR was choppier but ended the period just below baseline, even as merged PR volume grew 80% (Figure 2).
Cost per merged PR varies month to month as the use-case mix shifts. June skewed toward low-cost CI/CD work; July toward complex migrations and re-platforming, which pushed costs back up.
Why does efficiency improve as cloud agents scale?
Cloud agents give organizations what local tools cannot: a way to make AI usage more efficient over time.
Local tools give admins zero telemetry. Nobody can monitor usage, so there is no way to target training, apply governance, or funnel spend to productive use cases. Skills that drive efficiency or productivity remain siloed with individuals or teams.
Top-down usage caps don’t improve ROI, either. Caps throttle your most productive, highest-consumption users and leave everyone else untouched. You trade business impact for the appearance of efficiency.
In contrast, organizations that effectively deploy cloud agents see rising usage on productive use cases while cost per outcome falls. The main driver is that cloud agents enable centralization: a playbook or integration that makes one team faster compounds across the whole organization.
Centralized enablement gets the right training to the right users and moves the entire population toward the most efficient patterns. Deployment, integrations, and context management are handled in one place and raise the baseline for everyone. And unlike blunt caps, spend can be steered toward the work that earns it. Devin also scores every session on the quality of its use case and prompt, so admins can see exactly who needs coaching, and on what.
In short, centralization does three things:
- Individual users get better with enablement and experience.
- Central governance shifts the whole organization toward high-value, high-efficiency use cases (more on this in Finding 3).
- Teams compound what they learn about common workflows.
Finding 2: Cloud agents accelerate the entire SDLC
When AI only accelerates code generation, the bottleneck just shifts elsewhere and cycle time doesn’t accelerate. Backlogged strategic projects get a minor boost, at best, while AI adds slop and complexity. Code quality degrades and security risk climbs.
Ultimately, the business impact gets blunted and the engineering organization isn’t more effective.
We built Devin to remove bottlenecks across the entire SDLC, and we measure the results in productive hours: an estimate of engineering time saved grounded in real-world, publicly available benchmarks.
The data shows that cloud agents accelerate every stage (Figure 3). Plan and Design and Test each account for 27% of all productive hours, edging out Build’s 23%, with the remainder in Secure and Deploy.
Similar to the efficiency story, centralization enables cloud agents to accelerate the entire SDLC.
Governance and enablement push agents into Plan and Design and every stage after Build. Stage-specific agent personas fit the work at each stage, and automations cover use cases like CI/CD and security scanning. The later stages accelerate right along with Build.
Finding 3: Spend and output focus on high-value work
While usage spreads across the SDLC, dollars concentrate in Build and Test because that is where customers point cloud agents at complex strategic projects with high ROI.
High ACU spend and merged PRs converge on complex, high-value Build and Test use cases such as refactors, migrations, and feature development. The two stages account for 66% of spend and 74% of merged PRs (Figure 4). Spend is concentrated, and output is concentrated even more.
Set against the previous finding, it’s clear that sessions and total productive hours spread evenly across the lifecycle; money and merged code do not. Centralized governance steers Devin spend toward high-value outcomes rather than letting it leak into low-value or unintended use cases.
Build and Test use cases also deliver the most productive hours per session (Figure 5).
While total productive hours spread uniformly across the lifecycle, the sessions that generate the most productive hours are aligned to projects that generate business impact. The use cases in dark blue below fall into the ‘Build’ stage and generate high productive hours per session. Feature development, migrations, and refactoring are where Devin most often runs strategic projects, and those projects come with repeatable patterns and playbooks.
Conclusion
Many enterprises are slowing AI adoption because public data points to runaway costs. The caution is understandable: nearly all of that data comes from local agents, which still dominate AI software engineering.
The data in this report paints a different picture.
Cloud agents deployed centrally got more efficient as they scaled, with cost per session falling 21% while session volume more than doubled. The gains reached every stage of the SDLC, and spend concentrated in the complex Build and Test work that drives business outcomes.
Devin is built to drive productivity and ROI across the entire SDLC. With persistent context and central management, It handles long-horizon execution that carries projects all the way to merged PRs.
We’re still early in enterprise AI adoption, and we hope these benchmarks help everyone move faster.
Reach out to us if you’d like to compare notes or brainstorm how cloud agents can work inside your organization.




