All articles

The Data Behind Why Big AI Agent Deals Go Hybrid, Not Off-the-Shelf

Anthropic's 2026 State of AI Agents Report found 47% of enterprises now blend off-the-shelf agents with custom builds. MIT and Caylent data explain why: internal-only builds fail twice as often, and 98% of buyers have hard conditions before an agent touches production.

6 min read
Felt puppet character in a navy suit standing in a glass-walled boardroom, presenting a wall display of abstract dashboard panels and a connected node diagram, holding a felt tablet

Introduction

Every enterprise buyer evaluating AI agents in 2026 is running into the same fork in the road: build it internally, buy it off the shelf, or bring in a partner who does both. The data on that decision is now big enough to be conclusive, and it points in one direction for the kind of large, complex deployments that actually move a P&L. It also happens to describe exactly the kind of work Workmate does — building and deploying custom AI agents for teams, not selling a generic chatbot license.

Three separate research efforts published in 2026 line up on this. Here's what each one found, and what it means for a team sizing up a real AI agent contract rather than a pilot.

Most large organizations have already stopped choosing sides

Anthropic's 2026 State of AI Agents Report, based on more than 500 technical leaders and real deployments at companies like Novo Nordisk, Doctolib, L'Oréal, and Shopify, found that 47% of organizations now take a hybrid approach — combining off-the-shelf agents with custom-built components. Only 21% rely entirely on pre-built tools, and just 20% build everything in-house using raw APIs or open-source models.

The report is blunt about why the hybrid model dominates: "no single approach delivers everything organizations need. Off-the-shelf agents get teams running quickly but often lack the customization required for specific workflows or proprietary systems. Fully custom builds offer control and differentiation but require" more than most internal teams can sustain alone.

That's not a hedge. It's a description of a market that tried both extremes and settled on neither. It also describes the actual work of embedding custom agents into a company's existing tools and permissions — connecting the packaged pieces a team already trusts to the workflow-specific logic that only a bespoke build can provide.

The gap between internal builds and outside expertise is not small

Separately, MIT's Project NANDA published The GenAI Divide: State of AI in Business 2025, a study built on interviews with 52 organizations, a survey of 153 senior leaders, and a review of over 300 publicly disclosed AI deployments. Its headline finding — that 95% of generative AI pilots show zero measurable P&L impact — has been widely quoted. The more useful number for a buying decision is buried a few pages in: external partnerships succeed at roughly twice the rate of internal builds using generic tooling.

The report calls this the "implementation advantage," and identifies it as one of four defining patterns in what it calls the GenAI Divide — the split between the small number of organizations extracting real value from AI agents and the much larger group stuck with pilots that never scale. MIT's own explanation for the gap isn't model quality. It's what they call the "learning gap": most generic systems don't retain feedback, adapt to a company's actual workflow, or improve after they ship. Vendors who build with a specific customer's processes in mind close that gap; internal teams reaching for a generic framework usually don't.

Enterprise buyers already know what they'll demand before they say yes

The third data point is about what happens once a deal is on the table. Caylent's 2026 Enterprise Readiness for Agentic Engineering & Autonomous Cloud Operations survey, conducted by Censuswide among 200 senior enterprise leaders at organizations with 1,000+ employees, found that 98% of respondents have specific conditions they require before letting an AI agent run autonomously in production. Just 2% say no set of conditions would satisfy them at all.

The conditions aren't really about whether the model is smart enough. 83% of respondents told Caylent that guardrails matter as much as, or more than, model intelligence when deciding whether to scale a deployment. As Caylent's CTO Randall Hunt put it: "the question of whether enterprises will adopt agentic AI is settled. What's left is authority, not accuracy." In practice, that means the real work in landing a large agent deployment is the architecture around it — who approves what action, what's logged, what a human has to sign off on before an agent touches a live system — not the underlying model.

What this means for a team building custom AI agents

Put the three findings next to each other and they describe a specific kind of buyer: one who has already tried the generic tool, is skeptical of a pure in-house build, and has a checklist of governance requirements before anything gets anywhere near production. That's a different sales conversation than "our AI agent is smarter than theirs." It's a conversation about integration into the systems a company already runs, permissioning that satisfies a security review before it ever gets asked, and a build process that treats the customer's actual workflow — not a generic use case — as the spec.

For teams sizing up a big AI agent engagement in 2026, the numbers are less a validation of the AI agent market broadly and more a description of exactly what a large deal now requires to close: hybrid rather than all-or-nothing, built with outside expertise rather than assembled from a generic toolkit, and designed around guardrails from day one rather than bolted on after the fact.

Frequently Asked Questions

Why do hybrid AI agent approaches outperform fully custom or fully off-the-shelf ones?

Off-the-shelf tools get teams running fast but rarely fit a specific company's workflows or proprietary systems without modification. Fully custom builds offer control but demand engineering capacity most internal teams don't have to spare. A hybrid model — packaged components plus workflow-specific custom logic — captures the speed of the first and the fit of the second, which is why Anthropic's 2026 data found it's now the majority approach among enterprises.

What is the "learning gap" MIT's report identifies?

MIT's Project NANDA found that most generative AI systems that fail to scale share one trait: they don't retain feedback or adapt to how a specific team actually works. They're static tools bolted onto dynamic workflows. Systems built or customized around a particular organization's processes close that gap, which is part of why vendor-led implementations in the report succeeded roughly twice as often as internal builds.

What guardrails do enterprise buyers expect before approving an AI agent for production?

Caylent's 2026 survey found enterprise leaders weight guardrails — approval workflows, audit trails, defined limits on what an agent can act on autonomously — as highly as or higher than raw model capability. In practice that means clear logging, human-approval checkpoints for consequential actions, and scoped permissions are expected as part of the build, not optional add-ons requested later.

Does "hybrid" mean a company still needs an outside implementation partner?

Not automatically, but the same 2026 data suggests it usually helps. MIT's research links outside partnerships to roughly double the success rate of internal-only builds, largely because specialized implementers bring both the engineering capacity for custom integration and the experience to design the guardrails enterprise buyers require before signing off.

Workmate

See what an agent team would do for your business.

Talk to us about Workmate