Nobody Is Grading the Agents: The Eval Gap Inside Your Revenue Stack

Written by: Sarah Mitchell Updated: 08/04/26
11 min read
Nobody Is Grading the Agents: The Eval Gap Inside Your Revenue Stack

Ask your RevOps lead a simple question this week: how good is our AI?

Not "how much are we using it." Not "how many seats did we provision." Not "what did the vendor's ROI calculator say." Just: of the outputs our AI produced last month — the drafted follow-ups, the summarized calls, the scored accounts, the auto-routed leads, the renewal risk flags — what percentage were actually correct?

Most revenue organizations cannot answer that question. Not because the answer is embarrassing, but because nobody is measuring. There is a dashboard showing how many times the AI ran. There is no dashboard showing whether it was right.

That absence has a name in the engineering world, where teams have been fighting it for two years: the eval gap. And it has quietly become the single largest source of unrecognized risk in B2B go-to-market.

For CROs, RevOps leaders, marketing operations, and customer success executives who have deployed AI into customer-facing workflows and have no systematic way to know whether it is performing. This is not an article about whether to use AI. That argument is over. This is about the measurement layer that should have shipped alongside it and mostly did not.

The 37-point gap

The clearest evidence comes from the engineering teams building this stuff, not from the buyers.

LangChain's 2026 State of Agent Engineering report surveyed more than 1,300 practitioners — engineers, product leaders, and executives working on production AI systems. Fifty-seven percent of organizations now have agents running in production, up from 51% the year before. Another 30% are actively working on deployment. The technology has crossed from pilot into the operational core.

Then comes the number that should stop you. Among teams with agents in production, 89% have implemented observability — but only 52% have evals.

That 37-point gap is the whole story, and it is worth being precise about the distinction, because the two words sound interchangeable and are not.

Observability tells you what the AI did. Evaluation tells you whether what it did was any good. Observability is the flight recorder: every tool call, every prompt, every token, every latency spike, logged and searchable. It is genuinely useful — when something breaks loudly, you can trace it. Evaluation is the grading rubric: a defined standard of correctness, applied systematically to outputs, producing a score you can track over time and regress against.

Nine in ten teams built the flight recorder. Barely half built the rubric. Which means the majority of production AI systems can tell you exactly what happened and have no principled basis for saying whether it should have.

It follows, then, that quality is the top barrier to deployment in that same survey — cited by 32% of respondents, the largest single blocker, unchanged from the prior year. Teams are not stuck because the models are too weak. They are stuck because they cannot prove the models are strong enough. And the pilots reflect it: LangChain's data puts the share of agent pilots that never reach production at 88%.

Why revenue teams are uniquely exposed

Engineering teams at least know the eval gap exists. They have a vocabulary for it, conference talks about it, and a tooling category forming around it. Revenue organizations mostly do not — and they are structurally more exposed for three reasons.

First, the output is inherently subjective. When AI writes code, the test suite either passes or it does not. When AI drafts a prospecting email, a deal summary, or a renewal risk assessment, there is no compiler. "Good" is a judgment call, which makes it feel unmeasurable — so nobody measures it. That instinct is wrong, and we will come back to why.

Second, the output goes straight to a customer. A bad internal artifact costs you rework. A bad customer-facing artifact costs you a deal, a renewal, or a reputation. Revenue AI operates at the exact boundary where errors become external and permanent, and it typically does so without the review gates that any other customer-facing process would carry.

Third, and most damaging: revenue AI is built on data nobody trusts in the first place. This is the compounding problem. Only 35% of sales professionals say they completely trust the accuracy of their pipeline data, and 76% of CRM users report that less than half of their CRM data is accurate and complete. You cannot evaluate an AI system's output without a reliable notion of ground truth, and for most revenue teams the system of record is not one.

The downstream effect is measurable in the one place revenue leaders already look. Forecasting research consistently finds AI-assisted forecasts land at 85% to 95% accuracy when built on clean milestone data — and 50% to 60% when they are not. Same model, same vendor, same dashboard. A forty-point spread determined entirely by inputs the AI cannot see and the buyer never audited. And only 7% of sales organizations achieve forecast accuracy above 90% at all.

AI does not fix bad data. It industrializes it, scales it, and attaches a confidence score to the result. The confidence score is what makes it dangerous. A messy spreadsheet looks messy. A confidently wrong AI output looks like an answer.

The governance mistake: treating trust as binary

If you want to understand how this fails in practice, Gartner has been unusually blunt about it.

In a May 2026 release, the firm warned that applying uniform governance across AI agents will itself cause enterprise AI failure — and predicted that by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps discovered only after a production incident.

Read that qualifier again: identified only after production incidents occur. The gap was always there. The organization found out when a customer did.

Shiva Varma, Senior Director Analyst at Gartner, put the root cause plainly: "Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure." Agents, he noted, operate at different autonomy levels and across different trust boundaries — and governing them all identically means over-controlling the harmless ones and under-controlling the dangerous ones.

Revenue leaders will recognize this pattern instantly, because it is exactly how AI got deployed across GTM. The call summarizer and the autonomous outbound agent went through the same approval process — which is to say, roughly none. A tool that drafts an internal meeting recap and a tool that sends a pricing statement to a prospect were treated as the same category of thing: "an AI feature in a tool we already own."

They are not the same category of thing. One produces a document a human reads and discards. The other produces a commitment your company may have to honor.

The macro numbers say almost nobody has sorted this out. McKinsey's State of AI Trust 2026 survey — roughly 500 organizations, fielded December 2025 through January 2026 among people directly responsible for AI governance and investment — found that only about a third of enterprises meet governance standards for autonomous agents, and roughly the same share reach a meaningful maturity level in agentic governance. Meanwhile 62% are at least experimenting with agents and 23% are scaling them somewhere. Deployment is running well ahead of the ability to supervise it. Nearly two-thirds of respondents named security and risk as the top barrier to scaling agentic AI — ahead of regulatory uncertainty, ahead of cost.

And the surface area keeps expanding. Gartner projects that up to 40% of enterprise applications will ship with task-specific AI agents built in during 2026, up from under 5% in 2025. Note what that means for governance: you will not choose to adopt most of these agents. They will arrive in a release note. Your CRM, your marketing automation platform, your support desk, and your CPQ tool will each grow autonomous capabilities on their own roadmap, inside contracts you already signed.

The question is no longer whether AI is operating in your revenue workflows. It is whether anything is grading it.

What an eval actually is

Here is where revenue leaders tend to check out, assuming this is engineering territory. It is not. The core practice is closer to sales QA than to machine learning, and any RevOps team can build a workable version in a few weeks.

Start with a golden set. Take 50 to 100 real cases from your actual workflow — real accounts, real calls, real inbound leads — and have your best humans produce the correct output for each. The right renewal risk rating. The accurate call summary. The follow-up email a strong AE would actually send. This is your answer key. It is tedious to build exactly once and then it is an asset forever.

Define what correct means, in writing. Vague standards produce vague scores. "Good summary" is not a criterion. "Captures every stated objection, names the economic buyer if mentioned, and introduces no fact not present in the transcript" is. Most revenue AI failures are not exotic — they are hallucinated specifics, omitted objections, and confident claims about things the customer never said. Write criteria that catch those.

Score systematically, including with AI. Running a strong model as a judge against your written criteria — the practice engineering teams call LLM-as-judge — lets you grade hundreds of outputs cheaply. It is not perfect, and it needs calibration against human raters on a subset. But an imperfect systematic score you track weekly beats a perfect intuition you never collect.

Sample production continuously. Emerging enterprise practice puts random human review at 5–10% of medium-stakes AI outputs, with 100% review for high-risk actions — anything touching price, contractual language, legal commitments, or a named executive relationship. That tiering is the practical answer to Gartner's binary-trust critique: match the control to the consequence, not to the tool.

Watch for drift. This is the step almost everyone skips. Your AI's quality is not fixed. The vendor updates the underlying model. Your product changes. Your competitors change. Your ICP shifts. A battlecard agent that was excellent in January is confidently outdated by June, and nothing about the interface will tell you. The only thing that catches drift is re-running your golden set on a schedule and watching the score move.

Close the loop. Every failure your review catches should become a new case in the golden set. Over time the answer key accumulates precisely the edge cases your business actually generates, and your evaluation gets sharper in exactly the places your AI is weakest.

Who owns this

In most revenue organizations, the honest answer today is nobody, and that is a decision by default rather than by design.

It should not sit with the vendor. Vendor-reported quality metrics measure the model in general conditions, not your data, your ICP, your pricing, and your definition of a good deal summary. It should not sit with individual reps and CSMs either — they will catch the errors that inconvenience them personally and silently absorb everything else, which is how quality problems become invisible.

The natural home is RevOps, for the same reason RevOps owns CRM hygiene, territory logic, and forecast methodology: it is a cross-functional quality standard applied to a shared system. The practical ask is modest — a named owner, a golden set per high-value use case, a monthly score, and a tiered review policy that distinguishes an internal note from a customer commitment.

The reporting change matters as much as the measurement. Bring an AI quality number to the same meeting where you report pipeline coverage and forecast accuracy. Not usage, not seats, not hours saved — accuracy. The moment AI quality has a number and an owner, it stops being a vibe and starts being managed. Gartner's own trajectory points here: the firm expects explainable AI to drive LLM observability investment sharply upward through 2028, and the enterprises that get there first will be the ones who started tracking a score rather than a spend.

The competitive argument, not just the risk argument

It would be easy to read all of this as defensive — a governance chore to avoid an incident. That framing undersells it.

Consider what an eval layer actually gives you that your competitors lack. You can adopt new models and new vendors fast, because you can measure whether they are better on your work in a day instead of arguing about it for a quarter. You can safely expand agent autonomy where the score justifies it, while your competitors keep everything behind a human bottleneck because they have no evidence to justify letting go. You can tell the difference between an AI problem and a data problem, which is where most of the wasted spend in this category lives.

That last one is the quiet advantage. When an AI initiative underperforms, the default corporate reflex is to replace the tool. An eval layer tells you, in most cases, that the tool was fine and the CRM was not — which redirects budget from a vendor switch that would have changed nothing toward the data foundation that changes everything.

Meanwhile the failure mode for everyone else is already visible in the forecasts: 40% of agentic projects failing by 2027 on unclear value and inadequate controls, 40% of enterprises pulling agents back after incidents they could have caught earlier, 88% of pilots never reaching production at all. Very few of those failures will be caused by a model that was not smart enough. Most will be caused by an organization that deployed something it could not measure, could not defend, and eventually could not trust.

The takeaway

Your revenue team has AI in production. It is drafting things customers read, scoring accounts your reps prioritize, and flagging risks your CSMs act on. Somewhere in that stack, it is also getting things wrong — not catastrophically, not visibly, but steadily, in ways that leak deals and erode trust one artifact at a time.

The gap between the 89% of teams who log what their AI did and the 52% who measure whether it was right is the defining operational gap of this phase of AI adoption. It is not a technical gap. It is a management one.

You would never let a new sales hire talk to customers for a year with no call reviews, no ramp scorecard, and no manager checking their work. You have almost certainly done exactly that with your AI.

Build the answer key. Score the work. Tier the trust to the stakes. Then put the number on the board next to your pipeline coverage, and manage it like everything else that determines whether you hit the year.

Share this article:
Copied!
S

Sarah Mitchell

Chief Marketing Officer

Sarah is a veteran B2B marketer with over 15 years of experience helping SaaS companies scale their marketing operations.

View all articles

Newsletter

Get the latest business insights delivered to your inbox.