AI RopewayAI GTM Engineering
All posts

Sales Automation & RevOps

Why AI SDR Pilots Fail in the First 60 Days

Bharat Gulati·
Why AI SDR Pilots Fail in the First 60 Days

Almost every AI SDR pilot I have watched fail did the same thing in week one. Somebody switched it on, pointed it at a list they already had, and waited.

Nobody owned it. Nobody read the replies. Sixty days later the verdict was "AI outbound doesn't work for us", which is a bit like concluding cars do not work after leaving one in the drive.

Here is the short version. AI SDR pilots fail for reasons that have almost nothing to do with the AI. They fail because domains were not warmed, because the list was firmographic rather than signal-based, because nobody was accountable for replies, and because 60 days is often shorter than the time it takes for the system to be worth judging.

How bad is the failure rate, really

Honestly, nobody knows, and you should be suspicious of anyone who says they do.

The most-quoted figure is 50% to 70% churn on managed AI SDR contracts. That number traces back to UserGems, a vendor in the category, and is repeated in a failure-forensics piece by Leadgen Economy which then argues most of that churn lands inside the first contract cycle, based on operator post-mortems on review sites and Reddit rather than a measured dataset.

So: high churn is real and widely reported. The precise shape of it is an estimate built on other people's estimates. I use it as directional evidence that this category has an onboarding problem, not as a statistic.

Cause 1: the domains were not ready

This is the most common one and the most mechanical.

Cold sending from a fresh domain at volume is the fastest way to a spam folder. Google asks bulk senders to keep spam complaint rates below 0.10% in Postmaster Tools and warns that 0.30% or higher increases spam classification. Senders above 5,000 messages a day need SPF and DKIM together, DMARC, and one-click unsubscribe, under Google's sender requirements in force since February 2024.

A proper warm-up is four to six weeks. If your 60-day pilot spends the first four weeks warming, you are judging the system on 30 days of real sending, and the first two of those are usually spent finding out the list is wrong.

There is a second wrinkle specific to AI-written mail. In a study of 100,000 paired cold emails run from October 2025 to April 2026, AI-generated emails were spam-flagged at 8% against 3% for human-written ones, according to Digital Applied's analysis. It is a vendor's own dataset, but the methodology is published and the deliverability figures come from Gmail Postmaster Tools and Microsoft SNDS. Take the direction seriously even if you discount the decimal places.

Cause 2: the list was a filter, not a signal

Most failed pilots are pointed at a list built from job title, headcount and industry. That describes who might have the problem. It says nothing about who has it this month.

Signal-based lists perform differently because the timing is different, not because the copy is better. A company that just posted three SDR roles, or lost a VP Sales, or announced a funding round, is in a different state to one that merely matches your ICP. We set out how to build those lists in the signal-based outbound guide.

If a pilot is going to fail on one thing, it fails on this one.

Cause 3: nobody owned the replies

An AI SDR generates replies. Replies need a human within a few hours, and in most failed pilots that human does not exist.

What actually happens is that replies queue up in a shared inbox, get triaged twice a week, and the interested ones go cold. The system did its job. The organisation did not.

Before a pilot starts, name the person, agree the response window, and put the reply volume in their calendar. If nobody has capacity, the honest answer is to delay the pilot rather than run one you cannot service.

Cause 4: the pilot measured the wrong thing

Meetings booked in month one is a poor success criterion for a system whose first month is infrastructure.

Better criteria for a 60-day pilot: deliverability holding (spam rate under 0.10%, bounce under 3%), reply rate trending, positive-reply quality, and time-to-first-response on inbound replies. Meetings are the output of those four. Judge the inputs first, then the output in month three.

And be careful with cost per meeting inside a pilot window. Front-loaded setup makes month one look terrible and month three look artificially good. The calculation that survives that distortion is in cost per booked meeting.

Here is what most people get wrong

They run a pilot to evaluate a tool. The tool is rarely the variable.

Give two companies the same AI SDR, the same budget and the same 60 days, and the one with clean CRM data, a signal-based list and a named human on replies will get a completely different result. The pilot did not test the software. It tested whether the surrounding process exists.

Which is why the question worth asking before you start is not "which AI SDR should we pilot" but "if this works, who runs it in month four, and on what data". If your CRM is telling you a win rate that is not true, every decision downstream of the pilot is built on sand, a problem we picked apart in fractional RevOps: when the problem is your data, not your reps.

What a 60-day pilot should actually look like

  • Weeks 1 to 4: secondary domains registered and warmed, SPF, DKIM and DMARC verified, list built from signals rather than filters, reply owner named with a response window agreed
  • Weeks 5 to 6: low volume live sending, watch spam complaint rate and bounce rate daily, iterate copy on reply data rather than opinion
  • Weeks 7 to 8: scale volume only if deliverability held, review positive-reply quality, decide on month three rather than declaring victory or defeat

Anything that compresses this is not a faster pilot. It is a shorter one.

FAQ

How long should an AI SDR pilot run?

Ninety days is a fair test. Sixty is workable if domains are already warmed. Thirty tests your warm-up schedule and nothing else.

What percentage of AI SDR pilots fail?

Widely quoted figures put churn on managed AI SDR contracts at 50% to 70%, originating with a vendor's own reporting and repeated across the category. The direction is well evidenced; the precision is not.

What is the first thing to check when a pilot underperforms?

Deliverability, before copy. Check spam complaint rate and bounce rate in Postmaster Tools. A reply rate collapse is far more often an inbox placement problem than a writing problem.

Do AI-written emails perform worse than human-written ones?

On the published paired-email data, yes but not catastrophically: 4.1% reply rate against 5.2%, and an 8% spam-flag rate against 3%. The gap has narrowed year on year, and targeting still moves results more than authorship does.

Should I pilot an AI SDR or hire a person?

Different failure modes, different costs. We compared the three realistic options in fractional VP Sales vs AI SDR vs agency, and the first-year cost of the human option in what a £45k SDR actually costs.

Can I run a pilot from my main domain?

No. Use secondary domains. A spam reputation problem on your primary domain reaches your invoices and support mail, and it takes months to repair.

When is an AI SDR the wrong answer entirely?

More often than the category admits. When not to use an AI SDR covers the cases where the honest recommendation is to not buy one.

Where to start

Before you shortlist a vendor, write down who owns replies, what your list signal is, and how long your domains have been warm. If any of those three is blank, fix it first. The pilot will fail on that gap regardless of which system you buy.

If you want the gaps found before you spend the budget, book a GTM audit.

Ready to put AI to work?

Book a free AI audit — get a custom roadmap in 48 hours.

Claim free audit

From the blog

AI insights & playbooks

Claim Your Free AI Audit