Home › Guides › Does AI personalisation actually work in outbound?
Analysis

Does AI personalisation actually work in outbound?

The best public dataset on this says AI-written messages lose to human-written ones. Our controlled tests say the same thing about text, and something very different about where AI actually pays.

Paul CassidyBy Paul Cassidy, founder of Prospectio.ai. Ex Google, Salesforce and Twilio sales leader. Updated 10 September 2026.
In short: the evidence says AI-written outreach copy does not beat a good human template, and we agree. It also says AI applied to targeting, media and response speed produces some of the largest effects ever measured in outbound. Most tools point their AI at the one application that does not work.

The finding that started this

In August 2026 Expandi published The State of LinkedIn Outreach H2 2026, built on 13.2 million connection requests and 6.7 million messages. It is the largest public dataset in this category and it is worth reading.

Their most quoted finding is that AI personalisation does not work. Across 118 campaigns from 73 accounts, comparing each account’s AI-generated campaigns against that same account’s human-written ones, AI copy accepted about 12 percent worse. Reply rate was flat.

The within-account design is the part that makes it credible. A naive comparison showed AI campaigns beating non-AI campaigns, but that was selection bias: the accounts adopting AI were stronger operators to begin with. Comparing each account against itself removed that, and the effect reversed. That is careful work and they were right to publish it.

We sell an AI outreach product. The finding is inconvenient. It is also, as far as it goes, consistent with our own data.

Why we think they are right about text

Our dataset is different in shape: 389,890 prospects and 15,018 meetings across 41 client programmes, run as fifteen controlled tests with prospects randomised inside each account. Smaller than theirs, and experimental rather than observational.

We never ran AI-written against human-written copy as a controlled test, so we cannot confirm their number directly. What we can say is that everything we did test about copy points the same way.

TestLosing variantWinning variant
Message structureLonger message with solution detail, 18.6%2.5 sentences, question-led, no pitch, 44.5%
Copy focusCompany-led, 10.8%Question-led, 33.2%
Message lengthOver 700 characters, 11.2%240 to 420 characters, 38.6%

Every one of those rewards taking words out. Generated personalisation adds them. It produces an opening line about the prospect’s recent post, their company’s funding round and their role, and each of those clauses is a sentence the reader has to get through before reaching the question that would have made them reply.

There is a second mechanism that the length data does not capture. A pattern the recipient has seen before reads as a pattern, however many variables were filled in. Generated first lines converged on a recognisable shape some time in 2025, and once a reader can spot the shape, the personalisation stops being evidence that a human was involved. Which was the only thing it was ever doing.

Where the lift actually is

Here is where we part company with the conclusion most people are drawing from their report, which is that AI has no role in outbound.

The three largest effects in our dataset have nothing to do with writing.

Deciding who to contact was worth 4x. Across 6,870 enterprise prospects, an identical sequence produced a 13.0 percent reply rate on a list built from job titles and 51.9 percent on a list where subject-matter fit had been verified per person. No copy change we have ever measured comes close. Reading a profile and judging whether a specific problem is actually that person’s remit is slow, boring, and exactly the kind of judgement a model can now make at volume. The full test is here.

Media was worth 40 percent. A personalised video first touch replied at 40.4 percent against 28.8 percent for text-only across 8,385 prospects, and produced 57 percent more meetings. It held across three accounts and showed no decay over eleven weeks, which is unusual. A voice note at step two lifted meetings from 4.1 to 5.0 percent for a fraction of the production cost. Video is expensive to fake at scale, and that is precisely why it still signals effort in a way that generated text no longer can.

Speed was worth 6x. Replies answered inside 24 hours converted to a meeting at 21.4 percent. Replies answered after 48 hours converted at 3.6 percent. That is the largest single effect in our dataset and it is an operational problem, not a writing one.

Four possible applications of AI in outbound, then. Selection, media, speed, and prose. The evidence says the first three carry large effects and the fourth is at best neutral. Most of the category has pointed its AI at the fourth.

What we changed because of this

We are not neutral here, so it is worth saying plainly what this means for how we build.

Prospectio scores every prospect against the stated ICP before a message is sent, and skips the ones that do not fit. That is the 4x. It personalises one recorded video or voice note per prospect rather than generating novel prose for each. That is the 40 percent. Its auto-responder exists to close the 24-hour window. That is the 6x.

What it does not do is write four sentences of generated observation about someone’s recent activity and call it personalisation. On the evidence, that is the one thing in outbound that AI makes worse.

A note on comparing the two datasets

Do not put our reply rates next to Expandi’s. They are not comparable and treating them as though they were would be the sort of thing that makes benchmark reports useless.

Expandi reports reply rate per outbound message, drawn from platform telemetry across every account using their tool. We report reply rate across a full sequence and meeting rate measured to a calendar invite the prospect accepted, drawn from campaigns we ran for 17 clients. Different denominators, different populations, different measurement.

They are also different kinds of evidence. Theirs is observational at enormous scale, which is excellent for describing what is happening across a platform and, as they state in their own methodology, cannot establish cause. Ours is experimental at much smaller scale, which can isolate cause for the specific things we tested and tells you nothing about anything we did not test.

Neither is better. They answer different questions, and the only honest way to read them together is direction by direction.

What to do on Monday

If you take three things from all of this:

Stop paying for AI that writes your first line. Test your own template against it inside your own account, the way Expandi did, and use whichever wins.

Point the effort at the list instead. If you cannot say in one sentence why a specific problem belongs to a specific person, they are a lookalike, not a prospect.

Put someone’s name against the reply inbox, with a rule that nothing waits more than a day. It is free, it needs no software, and on our numbers it is worth more than every copy decision on this page combined.

Frequently asked questions

Does AI personalisation improve cold outreach reply rates?

Not when it is used to write the message. Expandi's within-account analysis of 118 campaigns found AI-hyperpersonalised copy accepted about 12 percent worse than the same accounts' own human-written templates, with reply rate flat. Our controlled tests point the same way: the message that won our copy test was shorter and said less, which is the opposite of what generated personalisation produces.

So is AI useless in outbound?

No, but the lift is not in the writing. In our controlled tests the largest effects came from AI applied to selection rather than prose: verifying subject-matter fit before adding someone to a list took reply rate from 13.0 to 51.9 percent on an identical sequence. Personalised video in the first touch lifted replies from 28.8 to 40.4 percent and meetings from 4.1 to 6.4 percent. Answering an inbound reply inside 24 hours converted at 21.4 percent against 3.6 percent after 48 hours.

Why does AI-written personalisation underperform?

Two reasons show up in the data. Generated personalisation adds words, and every copy test we have run rewards taking words out: a message of two and a half sentences with no pitch replied at 44.5 percent against 18.6 percent for the same message with solution detail added. And a pattern that a recipient has seen before reads as a pattern regardless of how many variables were filled in.

What should you use AI for in outbound then?

Deciding who to contact, producing media at scale, and responding fast. Those are the three places our tests found large effects. Writing the first sentence is where most tools point their AI and it is the weakest of the four applications.

Can you compare Expandi's reply rates to Prospectio's?

No, and anyone who does is misleading you. Expandi reports reply rate per outbound message across platform telemetry; we report reply rate across a sequence and meeting rate measured to an accepted calendar invite. The denominators are different, so the absolute numbers are not comparable. Only the direction of each test is.

See it on your own pipeline

Book a 20-minute demo or start a free trial. No card needed.