The uncomfortable finding first
In August 2026 Expandi published a comparison of AI-generated outreach against human-written templates, drawn from 118 campaigns across 73 accounts. Comparing each account against itself, human copy came out about 12 percent ahead on acceptance and reply rate was flat.
We sell an AI outreach product. That result is inconvenient and, as far as it goes, consistent with our own data.
We never ran AI copy against human copy as a controlled test, so we cannot confirm their number. What we can say is that everything we did test about copy points the same direction. A longer message with solution detail replied at 18.6 percent; the same message cut to two and a half sentences with no pitch replied at 44.5 percent. Company-led copy replied at 10.8 percent, question-led at 33.2 percent. Messages over 700 characters replied at 11.2 percent against 38.6 percent in the 240 to 420 band.
Every one of those rewards taking words out. Generated personalisation adds them. Ask a model for a personalised opener and you get the prospect’s recent post, their funding round and their job title, which is three clauses the reader has to survive before reaching the question that would have made them reply.
There is a second mechanism the length data misses. A pattern the recipient has seen before reads as a pattern regardless of how many variables were filled in. Generated first lines converged on a recognisable shape some time in 2025, and once a reader can spot the shape, the personalisation has stopped doing the only job it had.
Where the lift actually is
Four things you could point AI at in outbound: who you contact, what media you send, how fast you respond, and what words you use. The evidence says the first three carry large effects and the fourth is roughly neutral.
Deciding who to contact. 4x. Across 6,870 enterprise prospects, an identical sequence produced 13.0 percent replies on a list built from job titles and 51.9 percent on a list where subject-matter fit had been verified per person. No copy change we have ever measured comes close. Reading a profile and judging whether a specific problem genuinely sits in that person’s remit is slow, repetitive work that a model does well and at volume. The full test is here.
Media. Plus 40 percent. A personalised video first touch replied at 40.4 percent against 28.8 percent for text-only across 8,385 prospects, and produced 57 percent more meetings. It held across three accounts and showed no decay over eleven weeks, which is unusual: most format effects in outbound fade inside a month. Video is expensive to fake at scale, which is exactly why it still signals effort when generated text no longer can. A voice note at step two lifted meetings from 4.1 to 5.0 percent for a fraction of the production cost.
Response speed. 6x. Replies answered inside 24 hours converted to a meeting at 21.4 percent. Replies answered after 48 hours converted at 3.6 percent. Largest single effect in the dataset, and it is an operations problem rather than a writing one. An agent that reads the inbox each morning and drafts responses closes precisely that gap.
Writing the message. Roughly nothing. Which is where most of the category has pointed its AI.
What good use looks like in practice
Set the constraints rather than asking for quality. “Write a good LinkedIn message” produces a bad one. “240 to 420 characters, two and a half sentences, ends in a question, names their own product, says nothing about mine” produces the one that replied at 44.5 percent. The constraints are the value, and they come from testing rather than taste.
Use it before the message exists. If you cannot state in one sentence why a specific problem is a specific person’s responsibility, they are a lookalike rather than a prospect. Getting a model to make that case for every name on a list, and to say “no case” when it cannot, is worth more than every copy prompt combined.
Keep a human on the send button. Not forever, but until you have read enough output to know the constraints are holding. The failure mode is quiet: the drafts look fine and drift longer and more promotional across a batch.
Let it handle the inbox. Reading, classifying and drafting replies is the highest-return automation available and it needs no creativity at all.
Can Claude touch LinkedIn directly?
No, and you should be wary of anything claiming otherwise. LinkedIn publishes no outreach API and its user agreement prohibits using “bots or other unauthorized automated methods” to access the service, send messages or drive engagement.
What works is one step removed. An outreach platform holds the LinkedIn session and enforces its own limits, and Claude drives the platform through an MCP server. Claude never touches LinkedIn. If you want that running, the setup takes about five minutes and gives Claude 33 tools covering campaigns, prospect imports, sequences and the inbox.
Worth saying plainly, because a lot of pages in this category will not: routing through a platform reduces detection risk relative to a browser extension. It does not make automated outreach authorised. Anybody telling you their tool is LinkedIn-approved is not being straight with you.
The short version
Point AI at the list, the media and the response time. Write the message yourself, or make the model write it under constraints tight enough that it cannot do the thing it wants to do.
All fifteen tests, sample sizes and method are on the 2026 outbound benchmarks page. Free, ungated, and you are welcome to cite it.
