Seegnals

Outbound strategy · 25 August 2026 · 7 min read

A/B testing cold email subject lines and variants without fooling yourself

Small lists make most cold email A/B tests inconclusive. Learn what to test, how to read a variant table honestly, and when a difference is real enough to act on.

Two subject lines, one campaign, a table showing that variant B got more opens. Most cold email A/B testing ends there, with a winner declared on a difference that would evaporate if the same test ran again next week.

The problem is not the idea of testing. It is that small B2B teams send to small lists, measure the wrong outcome, and change three things at once. Under those conditions an A/B test does not tell you what works; it tells you what happened to work on eighty people last Tuesday.

You can still learn from variants. You need to test one thing, measure interested replies rather than opens, and repeat the comparison until the same answer shows up more than once. This article covers how to set the test up, how to read the variant table without kidding yourself, and which tests are worth your time.

What a variant test can and cannot tell you

An A/B variant splits a step’s recipients at random and sends each half a different version. If the two halves are drawn from the same list in the same window, any difference in outcome is down to the copy or to chance. The whole discipline is about telling those two apart.

With large lists, chance averages out and small differences become believable. With the lists a two-person team actually sends, a few hundred people at most, chance stays loud. Three interested replies against one is not a result; it is a coin that came up heads twice.

This does not make testing pointless. It changes what you are allowed to conclude. You are looking for differences that are large, that repeat across campaigns, and that make sense given what you changed. Anything else is a hunch to test again.

Set the test up so it can be read

Three rules make a variant table interpretable.

Change one thing. If variant B has a different subject line and a shorter body and a different ask, you cannot attribute the outcome to any of them. Pick the variable. If you must move fast, run one test on the subject and a separate test on the body in different steps.

Same segment, same window. Both variants must go to people from the same list, in the same campaign, during the same sending schedule. Sending variant A this week and variant B next week is not a test; it is two campaigns.

Write the hypothesis first. One sentence: “A subject line that names their industry will produce more interested replies than one that names our product.” A hypothesis tells you what to look at and stops you from finding a story in whatever the numbers turn out to be.

In the Seegnals sequence editor, each step can carry A/B variants of subject and body. Recipients are split between them, and the campaign Stats tab shows a variant comparison table for that step. The mechanics are the easy part. The discipline is in what you change.

Sequence editor with a step opened to show its A and B variants One step, two variants. Everything outside the variant box (delays, segment, schedule) is shared, which is what makes the comparison fair.

Prioritise replies over open rates

Open rate is the default metric for subject line tests because subject lines are supposed to affect opens. There are two problems with it.

First, the measurement is unreliable. An open is a tracking pixel loading. Apple Mail Privacy Protection loads it on behalf of the reader whether they looked or not, and many corporate mail clients block images entirely. Seegnals shows open rates with a note about that inflation, and tracks opens only through a custom tracking domain you own. There is a full explanation in the piece on open rates after Apple Mail Privacy Protection.

Second, and more important, opens are not what you want. A subject line that gets opened by curious people who then close the email has not helped. The outcome that matters is an interested reply, and it is what the variant table should be sorted by.

So read the table in this order:

  1. Interested per variant. The primary measure.
  2. Opted out and Not interested per variant. The cost. A variant that wins on Interested but also produces more unsubscribe requests is not a clean win.
  3. Replied in total, as a sanity check that the classifier has seen everything.
  4. Opened, as a tie-breaker when the reply counts are identical and you need a hint about whether the subject line was noticed.

The reasons to demote reply rate as a whole are covered in reply rate is the wrong north star; the same logic applies inside a variant table.

Campaign Stats tab with per-step statistics and the variant comparison table The variant comparison table lists each outcome per variant. Read the Interested column first and treat single-digit differences as no difference.

When is a difference real?

There is no honest formula for a list of a hundred people. Instead, use three tests of judgement.

Is it large? If variant A produced interested replies from a noticeably larger share of its half than variant B did, and the halves were of similar size, pay attention. If the difference is one or two people, ignore it for now.

Does it repeat? Run the same two variants in the next campaign to a comparable segment. If the same variant wins again, you have something. If the winner flips, you have noise. Two or three repetitions is a realistic standard for a small team.

Does it make sense? A result you can explain (the winning subject named their sector; the winning body asked a smaller question) is more likely to hold than one you cannot. If you have no idea why B won, be suspicious.

Only when a variant passes all three should it become your default. Then start the next test.

Tests worth running

Subject line word-swaps are the most common test and usually the least informative. The tests below tend to produce differences large enough to see.

The ask. A call versus a two-line reply. A meeting versus permission to send a short document. Changing what you ask for is the biggest lever in the body.

The angle. Leading with a problem versus leading with an observation about the company. Leading with cost versus leading with time.

Length. A four-sentence body against an eight-sentence one, same content order.

Personalisation depth. A snippet-driven line about their company versus a plain opener. The article on personalisation with snippets and custom fields explains how to build the snippet so the test is fair.

Subject line type. Not word A versus word B, but a category shift: a question against a statement, their company name against yours, lower case against sentence case.

Whichever you choose, keep the other variables fixed and run it across at least two campaigns before you act.

Variants and conditional steps are different tools

A conditional step (the fork) sends a different message depending on what happened, for example one follow-up for people who clicked and another for people who did not. That is not a test; both branches are meant to be sent, and the two groups are different by construction. Do not compare branch outcomes as if they were variants. Use variants inside a branch if you want to test copy for that group.

Keep a testing log

Small teams lose most of the value of testing because nobody wrote the result down. Keep a plain document with one line per test: date, segment, step, what changed, hypothesis, interested replies per variant, decision. After a quarter you will have a list of things that repeatedly worked for your market, which is worth more than any single campaign’s table.

Export the campaign statistics as CSV when a campaign finishes so the raw numbers survive alongside your notes.

What to do this week

  1. Pick one running or planned campaign and choose one variable to test, preferably the ask or the angle in step one. Write the hypothesis in a sentence.
  2. Create the A and B variants so that only that variable differs. Confirm both go to the same segment in the same window.
  3. When the campaign finishes, read the variant table by Interested first, then Opted out, and decide whether the difference is large, repeatable and explainable. If it is not all three, do not act yet.
  4. Repeat the identical test in the next comparable campaign before adopting a winner.
  5. Start a testing log today, and record this test in it whatever the result.

Questions people ask

How do I A/B test cold email subject lines?

Write two subject lines for the same step, keep the body identical, and send each to a random half of the same segment in the same window. Then compare interested replies per variant.

How many emails do I need for a valid A/B test?

More than most small teams send in one campaign. Rather than chasing a statistical threshold, repeat the same comparison across several campaigns and only act when the same variant wins each time.

Should I test the subject line or the email body first?

The body, specifically the angle and the ask. Subject line changes tend to move opens, which are hard to measure reliably; body changes move replies, which are what you want.

What should I measure in a cold email A/B test?

Interested replies per variant as the primary measure, unsubscribe requests and not-interested replies as the cost side, and opens only for a rough read on whether the subject line was noticed.

Written by

Tomasz Wierzba

Writes about outbound and B2B sales. Covers sequences, follow-ups and the account-based side of cold email. Runs the numbers before recommending anything.

See it on your own list

Connect a mailbox, import a CSV and send the first campaign. Free trial, no card.