How to test marketing ideas when your sales cycle takes ninety days

The four station loop: bet, build, expose, kill or scale
Revenue in a long cycle B2B business arrives 60 to 180 days after the decision that produced it, so any test judged on revenue at 30 days looks like a failure. This guide sets out the ladder of signals to judge tests on instead, the three rules that stop the goalposts moving, what belongs on a one page test card, and what it looked like on a live industrial programme.

Judge the test on the earliest signal you have proven predicts revenue in your business, and name that signal before the test goes live. On a ninety day cycle, revenue arrives 60 to 180 days after the decision that produced it, so anything reviewed on revenue at 30 days reads as a failure whether it worked or not. The discipline sits in what you agree to measure, and when, in advance.

Why does testing break when the cycle is long?

Because the answer arrives after the decision has to be made.

Almost every published guide to marketing experimentation is written for businesses where the conversion happens in the same session as the click. Sample sizes in the thousands, results by Friday, a significance calculator. Take that model into a business selling a £180,000 system to a buying group over four months and it comes apart on contact. You have forty enquiries a month, not forty thousand. The deal a January campaign started closes in May.

Two things happen next, and both are expensive.

The first is that the team stops testing. "Our cycle is too long to test properly" becomes received wisdom. The plan gets set once a year, and the only thing anybody learns is whether the year went well.

The second is worse. The team does test, then judges the test on revenue at 30 days, which is the one number that could not possibly have moved yet. Everything looks like a failure. Good work gets stopped at exactly the point it was starting to compound, and because no record was kept of what was stopped or why, the same idea comes back in eighteen months and gets stopped again.

In both cases the fault sits in the review date rather than the cycle length.

What do you judge a test on, if not revenue?

On the earliest signal you have proven predicts revenue in your business. And you decide which signal that is before the test goes live, not after.

That gives you a proxy ladder. Four rungs, read at four different points, each doing a different job.

Tier What you are reading Read after What it is for
1 Attention: click through, scroll depth, video completion, time on page Same day Obvious failures only
2 Intent: form starts and completions, gated downloads, pricing page visits, demo requests, outbound reply rate 3 to 7 days The primary kill layer
3 Qualification: ICP fit, meetings booked, meetings held, sales accepted leads 14 to 30 days Confirms you attracted the right people, not merely more people
4 Revenue: pipeline created, opportunities, closed won 60 to 180 days Recalibration only, never a kill criterion

Tier 2 does most of the work. Intent signals arrive within a week, there are enough of them to read even in a low volume business, and they move when a message lands. Tier 1 only catches the obvious: an advert nobody clicks, a page nobody scrolls past the first screen. Tier 3 is the honesty check, because it is where you find out whether you attracted the right people or simply more of them.

Tier 4 is where the discipline is hardest to hold. Revenue is the number that matters and it is the number you must not kill on, because by the time it reads, the decision window shut months ago.

Revenue has a different job. Use it to check whether the tier 2 signal you killed on actually predicted anything. That is what turns a borrowed method into your own data. After roughly fifteen closed tests, the log starts telling you which of your early signals track revenue in your business and which are noise. Nobody else has that, it cannot be bought, and it takes about a year to build.

What are the three rules that keep it honest?

Lock the kill tier before the data arrives. Write down, in advance, which tier you will judge on and at what level. A threshold moved after you have seen the numbers is a negotiation with yourself, and you will win it every time. This one rule does more work than the rest of the method combined.

Log the 30 day read even on tests you already killed. A test you stopped on day five still produced leads, and those leads still generate qualification data three weeks later. Record it. That is the part that compounds, and it is the part everybody skips, because the test is over and the next one is already live.

Carry an anti metric on every test. A guardrail number that must not degrade while you chase the one you are trying to improve. A variant that doubles form fills and halves lead quality is a loss wearing a medal, and without a guardrail it gets rolled out across every form you own before anyone notices.

What goes on a test card?

One page, written before any asset exists. Six fields.

Hypothesis. Written as: if we do X for audience Y, then Z will move, because. The "because" is what makes the test worth running whether it wins or loses, since it is the part you learn from either way.

Kill tier and threshold. Which rung of the ladder, what number, read on what day.

Guardrail. The anti metric, and the level at which it trips.

Budget cap. A figure that cannot be exceeded without a new card.

Decision date. An actual date, in a diary, before launch.

Decision owner. One named person. A committee cannot kill anything, which is why tests owned by committees run forever.

If the card does not fit on a page, you are not describing a test. You are describing a project, and projects get judged annually.

Two things about the card matter more than they look. It is written before the assets, so the creative work gets built to answer a question rather than the question being reverse engineered from creative that already exists. And "let it run a bit longer" is not one of the six fields, which means it is not available on the decision date.

How do you find tests worth running?

Take one problem and turn it into twelve ideas across the whole funnel. Twelve ideas, not twelve variations of a headline.

Twelve is a deliberate number. Ask a team for three ideas and you get the three safest, then spend the meeting defending your favourite. Ask for twelve and by about the eighth the safe ones have run out and the interesting ones start arriving.

Then apply a filter that sounds unserious and is not. If none of the twelve would make a sensible, competent person mildly uncomfortable, reject the round and go again. A list everybody approves of instantly is a list of things you already know the answer to, and testing those costs you the slot.

Run this weekly rather than quarterly. A quarterly cadence gives you four attempts a year, which is nowhere near enough repetition for the log to start earning its keep.

What does this look like on a live programme?

HMS Networks sells industrial communication hardware to engineers across the Anybus, Ewon, Red Lion and N-Tron brands. Long cycles, technical buyers, a buying group in which several people will never fill in a form.

Two lead numbers ran alongside each other rather than being blended into one: 1,588 sales ready leads from the website, and 201 confirmed leads into CRM, made up of 157 from the HMS own brands and 44 from other divisions. The website number reads within days. The CRM number reads weeks later. Keeping both, and always knowing which one you were looking at, is what made early decisions possible without anyone pretending the early number was the real one.

Measurement ran on infrastructure HMS owns rather than a rented platform. That matters more for testing than it does for reporting. A test register is only worth something once it accumulates, and a register held inside an advertising account disappears the day the contract ends.

The commercial result was a 35:1 return on media, with cost per qualified lead falling from around £760 through trade events to around £72.

Where to start this quarter

Three things, none of which need a purchase.

Write one test card for something already running. Pick a campaign nobody has questioned in six months and fill in the six fields retrospectively. Most teams find they cannot name the decision owner or the threshold. That is the finding.

Pick your tier 2 signal and time it. Take your last twenty closed won deals and find the earliest recorded action that shows up in most of them. That is your candidate kill signal, and now you have a reason to trust it rather than a preference.

Start the log. A spreadsheet, one row per test: hypothesis, kill tier, decision date, outcome, 30 day read. Fifteen rows from now it will be the most useful marketing document in the business.

If you want the model applied to your own numbers, our B2B marketing strategy work starts by finding which of your early signals actually predict revenue, and our conversion and attribution intelligence work puts the measurement somewhere you own it.

Frequently asked questions

How do you A/B test marketing with a long B2B sales cycle? By testing against leading signals rather than revenue. Decide before launch which signal you will judge on, usually an intent signal read at three to seven days or a qualification signal read at 14 to 30 days, and set the threshold in writing. Revenue at 60 to 180 days is used afterwards to check whether that early signal predicted anything.

What is a leading indicator in marketing? A leading indicator is an early measurement that reliably precedes the outcome you care about. In B2B that usually means form starts, gated downloads, pricing page visits, demo requests, meetings booked or sales accepted leads. It becomes genuinely useful once you have checked, across your own closed deals, that the signal actually shows up before revenue in your business rather than in someone else's.

How long should a B2B marketing test run before you decide? Long enough to read the tier you nominated, and no longer. Attention signals read the same day, intent signals at three to seven days, qualification signals at 14 to 30 days. The decision date belongs in the diary before launch. Extending a test after you have seen the numbers is how a test quietly turns into a permanent campaign.

What should a marketing test card contain? Six fields on one page, written before any asset is made: the hypothesis with its reasoning, the kill tier and threshold, the guardrail metric that must not degrade, a budget cap, a decision date, and a single named decision owner. If it needs more than a page, it is a project rather than a test, and it will be reviewed annually rather than weekly.

How many marketing tests should you run at once? Enough that no single test carries the quarter, and few enough that you can read them apart. Turning one problem into twelve ideas across the funnel and running them in a weekly rhythm gives most B2B teams sufficient repetition to learn from. Four attempts a year, which is what a quarterly cadence produces, is not enough for any pattern to appear.

Need a more tailored conversation with our team?

Get In Touch
Contact Us