
Core Web Vitals explained: what LCP, INP, and CLS measure, Google's official thresholds, how to run a Core...
Download our own internal B2B Playbook here

A/B testing means showing two versions of the same page, email or form to two randomly assigned groups of visitors, then keeping the version that converts better. It sounds simple. In practice, many A/B tests crown "winners" that vanish the moment they ship, because the test lacked traffic, time or a real hypothesis. This guide walks you through the full method, from the basic math to the tools, so every test teaches you something you can use.
A/B testing (also called split testing or bucket testing) is a controlled experiment. You keep the current version of an element, version A or the "control," and build a version B that differs in one specific way. A testing tool randomly assigns visitors to each version, and you compare a single metric: demo request rate, click-through rate, sign-up rate.
Randomization is the whole point: because the two groups differ only by chance, any gap in results can be attributed to the change you made. That is what separates an A/B test from a before-and-after comparison, which is skewed by seasonality, running campaigns or news in your industry.
You can test almost anything you can measure: headlines, value propositions, forms, pricing pages, email subject lines, Google Ads copy, meeting booking flows. The method also applies to apps and to entire conversion funnels.
These terms get mixed up constantly. Here is how they compare.
In B2B, every lead is worth a lot and traffic is usually limited. A/B testing exists to make decisions with evidence instead of the highest-paid person's opinion. It also de-risks change: rather than rolling out a new pricing page to everyone, you expose it to half your traffic and measure.
It builds lasting customer insight too. A test showing that your visitors respond better to industry-specific proof than to a generic promise tells you something about your market, not just about one page. That insight carries over to your ads, emails and sales talk tracks.
Finally, A/B testing is the engine of any conversion rate optimization program. Without testing, optimization is just opinions.
This is the A/B testing framework we use on B2B websites. Each step determines how reliable the next one is.
The hypothesis is the step most teams rush, and it is the step that turns a test into a learning. A hypothesis with no observation behind it ("let's try a green button") will never tell you why a version wins.

"Session recordings show that visitors on the demo page scroll down to the customer logos before filling out the form. We believe placing three logos and a customer quote next to the form will increase the demo request rate without lowering the share of sales-qualified leads."
This is the most technical part, and the most often skipped. Three concepts prevent most mistakes.
The confidence level (typically 95%) controls the risk of declaring a winner when there is no real difference. Statistical power (typically 80%) is the test's ability to detect a real difference when one exists. Set both before the test starts, never after you have seen the data.
A standard rule of thumb for 95% confidence and 80% power gives, per variation: n ≈ 16 × p × (1 - p) / d², where p is your current conversion rate and d is the absolute difference you want to detect.
Worked example: with a 3% conversion rate and a 0.6-point difference to detect (3% to 3.6%), you need about 12,900 visitors per variation, or nearly 26,000 visitors on the tested page. Most testing tools include a calculator, but knowing the order of magnitude keeps you from launching tests that can never reach significance.
Run every test for at least two full weeks, always in whole weeks, to cover weekday and weekend behavior. In B2B, match your buying cycle: if prospects visit several times before requesting a demo, two to four weeks is a reasonable minimum. Avoid overlapping with a trade show, a launch or an unusual campaign.
With limited traffic, you cannot test everything. Prioritize pages close to conversion and changes bold enough to produce a measurable difference. A simple ICE score (impact, confidence, ease, each rated 1 to 10) is enough to rank ideas.
If a page has real usability or speed problems, fix them before testing. No test makes up for a form that breaks on mobile. A UX design review upfront often surfaces the best hypotheses.
On the planned end date, check test quality first: balanced split, no bugs, target sample reached. Only then look at the primary metric and its confidence interval.

There are three possible outcomes:
In B2B, always check the quality of leads each version produces, not just the volume. A shorter form can generate more requests and fewer qualified meetings. Connect your testing tool to your CRM (HubSpot, Salesforce) to follow leads through to opportunity, which requires clean marketing analytics and tracking.
Keep a test log: hypothesis, screenshots, dates, sample sizes, outcome, decision. After a few months, that log becomes one of your most valuable marketing assets.
Google Optimize was sunset in September 2023, which pushed many teams to switch tools. Here are the main categories.
The right tool depends on your traffic, engineering resources and CMS. Also check the script's impact on page speed (flicker) and how it works with your consent banner: a test that slows the page down skews its own results.
Yes, when the basics are respected: a clear hypothesis, enough traffic, a fixed duration and a pre-defined analysis. It fails when teams stop tests early, test trivial changes on small samples or never connect results to revenue.
No. What has faded is low-value testing of cosmetic details. Privacy changes and AI-generated variations have changed the tooling, but randomized experiments remain the most reliable way to measure the effect of a change.
Marketing teams, product teams, UX designers, growth teams and ecommerce managers. In B2B, it is common on demo pages, pricing pages, email nurture sequences and paid search landing pages.
It is the repeatable process behind your tests: research, hypothesis, prioritization (for example ICE), sample size calculation, QA, launch, analysis and documentation. The 7 steps in this guide are a complete framework you can adopt as is.
Pick one primary metric tied to the goal (conversion rate, demo request rate, click-through rate) and a few guardrails such as bounce rate, average order value or lead-to-opportunity rate, so a "win" does not hide a loss elsewhere.
A/B testing only pays off when it is part of a program: research, prioritized hypotheses, properly sized tests and a shared log. That is what we build for B2B companies through our conversion rate optimization services, with senior experts and in-house AI agents to speed up analysis.
Not sure which tests to run first? Get your free action plan: we will pinpoint the pages and hypotheses with the most upside.

Core Web Vitals explained: what LCP, INP, and CLS measure, Google's official thresholds, how to run a Core...

RevOps & Automation
Lead scoring explained: fit vs. engagement, the main lead scoring models, a 7-step build process, a sample...

Nine B2B lead generation strategies compared on speed, cost and lead quality: SEO, lead magnets, cold...

Get your free action plan in 48 hours, with no commitment on your side.
Ask Now