For six weeks, two versions have been running on the page. In the first, the headline promises a clear outcome; in the second, it explains the service in more detail. The table shows that version B has two more enquiries. Someone is already suggesting ending the test and rewriting the rest of the site to follow its example.
But tomorrow one enquiry may arrive in the other version - and the “win” will disappear. Not because the team broke something. With a small number of enquiries, data can easily fluctuate from one case to the next. If you call such a fluctuation proof, the process starts running itself like a broken traffic light: everyone watches the signal, but that does not make it safe to drive.
An A/B test is a comparison of two versions of a page or element that different visitors see at random. It tests one assumed cause of an action, rather than helping you guess which headline looks prettier.
1. Low traffic is not a reason to do nothing
When there are few visitors and enquiries, the worst extremes look alike. The first is not touching the page for months because “there is too little data for a test anyway.” The second is changing everything at once and believing every new number. In both cases, you will not learn what helped or got in the way.
Moving from a visit to an enquiry is not only about buttons. Bring together the person’s path, their queries, one important problem and consistent checks into one simple rule: launching a test is not a decision by itself.
Start with an ordinary sentence: “On this page, I want the person to do this.” In B2B, this may be more than a submitted form. A target action is a specific step that genuinely brings a work conversation closer: a request for a quote, booking a consultation, downloading a technical description after leaving contact details. If the team argues about what counts as success, the test no longer has one honest question.
From an action in a report to an action for the business
It is important not to confuse an action that is convenient for reporting with one that is useful for the business. A click on a phone number can be a valuable observation, but it does not confirm a conversation. Viewing a pricing page may indicate interest, but it is not the same as a request. Define the main action so that after the test you can explain why it matters. Then the team will not argue about the number that simply appeared in the table first.
Suppose you sell services to manufacturing companies. If people who are looking for a supplier come to the page, but the form sends an event after they click the button rather than after it is submitted successfully, you are measuring an intention to click, not a real enquiry. The headline has nothing to do with it: first fix the measurement point.
Here is another situation: a page about a complex service receives few enquiries, but some people move on to a requirements calculator. Such an early signal - an action that logically comes before the main enquiry - can be watched separately as a supporting clue. But do not call it the page’s victory or replace enquiries with it. It only shows where it is worth looking more closely.
2. First check what actually reaches the report
A conversion is a recorded target action by a visitor, such as a successfully submitted form. It exists in a report only when the site has sent the event correctly. So before drawing a conclusion about the page, check the chain itself: the person completed the action, the site recorded it, and the measurement system received exactly the event you expected.
Measurement Protocol is a way to send an event to Google Analytics from a website or server. Before going live, check the implementation and the event keys so that you do not get a beautiful but false chart.
In practice, this does not mean “see whether there is a number.” Take one real test enquiry and trace its path. Did the screen change after successful submission? Did the required event fire rather than a button click? Is it sent twice after the page is refreshed? Do its name and parameters match what the report shows? If you do not have a clear answer to these questions, it is too early to compare versions.
Compare not only events, but also the path
It is also useful to check whether the person’s path itself changed between versions. For example, one version opens the form on the page, while the other leads to a separate screen. Make sure successful completion is recorded the same way in both scenarios, otherwise the test will show a difference not in how convincing the offer is, but in where an event went missing. Measurement is not a technical detail: it gives you the right to draw a conclusion. On the analytics service, you can see why a decision starts with what is actually recorded. And the conversion tracking recovery case shows that a dashboard full of numbers will not become a reliable foundation until target actions are collected correctly.
Also look at who comes to the page. If one version is more often seen by people from a price-focused ad, while the other is seen by people looking for a service description, you are not comparing only pages. You are comparing different expectations. In a test, both branches should receive visitors at random; otherwise, the difference may result from the traffic source rather than a change on the page.
This is not a reason to give up advertising or postpone the page until an “ideal” time. Before drawing a conclusion, simply check whether different messages, campaigns or visitor groups have been mixed into one test. If they have, do not conclude anything about the headline. First make the comparison even, or state its limitation in the result record.
3. Instead of a random edit, build one strong reason
A hypothesis is a testable assumption about why a person does not take the required step and what change may help them. Not “let’s make the button brighter,” but, for example: “The person does not leave a request because before the form they do not understand what information they will receive in response; if we briefly name the contents of the response next to it, they will be more likely to complete the submission.”
To keep this thought from becoming a guess from a meeting, reduce it to several fields in an ordinary document:
- the target action you count;
- the page and visitor group for which you noticed the problem;
- the observation: where the person stops or what they do not understand;
- one change you propose;
- the expected direction, without promising an outcome;
- a way to check that the event is recorded correctly.
These fields are useful for more than paperwork. They separate three different problems that are often mixed together. First: people do not reach the page or arrive with the wrong expectation - then look at the ad message and whether the page matches the query. Second: the person reaches it but does not understand the offer or the next step - then work on the page content. Third: the action happens but does not reach the data - then do not touch the headline until you fix the measurement.
There is also a fourth situation: the main action is so rare that the test cannot yet provide a reliable conclusion. Here, you do not need to pretend a new colour will solve the problem. Keep the main goal, calculate the required volume for each branch and consider which early step genuinely comes before it. An early signal helps you learn about the person’s path, but it does not replace the business result.
A hypothetical example: you see that visitors read the service terms but rarely move on to the form. This does not yet mean the text is poor. Perhaps there is no explanation before the form of what will happen after a request; perhaps the traffic source promises something else. Write down the exact observation you have and choose only one cause to test. If you rewrite the headline, price, form and ad at the same time, the next number will not tell you what worked.
It is useful to separate evidence from a clue. A call recording, a manager’s response or an observation of actions on the page can suggest what to test. They do not prove that the change by itself increased the number of enquiries. Evidence of a causal result requires an honest comparison in which you did not also change five other things.
4. Calculate not "until you are tired of it," but how much data you need
A sample is the number of visitors or recorded actions on which you draw a conclusion. For an A/B test, the required volume is calculated separately for each branch before launch. It depends on the current share of target actions and on the difference you want to be able to notice. There is no universal number that fits every website.
This does not mean you need to build complicated formulas yourself. Before starting, it is important to record two things: what volume is required by the selected calculation, and what you will do if it is not reached at a pace acceptable for the business. For example, you will not keep the test running forever simply because it has already started; you will return to measurement quality, the target action or an earlier signal.
The required volume is not a planned end date. It only marks the point after which the data may allow a stronger conclusion. If visitor flow is small, choosing a winner can genuinely take longer. That is not a broken test. The breakage begins when the team adjusts the rule to suit the number it likes today.
Statistical significance is an assessment of how likely the observed difference is to be random noise. It helps you avoid confusing chance with a result, but it does not provide one hundred percent certainty. So phrase your conclusion according to the strength of the data. “Version B has a promising early signal, but the sample is insufficient” is more honest and useful than “B won.” Such a record does not slow work down: it protects the next decision from an overconfident mistake.
Do not stop at the moment when a number looks pleasant, and do not continue the test merely because it looks unpleasant today. Both decisions should rest on a rule you defined before reading the result. Otherwise, the test is not testing the hypothesis, but the team’s patience.
5. When a test is appropriate - and how to leave it honestly
An A/B test makes sense when you have one specific change, a correctly recorded target action, random visitor allocation and a calculation of a sufficient sample. This does not guarantee the result you want. It does create conditions in which the answer will not resemble a weather forecast made from the window.
If the test has not collected the required data, that is also a result, but neither a defeat nor a victory. Record what exactly was tested, what data had time to be collected, whether the measurement worked and what data was missing for a strong conclusion. Then choose the next action according to the cause: refine the hypothesis, check an early step, review traffic quality or postpone the comparison itself. You do not need to decorate uncertainty - name it so that the next person on the team understands the boundary of the data and does not turn an assumption into a “proven fact.”
The first step is simple: open the current test or the page you wanted to test, and write the target action in one sentence. Then verify it in one real journey from the form to the report. Only after that does it make sense to argue about the winner.
6. FAQ
7. Glossary
- A/B-тест
- A comparison of two versions of a page or part of it, where visitors randomly see different variants. It helps you test one specific change rather than guess its effect.
- Конверсія
- A recorded target action by a visitor, such as a successfully submitted form. It is important for you to distinguish it from a simple click so that you do not make decisions based on a false signal.
- Measurement Protocol
- A way to send events from a website or server to Google Analytics. It should be checked before going live so that reports rely on correct data about the person's action.
- Вибірка
- The number of visitors or recorded actions on which a conclusion is based. For comparing variants, the required volume is determined separately for each branch.
- Статистична значущість
- An assessment of whether an observable difference in a test may be random noise. It helps you read data more carefully, but does not provide one hundred percent certainty.
