What a Mature Experimentation Program Needs Before It Picks a Platform

Published: • 8 min read

Cartoon of a lab team in white coats celebrating a 0.3% lift from an A/B test of two nearly identical Buy Now buttons, with a maturity pyramid, a p = 0.048 chart, and a trophy reading Smallest Lift, Biggest Celebration

ChatGPT

A mature experimentation program is a set of team habits for testing marketing changes: a ranked list of ideas, enough traffic to get a real answer, rules for calling a winner set before each test, and quick retirement of ideas that lose. Buy the testing tool after those habits exist.

Key Takeaways

  • An A/B test shows two versions to different visitors and measures which does better. The team's habits decide whether the answer is right.
  • Traffic is usually the limit. If a test needs months to give a reliable answer, no tool fixes that.
  • Write down what counts as a win before the test starts. Deciding after you see the numbers turns random swings into false wins.
  • Reality check: most tested ideas don't win, including the ones the team was sure about. Good programs drop losers quickly.

Why the tool comes last

An experiment, in marketing, is a controlled comparison. You show half your visitors the current version of a page, email, or offer and the other half a changed version, then measure which group buys, signs up, or clicks more. Marketers call it an A/B test. Run them regularly, with rules, and you have an experimentation program. Pretty simple, in theory.

Most CMOs meet the topic when someone on the team asks to buy a testing platform , the software that splits visitors into groups and counts the results: Optimizely or Wingify (the combined VWO and AB Tasty), for example. Google offered a free one, Google Optimize, until it shut it down in September 2023 ; Google Analytics now connects to paid tools instead.

The pitch usually comes with a vendor comparison and one question: which platform should we buy? That question starts with the tool, and the tool matters less than the team’s habits . A team with good habits gets trustworthy answers from a basic tool. A team without them gets misleading answers from the most expensive suite on the market, then believes them because the suite was expensive.

You do need a testing platform, but you pick it last. First comes the experimentation program itself: a few working habits the team builds, and no vendor sells them.

Start with a ranked list of ideas to test

A mature program keeps a running list of ideas to test (a hypothesis backlog). Each idea names the change, the reason the team expects it to work, and the number it should move. For example, the team might bet that showing shipping costs earlier will lift completed orders, because they suspect shoppers abandon carts when the cost appears at the last step. The list is ranked by how much each idea could move the business and how much work it takes.

This list is the clearest sign of a mature program. It makes the team say what it expects before it tests, which is the only way to learn something whether the test wins or loses. A team without the list tests whatever the loudest executive wants this week, and those results teach nothing, because no one wrote down what they expected.

Ask to see the list. If no one owns it, nobody is running a program yet.

Do you have enough traffic to get a real answer?

Traffic is the limit that stalls many programs, and a platform comparison never mentions it.

A test needs enough visitors, and enough of them buying or signing up (conversions), to tell a real difference from chance. How many depends on two numbers: how often visitors convert today, and how small an improvement you want to catch. The standard rule of thumb fits on one line (1. Kohavi et al., 2009):

Visitors needed per version ≈ 16 × conversion rate × (1 − conversion rate) ÷ (the change you want to catch)²

I’m not a math guy, so here’s what it means in practice. Say 3% of the people who visit a page buy something, and you want to know whether a change would lift that to 3.3%. You’d need about 52,000 visitors to see each version, or about 103,000 in all. Optimizely, which sells testing software, gets nearly the same answer: about 51,000 per version (2. Optimizely, n.d.).

Two things move the number. A page that converts more often needs fewer visitors: at 10% conversion, the same 10% improvement takes about 14,000 per version. And a smaller improvement needs far more. Halve the improvement you’re hunting for and the visitors you need roughly quadruple.

Ron Kohavi, who led Microsoft’s experimentation platform team, gives a blunter rule of thumb for when testing is worth it: below tens of thousands of users, the math doesn’t work for most metrics a business cares about, and testing starts paying off around 200,000 users (3. Rachitsky, 2023). Most changes a team tests move results only a little, if at all (4. Kohavi et al., 2014). And the smaller the difference, the more visitors it takes to see it.

Here’s what that looks like. A team without enough traffic runs a test for 3 weeks, and the new version comes out 4% ahead. With so few visitors, a 4% gap can happen by chance alone, so the test hasn’t proved anything. The team ships the change anyway, and nobody knows whether it helped.

Before anyone evaluates a single platform, ask the team for 3 numbers: monthly visitors to the page you want to test, the share who buy or sign up, and the smallest improvement worth catching.

That last number is a business decision. It’s the smallest rise in sales or sign-ups that would be worth acting on, measured against today’s rate. If going from 3% of visitors buying to 3.3% would bring in enough revenue to justify the change, your number is 10%. The team wouldn’t bother shipping anything smaller, so the test doesn’t need to be big enough to see it. Testers call this number the minimum detectable effect.

Run the 3 numbers through the formula above, or any free sample-size calculator, then multiply by 2 and divide by the page’s weekly visitors. That’s how many weeks a test has to run before you can rule out chance (statistical significance). Optimizely also advises running every test for at least 7 days, to capture a full week of visitor behavior (5. Optimizely, n.d.). If the answer is “longer than anyone will wait,” a better tool won’t help. Test bigger changes, test where more visitors arrive, or accept that some decisions won’t be settled by a test.

Decide what counts as a win before the test starts

A mature program writes down its rule before the test runs: the one number that decides it, how big a change counts as a win, how long the test runs, and who makes the call. Then it sticks to that rule when the results come in ugly.

The common mistake is checking the test every day and stopping it as soon as it looks good (peeking). The team sees the new version pull ahead on day 4 and ships it. Each extra look gives random swings another chance to pass for a win, so a test checked daily declares far more false winners than one read once at the end (6. Johari et al., 2022). The platform’s live dashboard makes peeking easy, which is one more reason the rule has to come from the team.

This is the same habit as agreeing on what you measure before you measure it , applied to one test.

Act on losing tests as fast as winning ones

Most tested ideas don’t win. Across experiments run at Microsoft, the software company, about 1 in 3 ideas moved the key metric in the right direction, on average (4. Kohavi et al., 2014). The rest changed nothing or made things worse, and ideas the team was sure about land in that group all the time. A mature program expects that. It gets faster by dropping losing ideas quickly and moving to the next one on the list, and a faster tool does nothing for that.

An immature program shows a different pattern: tests that ran, lost, and shipped anyway because someone liked the idea. Speed in testing means how quickly the team settles ideas, winners and losers both. A team that acts on its results learns faster than a team that argues with them, whatever platform either one uses.

Where the platform matters

Once the team has those habits, the platform has 4 jobs:

  • Decide by chance which visitors see which version, so one group doesn’t end up with more of your best customers.
  • Record what every visitor in each group does, without losing any of it.
  • Do the math correctly when the test ends.
  • If you test changes inside an app or software product instead of on your website, show the change to some users and not others. These on-off switches are called feature flags.

Those 4 jobs also explain why “best platform” lists sort tools into a few groups. Tools built for testing marketing websites, sold as conversion rate optimization tools, include Optimizely and Wingify. Tools built for software teams, such as LaunchDarkly, fit tests inside an app or product. Some analytics tools include testing, which suits teams that already work from that tool’s data. Pick the group that matches where you test and how much traffic you have.

Sign off on buying a testing tool only after your team can show you 3 things in writing: a ranked list of ideas to test, the numbers showing your pages get enough visitors, and the rules for deciding which version won. Buying the tool comes last .

About the Author

Gene De Libero, Independent Martech Advisor, Digital Mindshare LLC

Gene De Libero has spent more than thirty years in marketing technology — as buyer, seller, builder, and advisor. He is the creator of the Marketing Technology Transformation® framework, sponsor of How Marketing Technology Works®, and Independent Martech Advisor at Digital Mindshare LLC, a New York consultancy serving CMOs whose stacks have stopped paying for themselves. He believes most martech investments fail not because the technology is wrong, but because the organization was never built to use it. He fixes that.

Frequently Asked Questions

What makes an experimentation program mature?

Four habits: a ranked list of ideas to test, each tied to a number it should move; enough traffic for a test to give a reliable answer; a rule for calling a winner, written down before the test starts; and the willingness to drop losing ideas quickly. Those habits are what make any testing tool’s results trustworthy.

Which experimentation platform is best for a mature program?

Start with the program instead. A disciplined team gets reliable answers from a basic tool, and an undisciplined team gets misleading ones from the most expensive suite. Once the habits exist, choose a tool that splits visitors at random, records what they do, calculates results correctly, and, for tests inside your product, can show a change to only some users.

How much traffic do you need to run experiments?

Enough that each version of the test collects the purchases, sign-ups, or other conversions needed to tell a real difference from chance, within a time your team will wait. If your key tests would need months, test bigger changes or test where more visitors arrive, such as the home page, instead of buying a better tool.
References
  1. Kohavi, R., Longbotham, R., Sommerfield, D., & Henne, R. M. (2009). Controlled experiments on the web: survey and practical guide. Data Mining and Knowledge Discovery, 18, 140-181. https://exp-platform.com/Documents/controlledExperimentDMKD.pdf
  2. Optimizely. (n.d.). Use minimum detectable effect to prioritize experiments. Optimizely Support Help Center. https://support.optimizely.com/hc/en-us/articles/4410283338253-Use-minimum-detectable-effect-to-prioritize-experiments
  3. Rachitsky, L. (2023). The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon). Lenny’s Newsletter. https://www.lennysnewsletter.com/p/the-ultimate-guide-to-ab-testing
  4. Kohavi, R., Deng, A., Longbotham, R., & Xu, Y. (2014). Seven rules of thumb for web site experimenters. Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1857-1866. https://exp-platform.com/Documents/2014%20experimentersRulesOfThumb.pdf
  5. Optimizely. (n.d.). How long to run an experiment. Optimizely Support Help Center. https://support.optimizely.com/hc/en-us/articles/4410283969165-How-long-to-run-an-experiment
  6. Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806-1821. https://doi.org/10.1287/opre.2021.2135