In 2007, President Obama’s campaign’s analytics lead, Dan Siroker, was certain that a video of the candidate would beat a plain photo on the sign-up page. He tested it anyway. Every video lost to every image. The winning combination, a family photo paired with a button that read “Learn More,” lifted sign-ups from 8.26% to 11.6%. That 40.6% jump was later tied to roughly 2.9 million extra email addresses and about $60 million in donations, as Siroker documented for Optimizely. The expert was wrong, and only a disciplined test caught it.
That is the entire promise of experimentation, and it is exactly where most programs fall apart. In Issue 08, we placed the trust signals that shape the first 5 seconds. Every one of those placements was an educated guess, and educated guesses are worth nothing until an experiment confirms them. Running that experiment properly is much harder than buying the tool that runs it. A team can own multiple testing tools and still have no structured experimentation program, in the same way that owning a stethoscope does not make someone a doctor.
At Microsoft, after running thousands of experiments across its major products, Ron Kohavi, formerly Vice President of Analysis and Experimentation at Bing, documented a finding that most teams misread as discouraging and should read as structural: only about one-third of experiments were successful at improving the key metric they were designed to improve, per Harvard Business Review’s “The Surprising Power of Online Experiments”, co-authored by Kohavi and HBS Professor Stefan Thomke.
One in three at Microsoft, with one of the most sophisticated experimentation programs in the industry. At the median organization, GrowthBook’s analysis of Kohavi’s published research puts it lower: roughly 10% of experiments move the metric they were designed to improve.
Experimentation is not a system for confirming good ideas. It is a system for finding out which ideas are actually good, including the ones nobody thought were worth testing.
Here are the 5 reasons programs fail:
In 2012, a Bing engineer suggested a small change to the way ad headlines were displayed. The idea sat in a backlog for months, ranked as low priority, because it looked trivial. When someone finally ran it as a test, revenue rose about 12%, worth more than $100 million a year in the US alone, with no damage to the user-experience metrics the team watched. It became the most profitable idea in Bing’s history, and as Kohavi and Thomke wrote in Harvard Business Review, nobody in the room could tell it apart from a hundred forgettable tweaks.
This is the rule, not the exception. At Microsoft, only about one in three well-built experiments actually improves the metric it targets. The rest come back flat or negative. So a program built on “let us ship the ideas we believe in” ships mostly losers, because human judgment about which idea will win is close to a coin flip.
A program built on “let us test the ideas we cannot rank” is the one that finds the rare, enormous winner hiding in the backlog. Every test should start from a question grounded in research, which is the point of the diagnosis work in our conversion funnel analysis, and never from the loudest voice in the meeting.
A test that looks like a clear winner on day three is usually just noise that has not settled yet. “Statistically significant” only carries meaning if you look once, at a finish line you set before the test started. Every extra peek is another chance for randomness to trip the wire. Check a test every morning and stop the moment it turns green, and roughly one in four of your “winners” is pure luck. This is why so many dashboards fill with victories while the real conversion rate refuses to move.
The fix is a rule, not a feature: Decide the sample size and the end date before launch, then read the result once. If you truly need to watch a test live, use a method built for continuous monitoring, but never fake it by peeking at an ordinary calculator.
Split your traffic evenly, and it comes back 52 to 48; that is not a close finish. That is a broken experiment. When the split drifts further from your plan than chance allows, something misfired: a variant that loaded slowly, a redirect that failed on one side, tracking that fired unevenly. The result is unreadable, yet most teams never run the quick check that catches it. Microsoft’s experimentation platform refuses to show results at all until this check passes, because a green result on a broken test is simply a bug you shipped to everyone. The usual culprit is measurement itself, the same rot behind these ten GA4 configuration mistakes.
Google once tested 41 shades of blue for its links and found a slightly purpler one that earned more clicks, reportedly worth about $200 million a year in ad revenue. It is the most famous A/B test ever run, and it is also a warning. Analysts have noted that the effect was tiny, under 1% of revenue, and that different screens render color differently, so the “magic blue” may have been a random flicker promoted to a discovery. A number can be perfectly real and still be the wrong number to chase.
Two habits protect you:
One great test is luck. A program is what you have when every test, winner or loser, becomes a lesson you keep. The Obama 2012 campaign ran more than 500 experiments across web and email in 20 months and lifted donation conversions by 29% and sign-up conversions by 161%, as documented in this review of the campaigns, because each result sharpened the next. Booking.com built the same habit into its culture and became famous for running well over a thousand experiments at once. The gap between those teams and a stalled one is not budget or tooling. It is memory. Teams that record their losers stop repeating them. Teams that delete them re-argue the same opinions every quarter and wonder why the number never moves.
Score your last ten tests, one point each:
Here is the part that makes all of this worth doing. Even the best teams in the world see only about one in three tests win. That is not failure, that is the game. A program built to expect it, and to trust only the wins that are real, compounds quietly into an advantage competitors cannot copy. A program built on green dashboards and gut feel just spins in place.
The UX mistakes from Issue 07 and the trust signals from Issue 08 are worth nothing until a disciplined test proves them. In Issue 10, we build the roadmap that makes tests compound, so each result feeds the next instead of vanishing into a dashboard.
Want an outside read on where your program sits on that ladder? Talk to our CRO and A/B testing team.
Minal Joshi is a content marketer at Krish with a flair for eCommerce and Digital Commerce aspects. She is a MarTech fanatic with a knack of writing with which, she helps brands to curate, create, & commence digital brand positioning. Sharing insights via articles, case studies, eBooks, Infographics, and other forms of content creation is what she lives for. Being an ardent traveler, when not writing, you'll find her sipping coffee into the mountains or petting a stray.
17 July, 2026 Most marketing conversations about machine learning (ML) collapse it into a single capability. Teams talk about "using AI" or "adding ML to the stack" as though it were one tool with one function. It is not, and treating it that way is why most ML initiatives either overreach and fail or underdeliver and disappoint.
Never miss any post, stay tuned!