The most consequential decision in product delivery isn't how well something gets built. It's whether it should have been built at all—and that decision usually gets made with far less scrutiny than the build itself receives.
What Happens When You Actually Check
Most organizations assume their product ideas are good ideas, more or less by default. The evidence says that assumption is wrong more often than it's right.
Ronny Kohavi, who built and ran large-scale controlled-experimentation platforms at Microsoft, Bing, and later Airbnb, has published some of the most extensively reviewed data available on this question—what actually happens when a shipped idea is tested against a real, randomized control group instead of just being assumed to work. The pattern holds with remarkable consistency across very different companies: at Microsoft, roughly two-thirds of tested ideas failed to improve the metric they were built to improve. At Bing, the failure rate ran higher still, around 85%. At Airbnb, roughly 92%. Booking.com, running more concurrent experiments than almost any company in the world, has reported a similar result—the large majority of ideas its own product teams believed would help did not, when actually measured.
Exhibit 1
Share of tested product ideas that failed to move the metric they were built for
Measured against a real, randomized control group — at three of the most data-driven product organizations in the industry.
Source: Kohavi, R. et al., Trustworthy Online Controlled Experiments; exp-platform.com.
Exhibit 2
An idea seems obviously good
Full delivery capacity gets committed
It ships as planned, on schedule
It's tested against real usage—or not tested at all
Most of the time, it doesn't move the metric it was built for
Why shipping on schedule and shipping something that works are two different outcomes.
These are not companies with weak product instincts. They are among the most sophisticated, data-driven product organizations in the industry, and their own numbers say the same thing: most ideas that look good enough to build turn out not to earn their cost once someone checks.
The ideas that survive contact with a controlled test are the exception, not the rule—for everyone, not just for teams that are getting it wrong.
"It Seemed Like a Good Idea" Isn't a Business Case
If even the best product organizations in the world are wrong most of the time about which ideas will work, the honest conclusion isn't that those organizations are bad at their jobs. It's that intuition alone was never going to be a reliable filter—and every organization that skips validation and commits delivery capacity straight from "this seems like a good idea" is making the same bet those companies' own data shows usually doesn't pay off.
Melissa Perri calls the organizational pattern that results from skipping this step the build trap: equating more shipped output with more success, and losing track of whether any particular thing that got shipped actually created value. Her proposed fix reframes the whole problem as a capital allocation question—fund product work the way a venture investor funds a portfolio, putting a small amount of capacity against many unproven ideas, and only committing serious capacity once an idea has evidence behind it.
That's also the core discipline behind Eric Ries's build-measure-learn loop: treat what you ship as a test of an assumption, not a finished commitment, until the data says otherwise. The goal isn't to move slower. It's to spend the smallest amount of capacity necessary to find out whether an idea is one of the roughly one-in-three that works—before spending the much larger amount of capacity it takes to fully build, harden, and maintain it.
The Economics of Checking First
This connects directly to the same economic logic that should govern any prioritization decision: capacity is finite, and every dollar of it committed to an unvalidated feature is a dollar not available for the smaller share of ideas that would have actually earned their investment.
The fix costs far less than the mistake. A validation step—a small experiment, a narrow release, a real test against real usage—costs a fraction of what building, shipping, and then indefinitely maintaining the wrong thing costs. The organizations with the best data on this question aren't the ones that guess less often. They're the ones that built the discipline to find out before they commit.
The Real Question, Asked Earlier
Not every feature deserves to be built—not because most product ideas are bad ones, but because most ideas, even from strong teams, don't turn out to be worth what they'd cost until someone actually checks.




