Photo By: Declan Sun
Live A/B testing has been at the center of digital growth for more than two decades.
E-commerce companies have tested checkout buttons. Subscription businesses have experimented with pricing. Product teams have compared different features and customer experiences. The basic process has remained the same: develop an idea, split customers into groups, measure the results, and roll out the option that performs better.
That approach is beginning to show its limits.
Digital customer journeys are becoming more complex. Privacy regulations are making customer tracking more difficult. Competition is also moving faster, with strategies that once took months now changing within days.
These changes raise an important question: What if testing every major business decision on real customers is no longer the most efficient approach?
The Hidden Costs of Live A/B Testing
A/B testing is built on a powerful statistical principle, but running experiments with real customers creates several challenges.
First, live experiments take time and money. Testing a new price, promotion, or product placement can require weeks of customer traffic before a company has enough data to make a reliable decision. During that period, some customers may be exposed to an option that generates lower profits or weaker results.
Second, live experiments do not always reveal why something worked.
Outside factors can influence the results. Seasonal demand, competitor stock shortages, changes to advertising algorithms, and broader market conditions can all affect customer behavior.
A company might see a short-term increase in sales and conclude that an experiment was successful. Months later, the company may discover that the promotion simply caused customers to purchase earlier than they normally would have.
A promotion can increase sales today while reducing full-price purchases later.
Third, live testing can discourage companies from exploring bigger ideas.
Real-world experiments carry real financial risk. That risk often pushes teams toward small, low-risk changes such as button colors, headlines, or page layouts. Larger ideas can be harder to test.
Those ideas might include new pricing structures, regional pricing strategies, dynamic product bundles, or entirely different subscription tiers.
The Rise of Calibrated Synthetic Simulation
A different approach is gaining attention: using computer simulations to evaluate business decisions before making changes in the real world.
Companies can build models that simulate how different customer groups might respond to a new price, promotion, or product strategy. These models can also account for different market conditions.
Recent research from Amazon Science examined this idea. Researchers tested whether AI models could simulate customer responses using data from 67 historical A/B tests.
The research highlighted both the potential of synthetic experimentation and one of its biggest risks.
The Problem With Uncalibrated AI
Simple AI simulations can produce overly positive results.
Asking an AI model to “act like a consumer” does not necessarily create an accurate picture of how people will behave in the real world. The model may produce responses that sound realistic without properly accounting for economic trade-offs, price sensitivity, or actual purchasing behavior.
The Value of Calibration
The results changed when researchers calibrated the simulation models using historical customer data and real-world constraints.
Those calibrated models produced results that were much closer to the outcomes of actual experiments.
The lesson is straightforward: Synthetic customer simulations can be useful, but they need to be grounded in real economic and behavioral data.
AI alone is not enough.
From Live Discovery to Targeted Validation
This approach could change the way companies think about experimentation.
The goal is not necessarily to eliminate A/B testing. The bigger opportunity is to change the role that live testing plays.
Shenbo Xu, Co-Founder and CTO of Kapnova, has worked on causal inference in complex observational data through research at MIT, along with quantitative work at Point72 and Scale AI. His work reflects a broader challenge in the field: connecting statistical simulation with practical commercial decisions.
Traditional growth teams use live customer traffic to discover what works. They may run dozens of experiments and see which ideas produce positive results.
A simulation-driven approach reverses that process.
1. Simulated Exploration
AI agents and quantitative models can evaluate thousands of potential scenarios before anything is shown to customers.
These scenarios can include different prices, promotions, product bundles, subscription structures, and other commercial strategies.
2. Causal Analysis
Specialized models can then identify which strategies are most likely to generate genuine additional revenue.
The analysis can account for factors such as price sensitivity, existing demand, customer behavior, and competitive responses.
3. Targeted Validation
A company can then use live A/B testing to validate the most promising decisions before making a larger investment.
A business might evaluate 1,000 possible strategies through simulation and narrow the list down to a handful of options. Live testing becomes the final validation step rather than the primary discovery tool.
A New Benchmark for Commercial Growth
Generative AI is making it faster and easier to create digital experiences, advertisements, promotions, and product variations.
That development is shifting the biggest challenge in business growth.
The advantage may no longer come from simply running more live experiments. It may come from understanding the likely financial consequences of a decision before committing significant resources.
The traditional approach relied heavily on trial and error with real customers.
The emerging approach is closer to calibrated decision-making.
AI can help identify opportunities. Quantitative models can estimate their potential impact. Live experiments can then provide real-world validation.
Consumer brands may ultimately gain more from evaluating fewer ideas more intelligently than from simply running more A/B tests.
Combining AI-driven discovery with quantitative simulation and targeted live testing can reduce some of the guesswork involved in high-stakes commercial decisions.
The future of experimentation may not mean eliminating A/B tests. It may mean using them more selectively, after much of the discovery work has already been completed in a simulated environment.




