A/B testing

Better Creative, Fewer Tests: A New Framework for Efficient Self-Improving Systems

  • ARF; MSI

This MSI working paper introduces TextBO, a novel AI framework designed to improve marketing decisions more efficiently by minimizing costly evaluation cycles. By combining large language models with Bayesian optimization principles, the approach enables AI systems to iteratively refine outputs—such as ad creatives—while requiring fewer real-world tests. The result: faster learning, better-performing outcomes, and a more scalable path to AI-driven decision-making.

Member Only Access

Learning Across Marketing Experiments: A Bayesian Approach to Improving Targeting

  • ARF, MSI

Companies run many marketing experiments, but most A/B tests are analyzed independently—limiting what firms can learn about how customers respond to interventions over time. This research introduces a hierarchical Bayesian framework that integrates data from many experiments simultaneously to estimate customer-level responsiveness to marketing. Using large-scale field experiments, the model decomposes treatment effects into customer, campaign and timing components and uses these insights to improve targeting decisions. The results show that most variation in marketing effectiveness comes from persistent differences in customer responsiveness, enabling firms to better identify who to target and when.

Member Only Access

Improving AI-Driven Marketing Content Using LLM-Generated Knowledge

  • ARF, MSI

As generative AI becomes a central tool for producing marketing content, firms increasingly rely on fine-tuning models using engagement data, such as A/B test results. This MSI working paper argues that optimizing only for “what works” risks reward hacking, clickbait and poor generalization. The authors propose a knowledge-guided alignment framework in which large language models (LLMs) generate and validate hypotheses about why content performs well, and then use this knowledge to guide fine-tuning. Using more than 23,000 A/B-tested news headlines, the study shows that knowledge-guided AI produces higher engagement, avoids clickbait and generalizes better—especially in low-data settings.

Member Only Access

The Persuasive Power of Pets: New Insights for Influencer Strategy

  • ARF
  • JOURNAL OF ADVERTISING RESEARCH

A recent Journal of Advertising Research study presents the first empirical evidence that petfluencers can outperform human influencers in driving engagement and willingness to pay. Across four studies—including a real-world A/B test and controlled experiments, the researchers show that pet influencers benefit from higher perceived sincerity, and that message framing matters. When temporal cues align with consumers’ propensity to anthropomorphize animals, the persuasive impact increases further.

Member Only Access

The Power of a Word: Measuring the Real Impact of Language in Marketing

  • ARF
  • MSI

How much impact can a single word have in a marketing message? A new study introduces a cutting-edge causal inference framework using language models to quantify the exact influence of words—such as “you” or “thank you”—on consumer engagement. The findings show that traditional A/B tests often miss these nuanced effects, while this new method isolates true word-level causal impacts, with big implications for advertising and fundraising success.

Member Only Access

LOLA: Revolutionizing Content Experiments with LLM-Assisted Online Learning

In the rapidly evolving digital content landscape, media firms and news publishers require automated and efficient methods to enhance user engagement. This study introduces the LLM-Assisted Online Learning Algorithm (LOLA), a novel framework that integrates Large Language Models (LLMs) with adaptive experimentation to optimize content delivery. Leveraging a large-scale dataset from Upworthy, which includes 17,681 headline A/B tests, the study investigates three pure-LLM approaches and finds that prompt-based methods perform poorly, while embedding-based classification models and fine-tuned open-source LLMs achieve higher accuracy.


LOLA combines the best pure-LLM approach with the Upper Confidence Bound (UCB) algorithm to allocate traffic and maximize clicks adaptively. Numerical experiments on data from the website Upworthy show that LOLA outperforms the standard A/B test method, pure bandit algorithms and pure-LLM approaches, particularly in scenarios with limited experimental traffic. This scalable approach is applicable to content experiments across various settings where firms seek to optimize user engagement, including digital advertising and social media recommendations.

Member Only Access
  • Article

From Data to Dollars: How Shiseido Leveraged Machine Learning to Create Meaningful Customer Experiences

Our research uses A/B testing of diverse (social, native, mobile, video, addressable TV, zone TV, etc.) digital options on top of the national layer of brand TV as it exists.

It is anticipated that some of the A/B tests will include bold options such as social-dominant digital allocation, native-dominant, mobile-dominant, etc.

Questions to be addressed include:

-What is the optimal form of digital/advanced platform advertising/native to use synergistically with traditional TV?

-How does this differ by product vertical (CPG, Auto, Rx, Tune-in, Other)?

-How does this differ by brands in the same vertical?

-Are there so many variables that each brand has to continually test for optimal digital/advanced mix because new creative or other factors could change the optimal digital mix?

Review the Audience Measurement program and register.

 

Member Only Access
  • Article

Using Daily, Campaign-Level Insights to Maximize ROAS with Machine Learning MMM

Our research uses A/B testing of diverse (social, native, mobile, video, addressable TV, zone TV, etc.) digital options on top of the national layer of brand TV as it exists.

It is anticipated that some of the A/B tests will include bold options such as social-dominant digital allocation, native-dominant, mobile-dominant, etc.

Questions to be addressed include:

-What is the optimal form of digital/advanced platform advertising/native to use synergistically with traditional TV?

-How does this differ by product vertical (CPG, Auto, Rx, Tune-in, Other)?

-How does this differ by brands in the same vertical?

-Are there so many variables that each brand has to continually test for optimal digital/advanced mix because new creative or other factors could change the optimal digital mix?

Review the Audience Measurement program and register.

 

Member Only Access