Skip to main content

Ce site est aussi disponible en français.

Voir en français
AI for marketing

Meta’s AdLlama: An LLM for Ad Text Variations

Generative AI makes it easy to produce more ads. Whether those ads perform better is a separate question. A study by Meta researchers on AdLlama, a model that writes variations of ad text on Facebook, offers one measured answer, with clear limits.

By EpicflarePublished Updated 6 min read

BaseLlama 2 Chat (7B)General-purpose language model
ControlImitation modelFine-tuned to imitate curated ad examples
TestAdLlamaFurther trained with performance feedback (CTR)

Evaluation: 10-week A/B test on Facebook (February to April 2024). Each advertiser was randomly assigned one of the two models in the Text Generation feature.

The two models compared. Both models start from the same base model. AdLlama adds a reinforcement learning step that uses historical click-through data as a reward signal. Source: Jiang et al. (Meta), arXiv:2507.21983.

Source and date

The results discussed here come from one primary source: a research paper by Meta researchers Daniel R. Jiang, Alex Nikulkov, Yu-Chia Chen, Yang Bai and Zheqing Zhu, “Improving Generative Ad Text on Facebook using Reinforcement Learning” (arXiv:2507.21983) (opens in a new tab), first posted in July 2025. Checked in October 2026: the figures are unchanged in the paper’s December 2025 revision.

How AdLlama Works: Moving Beyond Imitation

AdLlama is a language model integrated into Meta’s Text Generation feature for advertisers. An advertiser enters an ad text, the model suggests variations, and the advertiser chooses which ones to use.

To see what is different, start with how generative text models for ads are usually trained. They are shown many examples of “good” ads, written by people or by larger models, and learn to imitate them. The result is fluent, plausible copy.

The Limitation of Imitation

Imitation does not tell the model whether an ad actually drove clicks or conversions: it learns style rather than results. Meta’s previous Text Generation model was trained this way. For AdLlama, the researchers added a post-training method they call reinforcement learning with performance feedback (RLPF):

  1. Collect comparable ads

    Historical Facebook ads in which advertisers tested several versions of the body text while keeping the image, targeting and other settings identical.

  2. Compare them by click-through rate

    Within each ad, the text variations are compared by CTR, producing pairs of preferred and non-preferred texts.

  3. Train a reward model

    A reward model learns to predict which of two texts is likely to achieve the higher CTR. This is the performance signal.

  4. Post-train the text model

    The text model is then trained with reinforcement learning to write variations that the reward model scores higher, rather than only imitating examples.

Why it matters

AdLlama links text generation to a measured performance signal instead of imitation alone. The same principle applies to any team using generative AI in advertising: define the metric that matters, then evaluate outputs against it.

What the Test Measured

Meta evaluated AdLlama in a 10-week A/B test on Facebook, from February 16 to April 25, 2024. Each advertiser was randomly assigned either an improved version of the imitation model (control) or AdLlama (test). The analysis covers direct-response ads, aggregated per advertiser.

Results reported in the study: AdLlama vs. imitation model
MeasureReported result
Advertisers in the test34,849
Ad variations in the testAbout 640,000
Advertiser-level CTR+6.7%
Ad variations per advertiser+18.5%

Relative changes reported by Meta researchers for the 10-week A/B test on Facebook (February to April 2024). CTR rose from about 3.1% to 3.3% (p = 0.0296); ad variations from about 16.8 to 19.9 per advertiser (p < 0.01). Source: Jiang et al., arXiv:2507.21983. These are not Epicflare results.

In other words, advertisers using AdLlama obtained 6.7% more clicks per impression than those using the imitation model: a relative gain, from roughly 3.1% to 3.3% in absolute terms. The comparison is with an AI model trained by imitation, not with ads written without AI assistance.

Advertisers in the AdLlama group also created 18.5% more ad variations, while the number of ads they created stayed statistically the same. The authors interpret this as higher satisfaction with the suggestions; it is inferred from usage rather than measured directly.

The authors add that on mature, highly optimized platforms such as Facebook, even small increases in CTR are typically difficult to achieve.

CTR Is Not ROI

The paper states that the CTR gain “represents a substantial improvement in advertiser return on investment”. This is the authors’ interpretation of the CTR result, not a separate measurement: the study reports clicks, impressions and ad creation, not conversions, sales or return on ad spend.

The authors also note that CTR is typically considered a proxy for the true goal, conversion rate, and that they used CTR because conversion data is sparser and noisier. For an advertiser, a higher CTR is a useful signal; its effect on CPA or revenue still has to be measured against a defined baseline.

What the Study Does Not Show

  • Other creative elements: The model rewrites the body text of an ad written by the advertiser. Headlines, images, video and formats were not part of the test.
  • Other platforms and objectives: The results come from one feature on Facebook and from direct-response ads. They may not transfer to other platforms, formats or goals.
  • Continuous learning: Training used historical data once (offline reinforcement learning). The authors describe learning from new results as a possible next step.
  • Other qualities of the text: The model optimizes for performance. The authors note possible trade-offs with creativity and with following advertiser instructions such as tone.
  • The advertiser’s choice: Advertisers select which suggestions to use, a step the model does not take into account.

Beyond Ad Text

It is tempting to extrapolate from text to the whole ad, with headlines, images, video, calls to action and formats chosen automatically for each audience. The study does not test any of this. Applying the same principle to other elements would require its own data, training and controlled tests.

For a broader view of producing and testing creative variations for different audiences, see Achieving Creative Hyper-Personalization with AI.

What This Means for Media Buyers Today

For advertisers and media buyers, the practical takeaways are operational.

Treat Text Variations as Something to Test

The study indicates that the choice of text variation can affect CTR, even on a mature platform. Whether a higher CTR improves acquisition cost or ROI in your own account needs its own measurement.

The Creative Strategist Role Is Shifting

Less time goes into writing every variation by hand, and more into providing a strong core concept, then curating, refining and testing the AI’s suggestions, much like a creative director.

AdLlama shows, on a large platform, that training a text model on performance feedback can raise CTR compared with imitation alone. It is a useful data point about where ad creative tools are heading, not a guarantee of results for other advertisers, formats or metrics.

Go further

Agentic AI and media buying: An investment guide

A framework for assessing architecture, interoperability, operating costs and control when investing in AI for media buying.

PDF · 27 pages · Free · one short form unlocks every resource

Related reading