Meta’s AdLlama: An LLM for Ad Text Variations
Generative AI makes it easy to produce more ads. Whether those ads perform better is a separate question. A study by Meta researchers on AdLlama, a model that writes variations of ad text on Facebook, offers one measured answer, with clear limits.
By EpicflarePublished Updated 6 min read
Evaluation: 10-week A/B test on Facebook (February to April 2024). Each advertiser was randomly assigned one of the two models in the Text Generation feature.
Source and date
The results discussed here come from one primary source: a research paper by Meta researchers Daniel R. Jiang, Alex Nikulkov, Yu-Chia Chen, Yang Bai and Zheqing Zhu, “Improving Generative Ad Text on Facebook using Reinforcement Learning” (arXiv:2507.21983) (opens in a new tab), first posted in July 2025. Checked in October 2026: the figures are unchanged in the paper’s December 2025 revision.
How AdLlama Works: Moving Beyond Imitation
AdLlama is a language model integrated into Meta’s Text Generation feature for advertisers. An advertiser enters an ad text, the model suggests variations, and the advertiser chooses which ones to use.
To see what is different, start with how generative text models for ads are usually trained. They are shown many examples of “good” ads, written by people or by larger models, and learn to imitate them. The result is fluent, plausible copy.
The Limitation of Imitation
Imitation does not tell the model whether an ad actually drove clicks or conversions: it learns style rather than results. Meta’s previous Text Generation model was trained this way. For AdLlama, the researchers added a post-training method they call reinforcement learning with performance feedback (RLPF):
Collect comparable ads
Historical Facebook ads in which advertisers tested several versions of the body text while keeping the image, targeting and other settings identical.
Compare them by click-through rate
Within each ad, the text variations are compared by CTR, producing pairs of preferred and non-preferred texts.
Train a reward model
A reward model learns to predict which of two texts is likely to achieve the higher CTR. This is the performance signal.
Post-train the text model
The text model is then trained with reinforcement learning to write variations that the reward model scores higher, rather than only imitating examples.
Why it matters
AdLlama links text generation to a measured performance signal instead of imitation alone. The same principle applies to any team using generative AI in advertising: define the metric that matters, then evaluate outputs against it.
What the Test Measured
Meta evaluated AdLlama in a 10-week A/B test on Facebook, from February 16 to April 25, 2024. Each advertiser was randomly assigned either an improved version of the imitation model (control) or AdLlama (test). The analysis covers direct-response ads, aggregated per advertiser.
| Measure | Reported result |
|---|---|
| Advertisers in the test | 34,849 |
| Ad variations in the test | About 640,000 |
| Advertiser-level CTR | +6.7% |
| Ad variations per advertiser | +18.5% |
Relative changes reported by Meta researchers for the 10-week A/B test on Facebook (February to April 2024). CTR rose from about 3.1% to 3.3% (p = 0.0296); ad variations from about 16.8 to 19.9 per advertiser (p < 0.01). Source: Jiang et al., arXiv:2507.21983. These are not Epicflare results.
In other words, advertisers using AdLlama obtained 6.7% more clicks per impression than those using the imitation model: a relative gain, from roughly 3.1% to 3.3% in absolute terms. The comparison is with an AI model trained by imitation, not with ads written without AI assistance.
Advertisers in the AdLlama group also created 18.5% more ad variations, while the number of ads they created stayed statistically the same. The authors interpret this as higher satisfaction with the suggestions; it is inferred from usage rather than measured directly.
The authors add that on mature, highly optimized platforms such as Facebook, even small increases in CTR are typically difficult to achieve.
CTR Is Not ROI
The paper states that the CTR gain “represents a substantial improvement in advertiser return on investment”. This is the authors’ interpretation of the CTR result, not a separate measurement: the study reports clicks, impressions and ad creation, not conversions, sales or return on ad spend.
The authors also note that CTR is typically considered a proxy for the true goal, conversion rate, and that they used CTR because conversion data is sparser and noisier. For an advertiser, a higher CTR is a useful signal; its effect on CPA or revenue still has to be measured against a defined baseline.
What the Study Does Not Show
- Other creative elements: The model rewrites the body text of an ad written by the advertiser. Headlines, images, video and formats were not part of the test.
- Other platforms and objectives: The results come from one feature on Facebook and from direct-response ads. They may not transfer to other platforms, formats or goals.
- Continuous learning: Training used historical data once (offline reinforcement learning). The authors describe learning from new results as a possible next step.
- Other qualities of the text: The model optimizes for performance. The authors note possible trade-offs with creativity and with following advertiser instructions such as tone.
- The advertiser’s choice: Advertisers select which suggestions to use, a step the model does not take into account.
Beyond Ad Text
It is tempting to extrapolate from text to the whole ad, with headlines, images, video, calls to action and formats chosen automatically for each audience. The study does not test any of this. Applying the same principle to other elements would require its own data, training and controlled tests.
For a broader view of producing and testing creative variations for different audiences, see Achieving Creative Hyper-Personalization with AI.
What This Means for Media Buyers Today
For advertisers and media buyers, the practical takeaways are operational.
Treat Text Variations as Something to Test
The study indicates that the choice of text variation can affect CTR, even on a mature platform. Whether a higher CTR improves acquisition cost or ROI in your own account needs its own measurement.
The Creative Strategist Role Is Shifting
Less time goes into writing every variation by hand, and more into providing a strong core concept, then curating, refining and testing the AI’s suggestions, much like a creative director.
AdLlama shows, on a large platform, that training a text model on performance feedback can raise CTR compared with imitation alone. It is a useful data point about where ad creative tools are heading, not a guarantee of results for other advertisers, formats or metrics.

Go further
Agentic AI and media buying: An investment guide
A framework for assessing architecture, interoperability, operating costs and control when investing in AI for media buying.
PDF · 27 pages · Free · one short form unlocks every resource
Related reading
- AI for Ad Operations: From Request to Controlled Execution
Explore an AdOps workflow for deal requests, with explicit business rules, human approval and practical criteria for evaluating AI support.
- Bringing AdTech Expertise into AI Workflows
Turn platform knowledge, business rules and operational playbooks into AI workflow requirements, then test how reliably they are applied.
- AI Agents or Automated Workflows? Choosing the Right Approach
Compare AI agents and automated workflows through task requirements, control, evaluation and operating cost to choose a suitable starting point.
- Building an AI Workflow: Start Small, Evaluate, Expand
Scope an AI workflow around one operational task, define evaluation criteria and add complexity only when testing shows it is needed.