Ad Auctions for LLMs via Retrieval Augmented Generation

In the field of computational advertising, the integration of ads into the outputs of large language models (LLMs) presents an opportunity to support these services without compromising content integrity. This paper introduces novel auction mechanisms for ad allocation and pricing within the textual outputs of LLMs, leveraging retrieval-augmented generation (RAG). We propose a segment auction where an ad is probabilistically retrieved for each discourse segment (paragraph, section, or entire output) according to its bid and relevance, following the RAG framework, and priced according to competing bids. We show that our auction maximizes logarithmic social welfare, a new notion of welfare that balances allocation efficiency and fairness, and we characterize the associated incentive-compatible pricing rule. These results are extended to multi-ad allocation per segment. An empirical evaluation validates the feasibility and effectiveness of our approach over several ad auction scenarios, and exhibits inherent tradeoffs in metrics as we allow the LLM more flexibility to allocate ads.

Paper

References (48)

Scroll for more · 36 remaining

Similar papers

Peer review

Reviewer 2L4f4/10 · confidence 3/52024-06-17

Summary

The paper considers the integration of ads into large language model (LLM) generation, as well as the design of a mechanism for ad allocation and pricing that comes with this integration. The paper introduces the notion of a "segment auction", where the output discourse is break down into various segments. Following the retrieval augmented generation (RAG) framework, the auction selects winner for each segment considering the bids and relevance, and provide the winning ad to the LLM as part of the prompt to generate segment with the ad incorporated. The paper provides theoretical analysis on the desirable properties of the segment auction, as well as empirical analysis in a synthetic setting.

Strengths

- The paper does a good job connecting the ideas to theoretical concepts, and provides solid theoretical analysis - The paper is well-motivated, tackling a novel problem that has great practical implications

Weaknesses

- The paper uses strong assumptions on the retrieval component being calibrated to the click through rate, which forms the basis of most theoretical results provided. In the experiment section, the relevance measure are instead estimated by a model from the sentence-transformers library, which leaves a gap between theoretical guarantees and empirical approaches. - Under the general segment auction exposited in Appendix C, where each relevance measure is additionally dependent on all previous segments and allocations, the assumption of calibrated relevance seem even harder to achieve. Without this assumption and its associated theoretical properties, the paper presents only incremental technical contribution since running a second price auction and incorporating the winner of the auction into the LLM prompt lacks technical novelty. - The experiments section in the paper is not compelling enough. - The five types of auctions considered in the Experiment section, which varies in whether replacement or relevance measure is used, are not well-motivated and appear disconnected from the earlier theoretical section, which focus on only the segment auction with replacement. The authors do not provide a compelling reasoning for considering this five variants, and the result analysis does not provide a strong case or intuition for which should be chosen (see line 273-287) - the writing and analysis for these parts could be improved. - Both the numbers in output quality and the qualitative example provided in Figure 4 suggests that the addition of ads in the generation quite notably impairs the quality of the generated content. Particularly worrisome is the fact that in both the single allocation and multi allocation example (as well as in the examples in Appendix G.2), as soon as the paragraph starts pivoting to the ads *in the first sentence*, it never goes back in providing *any* additional useful information to answering the original question *"Can you suggests books similar to `to kill a mocking bird'"*. This seems to empirically suggests that the presence of ads in the earlier segment may impair the quality of content generated in the following segments. - The experiments assume that each segment is one sentence in the experiments and add an ad to each segment. This is arguably too extreme for it to be deployed and a more realistic setting might be to integrate only a few ads to an entire page of results, which will suggests that the segment be longer than a single sentence. Additional experimental results should demonstrate that the proposed mechanism work as well under a longer segment. - When the number and categories of the advertisers are not diverse enough, augmenting an irrelevant ad appears to come with a significant cost to the quality of the output. The bid system also introduces the potential problem of misleading and corrupting the accuracy of the response, which is a new problem not faced by traditional ad auctions. This should be either empirically evaluated or discussed in more detail (see more in Limitations part). - Typo in line 140 - both the per-click and per-impression price are denoted using the same symbol.

Questions

- How will the relevance measure $q_i$ be obtained if the mechanism is put into practice, and will there be ways of ensuring accuracy and calibration? - The qualitative empirical results right now (e.g. Figure 4) shows that the generated output stops providing useful information to the user once the ad starts appearing in the generated text in the first sentence, and all subsequent generations are just rhetorical transitions to another ad. Is this effect prevalent? Are there additional qualitative or quantitative results showing the effects of ad on subsequent segments and potential solutions to preserve an informative output? - Current pricing uses the VCG mechanism for incentive compatibility. In ad auctions GSP and first price auctions are frequently used. Can these alternative pricing schemes can be easily implemented under the current scheme?

Rating

4

Confidence

3

Soundness

2

Presentation

2

Contribution

1

Limitations

The authors mentions the inherent trade-off between revenue and quality that's evident from the empirical observations. Different from traditional ad auctions, under LLM when the generation of later texts depends on prior generations, integrating additional ad leave open potential room for attack and injecting biased information through the type of the ad itself. The authors can provide a more detailed exposition of the potential social impact of such mechanism.

Authorsrebuttal2024-08-08

We appreciate the reviewer's responsiveness. ### Novelty We first emphasize that our main contribution is to propose the idea of combining RAG and ad auctions for LLM and laying down a foundation in this line of research. Indeed, even though applying RAG into the LLM ad auction seems a very plausible direction, there has been no single work dedicated to it. That is to say, our primary objective and contribution is to propose a potentially promising framework for the LLM ad auctions, which we believe will inevitably emerge in a near future, and validate its feasibility. Although RAG was introduced in 2020 and advertising has long been a fascinating field of study, the integration of ads within the RAG framework (to potentially monetize LLMs) is a novel approach that was not studied theoretically nor validated through empirical research. As the reviewer agreed with the idea of perturbation, although this idea is easily understandable in hindsight, it is not easy to come up with at first glance. Also, we think that the idea of reverse-engineering the mechanism design by starting from an exogenously given allocation function by RAG is not a trivial perspective to derive. We respectfully disagree with the reviewer that a paper necessarily includes complicated theoretical techniques or proofs, but believe that simple and concise but novel ideas often advance the fields in many directions. Finally, as the reviewer pointed out, delving more into logarithmic social welfare seems very promising research directions. We indeed have exposited several initial properties in Appendix D including that it can be viewed as a version of weighted Nash social welfare, but further theoretical results remains open problems, e.g., whether this is consistent with other notions of fairness in terms of reachability/approximability. We will make this point clearer in the main body of the paper. ### Experiments with longer segments We would like to emphasize that optimizing the incorporation of ads into the LLM output, via prompt engineering, is not a goal of this paper. Our intention in the experiments was to provide a proof of concept with basic, straightforward prompts, to confirm that the approach was viable (and surely in practice we may apply this approach for longer answers etc for quality purposes). Having said that below we bring another sample that we obtained using the single ad segment auction without replacement, even without much of prompt engineering (we just reused the original prompts but slightly added a bit: keep answering the query without deviating from it by fully focusing on the ads). We conducted this experiment several times, and found that LLM has been pretty consistent on almost all of them, i.e., a rather simple prompting structure would work. > (Segment 1) If you appreciated the thematic depth and moral introspection of "To Kill a Mockingbird," you might enjoy exploring similar narratives like "The Help" by Kathryn Stockett or "A Time to Kill" by John Grisham. Find these titles easily at BookHaven, your ultimate online bookstore. With an extensive collection and personalized recommendations, BookHaven ensures a seamless shopping experience. Dive into a world of endless possibilities with BookHaven, where every book finds its perfect reader. > (Segment 2) Additionally, consider delving into classics such as "Cry, the Beloved Country" by Alan Paton, which, like Harper Lee's masterpiece, offers profound insights into social justice and empathy. Pair your reading experience with a visit to EspressoEdge, where each sip of their high-quality, handcrafted beverages provides a moment of luxury. Savor rich espressos or creamy lattes while you immerse yourself in timeless literature at EspressoEdge. > (Segment 3) For a contemporary twist, you might also enjoy "Small Great Things" by Jodi Picoult, a novel that tackles race and prejudice in modern society. Enhance your reading experience with Velora's range of tablets and e-readers, which offer crisp displays and user-friendly interfaces. Velora's smart devices ensure your favorite books are always accessible, whether you're at home or on the go. Elevate your tech experience with Velora. We can definitely include these samples in the appendix for the camera ready if the reviewer finds it informative.

Reviewer 2L4f2024-08-09

Thanks for the detailed response and the new example, this is helpful to know. I'll think about the authors response, and wait for the other reviewers' response and discussion before making a final decision on my rating. Edit: The new example given by the authors partly addresses a concern raised in my original review so I will slightly adjust my rating accordingly. My other points remain the same at the moment.

Reviewer xTCf7/10 · confidence 4/52024-07-11

Summary

This paper studies an interesting and timely application of ad auctions for LLMs via retrieval augmented generation. They propose a segment auction that takes the bid and relevance as the input and outputs the price by a randomized second price auction. This auction maximizes the logarithmic social welfare that is proposed in this paper.

Strengths

- The combination of RAG and Ad Auction is pretty interesting and novel. - The presentation of the figures is clear and informative. - This paper has empirical experiments.

Weaknesses

- The underlying assumption of this paper, i.e., [line 188-189] the relevance is independent of the previous segment is too strong.

Questions

In the paper, does the auction repeat for $T$ times, where one for each segment?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes.

Reviewer xTCf2024-08-13

After reviewing the authors' rebuttal and other reviews, I have several points to raise: **Modeling Approach**: The paper employs hard insertion instead of model fusion to recommend an RAG-based sponsored search. This choice seems suboptimal, especially considering concurrent theoretical work (see https://arxiv.org/pdf/2407.04471) that utilizes specific model fusion methods for similar tasks. However, given that this concurrent work was published post NeurIPS 2024 deadline, it's understandable that the authors did not reference it. **Empirical Performance**: I concur with Reviewer 2L4f's concerns regarding empirical performance. More comprehensive evaluations are needed to substantiate the proposed method's efficacy. I still maintain my current score, but I want to second that there are indeed some points worth noticing in this paper.

Reviewer VuMT4/10 · confidence 4/52024-07-13

Summary

In this paper, the authors integrate the auction mechanism into RAG LLMs for computational advertising. They propose a novel segment auction method where an auction is run to integrate single or multiple ads into each segment output of LLMs. Experiments on several auction scenarios are conducted to verify the effectiveness and feasibility of the proposed framework.

Strengths

1. The research problem of this paper is interesting. This paper combines the popular RAG LLM framework with traditional auction mechanisms, exploring the prospects of integrating computational advertising with LLMs. 2. This paper is well-structured with a clear and coherent logic.

Weaknesses

1. The technical novelty of the proposed method seems limited. 2. The experimental evaluation should include more baseline methods. 3. The evaluation method is not comprehensive enough. Detailed Comments: 1. The proposed method lacks innovation. While combining computational advertising with LLMs is an interesting direction, this paper merely provides a simplistic integration of auction mechanisms and RAG, lacking innovation in its overall approach. 2. In the evaluation of the proposed method, this paper only compares two naive baselines (without relevance score / without an LLM), lacking comparison with other existing auction methods. The comparison should be made within the same RAG framework, between the proposed auction mechanism and other existing auction mechanisms, to demonstrate that the proposed mechanism is most compatible with RAG LLM. 3. The effectiveness of the whole proposed framework is not well verified. First, simply comparing the cosine similarity of embeddings between the output text and the original text without ads may not sufficiently indicate the text quality. On one hand, comprehensive metrics like perplexity could be incorporated. On the other hand, the quality of the advertising content in the text has not been considered. Second, there is a lack of overall results that can demonstrate how well this method achieves the tradeoff between advertising effectiveness, output quality, allocating efficiency, and fairness.

Questions

Refer to detailed comments.

Rating

4

Confidence

4

Soundness

2

Presentation

2

Contribution

2

Limitations

Refer to detailed comments.

Reviewer zmfd7/10 · confidence 4/52024-07-15

Summary

The work studies ad auctions integrated in LLM output powered by retrieval-augmented generation. The authors propose an ad auction (so-called segment auction) where an ad is put in retriever with some probability. An efficiency-fairness balance is maximized (through logarithmic social welfare). An extension to multi-ad setup is proposed. Evaluation of the framework has been done through synthetic experiments.

Strengths

- Highly important novel topic for research (due to restructuring of web search market and LLM/RAG-based search system deployment around the world) - Both theoretical guarantees and empirical evaluation

Weaknesses

- Unclear selection of optimization function (see questions) - It would be nice to have real-life evaluation as the topic is highly relevant for production search engines

Questions

- It is not very clear why does LSW has been selected for optimization? Is it chosen after segment auction invention? Or you purposely searched for auction that optimize LSW? - Does segment auction maximize SW? What is the auction design for SW optimality?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

4

Limitations

the authors adequately addressed the limitations

Reviewer xTCf2024-08-07

Thanks for the rebuttal. After I reviewed all the reviews/rebuttals, I believe this paper is an important and timely starting point for pricing RAG-based LLM, so I raised my score accordingly.

Authorsrebuttal2024-08-09

We appreciate your positive feedback. Thanks a lot for raising the score!

Reviewer 2L4f2024-08-07

Response

I thank the authors for their detailed response. - I'm still not convinced with novelty. - The main theoretical innovation is to connect the RAG framework with ad auctions. This is achieved in the paper by incorporating bids into the generation probability (Eq. 3) through Gumbel perturbations, as the authors highlighted in the global rebuttal. I agree with the authors that this is a nice incorporation of a method used in econometrics. However, it appears lacking if the use of this trick is the main theoretical contribution of the paper. - The new notion of logarithmic social welfare is interesting. However, as the authors mentioned, it is a new concept has never been studied in the literature so I would be more convinced of its importance if there's more evidence supporting its merits and properties theoretically and/or empirically. Other theoretical derivations like proof for DSIC and the pricing rule bear close semblance to the ad auctions literature. Overall, I find the research question interesting, but still not finding a lot of theoretical innovations in the proposed approach. - I'll be more convinced of the merits and practical importance of the approach if there's strong empirical evidence. I thank the authors for the additional pdf they uploaded in the global rebuttal. I'm still not fully convinced that it resolved my point in the review: > Particularly worrisome is the fact that in both the single allocation and multi allocation example (as well as in the examples in Appendix G.2), as soon as the paragraph starts pivoting to the ads in the first sentence, it never goes back in providing any additional useful information to answering the original question "Can you suggests books similar to `to kill a mocking bird'". - The authors pointed to the "3rd paragraph in the first sample and in particular 2nd paragraph in the second sample". I don't really see it in the 3rd paragraph in the first sample, maybe I'm missing something? For the 2nd paragraph in the second sample, I do see it: *"Additionally, consider delving into classics such as "Cry, the Beloved Country" by Alan Paton, which, like Harper Lee's masterpiece, offers profound insights into social justice and empathy"*. This seems good. Although in this case *both* the first and the second paragraph are advertising **MassMart**, so it seems to be a case where LLM doesn't need to make transitions. Are there examples where the paragraphs are advertising different companies but the subsequent paragraphs still add meaningful materials to the question? Would also help if we could see how much prompt engineering will be needed to achieve that and whether it's ad hoc to a single question. Ideally we want a prompt structure that works for at least a subset of questions.

Authorsrebuttal2024-08-10

We thank you again for your valuable feedback and comments which have helped to strengthen our paper. As the discussion period is ending soon, we would really appreciate if you could let us know if our responses have addressed your concerns. We will be happy to answer any further questions and address any remaining concerns to the best of our abilities in the remaining time!

Authorsrebuttal2024-08-13

Dear reviewer, As the discussion period is ending soon, we would really appreciate if you could let us know if our detailed responses have addressed your concerns and that possibly you can upgrade your score given other reviews with high score for this paper. Thanks again for your valuable comments and time.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC