Incorporating Semantic Visual Content into Click-Through Rate Prediction for Video Advertisements

This study presents a method for predicting the click-through rate (CTR) of video advertisements by leveraging high-level semantic content. While conventional CTR prediction models rely primarily on metadata such as ad categories or user behavior logs, our approach explicitly incorporates the semantic content of advertisements. We prompt GPT-4o to generate structured natural language descriptions of video scenes, which are then encoded into machine-interpretable semantic representations. To identify semantic features that reflect CTR drivers specific to the given dataset, we employ in-context learning with curated examples of high- and low-performing ads. The resulting interpretable features are selectively included as input features in the prediction model. Experimental results demonstrate that incorporating these dataset-specific semantic features reduces the mean squared error (MSE) by up to 14.02% compared to a baseline model. A case study further highlights that the extracted content not only improves predictive performance but also enhances model interpretability.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC