DeepShop: A Benchmark for Deep Research Shopping Agents

Large language models (LLMs) increasingly mediate consumer brand evaluation in search and agentic commerce. This study demonstrates that LLMs systematically collapse brand perception to Experiential and Economic dimensions, rendering brands differentiated on narrative, ideological, cultural, or temporal dimensions metameric – structurally different but functionally equivalent in AI recommendations. Using PRISM-B across 21,350 API calls to 24 models spanning nine cultural traditions and 15 languages, a Dimensional Collapse Index (DCI; uniform baseline = .250) is computed. Mean DCI equals .291 for global brands ($p = .017$) and .357 cross-culturally ($d = 3.449$). Cross-model convergence is near-perfect (cosine = .977). Local brands collapse 25% more severely than global brands ($d = .878$, in five FMCG pairs tested). Native-language prompting is null in home markets ($p = .716$) but reduces collapse up to 9.5 DCI points in foreign-market frames ($p = .002$). Geopolitical context independently modulates weights ($p < .001$). Structured Brand Function specifications recover approximately 20% of lost dimensionality. Collapse is stable across model versions. For search advertising, these findings reframe strategy from platform selection to dimensional defensibility: advertisers must encode soft brand dimensions in verifiable, machine-readable form or risk effective erasure by LLM intermediaries. Patterns reflect the text-conditioned LLM observers tested here. Includes paper.yaml (Paper Spec v0.1.0) – a machine-readable specification of the paper's claims, assumptions, and dependencies. See https://github.com/spectralbranding/paper-spec for the standard.

Paper

Similar papers

© 2026 NYSGPT2525 LLC