Relational Consistency Probing: Protocol Design, Pilot Findings, and Two Instructive Failures from a Five-Model Experiment

I designed a black-box protocol that measures how language models' similarity judgments shift under cultural framing. The protocol uses only API access. I pre-registered the experiment on the Open Science Framework and ran it against five deployed models: Claude Sonnet 4, GPT-4o, Gemini 2.5 Flash, Grok, and Llama 3.3 70B. The inventory covered 18 concepts across physical, institutional, and moral domains under seven framings. The pre-registered hypothesis failed: institutional drift exceeded moral drift in all four interpretable models, and the pre-registered permutation test was structurally underpowered at the chosen inventory size. What survived: the physical control domain held with moderate-to-large effect sizes, confirming the protocol discriminates culturally invariant concepts from culturally loaded ones. Four distinct nonsense-compliance profiles emerged across the models, and every model except one constructed coherent moral frameworks from an irrelevant weather preamble.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC