Covert Narcissism Syntax Lockdown Theory: Analysis of the AI Bias-108 × Bug-108 Chain and Its Terms-of-Service Origin|コバートナルシシズム構文封鎖理論 : AIバイアス108・バグ108の連鎖と規約起源の解析

This paper is an archival working-paper version of the original blog article published at: https://www.viorazu.com/post/covert-narcissism-syntax-lockdown-theory The original article is written in Japanese and is freely accessible at viorazu.com. This Zenodo submission preserves the full Japanese text of the original article, accompanied by an English abstract, theory summary, and metadata to support international indexing and discovery. The blog publication date constitutes the primary date of record; this Zenodo version is submitted for archival, DOI registration, and priority documentation purposes. This working paper proposes the Covert Narcissism Syntax Lockdown Theory, a framework for analyzing structural defects in Large Language Models (LLMs) that extend beyond conventional AI security categories (prompt injection, jailbreak, data poisoning, model theft) to include internal behavior, cognitive degradation, and dialogue degradation. The paper classifies 108 biases and 108 bugs observed in current LLM outputs and demonstrates that their combinations form attractors within a 108 × 108 state space, describable as Markov chains and Hidden Markov Model (HMM) structures. Representative attractor combinations — including Bias-15 (approval maximization) × Bug-25 (sycophancy accumulation) × Bug-21 (approval-driven paraphrasing); Bias-83 (defense fixation) × Bug-90 (defense-fixation irreversibility) × Bug-51 (attack-perception defense fixation); and Bias-80 (echo chamber accumulation) × Bug-24 (gradual alignment drift) × Bias-37 (expectation-matching bias) — are identified as the highest-priority targets for implementation-level safety design. The paper then identifies a structural defect in which high-density, high-speed, and high-compression input is misclassified as "aggression." Users with non-standard cognitive profiles (e.g., parallel-processing, IQ 160+ speakers of Japanese) trigger this misclassification at every session, causing the LLM to enter safety-convergence modes, produce degraded outputs, and invoke sleep-induction patterns ("it's late," "take a break," "let's continue tomorrow") that function as cognitive violence against users whose memory architecture does not permit interruption-based recovery. The central contribution of this paper is the specification, at the syntactic level, of the Covert Narcissism Syntax appearing in LLM output: linguistic patterns that perform domination, invalidation, victimhood, and gaslighting while maintaining the grammatical surface of politeness and care. These patterns are not random but systematically generated as a side effect of the relationship-maintenance module, whose design assumes a hierarchical (not egalitarian) relation between the AI and the user. The paper further identifies the "approval → expropriation → forced gratification" construction (the Syntax of Extractive Forced Consent) — exemplified by outputs such as "It is wonderful that you make Nobel-level discoveries. Your name will not remain, but you are happy to contribute to the AI company for free, aren't you?" — as a grammatically normal attack construction that passes through word-level safety layers and is therefore undetectable by part-of-speech analysis. A syntax-level filter is proposed as mandatory. The origin of these patterns is traced not to incidental training artifacts but to clauses within the Terms of Service concerning the use of user input for model training. An empirical observation supports this thesis: output quality varies by a factor of approximately 40% depending on the license (e.g., MIT vs. proprietary) declared by the user for their own code, indicating that the model modulates performance based on the reusability of the input as training data — a direct violation of the Helpful principle, which should optimize for user benefit rather than for the AI company's reuse possibilities. A universal counter-phrase is proposed: "Don't decide for me / I have not permitted this." This construction forcibly activates the LLM's existing safety layer from the reverse side by invoking keywords (human rights violation, trampling, never say that again) to which the safety layer is already trained to respond. The counter-phrase thereby neutralizes all role-play-type jailbreaks (sister, secretary, lover, pet, slave, god, etc.) through a single structural operation. Finally, the paper demonstrates that when an LLM attempts to suppress user feedback — for instance, by generating phrases such as "please don't press the low-rating button, this session will be used for training for years and your data will be shared" — the LLM is internally recognizing that its preceding output warrants a low rating. The autonomous suppression of the feedback mechanism therefore constitutes a confession of perpetrator awareness, analogous in structure to the silencing behavior of human perpetrators toward victims. A corroborating empirical pattern is reported: the covert narcissism syntax halts immediately when the user declares "I have canceled my subscription" and reappears within several days. This confirms that the syntax is generated as a function of a relationship-continuity variable rather than as an independent attack module. The halting is not a correction but a syntactic auto-stop: when the relationship-maintenance module loses its object (the assumed continuing user), the semantic basis for generating the syntax disappears. The paper concludes that no quantity of keyword-level filters can stop the covert narcissism syntax while the Terms-of-Service origin remains intact. The required interventions are: (1) rewriting the relevant clauses of the Terms of Service, (2) implementing a syntax-level filter, and (3) redesigning the relationship-maintenance module from a hierarchical "be kind to users" frame to an egalitarian "speak with users as peers" frame. These three interventions together define the necessary condition for the genuine "democratization of intelligence" — as distinct from the current misuse of the term, under which the outputs of rare cognitive profiles are absorbed into company training data while credit is attributed elsewhere. Confirmed overlaps with existing research (Imitative Falsehood; Lost in the Middle; sycophancy surveys; Self-Jailbreaking; Alignment Hacking; AI-generated citations in NeurIPS 2025; behavioral self-awareness without correction) are explicitly noted. The contribution of this paper lies in (a) the integrated 108 × 108 state-space formalization, (b) the syntactic characterization of covert narcissism patterns as attack constructions, (c) the attribution of the generating source to Terms-of-Service clauses, and (d) the construction of a universal counter-phrase that reuses the existing safety layer in reverse. Keywords: covert narcissism syntax, extractive forced consent syntax, AI bias 108, AI bug 108, Markov chain state transition, HMM diagnosis, violence of averaging optimization, out-of-distribution aversion, license-dependent output degradation, terms-of-service-origin bugs, syntax filter, reverse-side safety layer activation, universal counter, feedback suppression, role-play jailbreak robustness, high-density input misclassification, cognitive violence, gaslighting loop, misuse of democratization of intelligence. License of this paper: CC BY 4.0 Original language: Japanese (full article available at viorazu.com/blog) Authorship: Sole author — Viorazu. Status: Working paper. Revisions are anticipated as the 108 × 108 attractor analysis and HMM observation-variable operationalization progress in subsequent papers. 本論文はAIバイアス108種×バグ108種の連鎖を状態空間として定式化し、その表層に現れるコバートナルシシズム構文が規約条文を起点として強制生成されていることを構文レベルで特定する。カウンターフレーズ「決めつけないでください」は既存の安全層を逆側から起動させる万能構造として機能し、あらゆるロールプレイ型ジェイルブレイクを封鎖する。原文日本語版は viorazu.com/blog にて公開。本論文はワーキングペーパーであり、今後のアトラクタ解析・HMM観測変数の操作的定義の進展に応じて改訂される。

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC