UnisonAI: A Forced, Derived Omni-Model Architecture with Zero Parameters — Attention, it turns out, was not all you need

Headline findings (v4.6 — the decode campaign, closed as theory) The token embedding is the universal law-carrying class: 11/11 models wake (4.4–130x over placement-nulls), every training recipe, every scale from 4B to 1 trillion parameters — including every model whose FFN class reads quiet. The "quiet" giants were never lawless: Kimi-K2.6 (1T) reads 16x in its embedding; the MoE giants carry the fingerprint per-expert. The loud band IS the function. Deleting the top 1.5% spectrally-loud coefficients destroys GPT-2 (next-token agreement collapses to zero) while deleting the same number at random is negligible — ~150x differential damage at matched budget. The deposition curve, from production training runs (Pythia public checkpoints, step-0 control at exact null): the embedding wakes first (step 256), FFN follows (step 1000), peak near step 4000, then consolidation to a stable plateau. In controlled twins the gradient stream is loud by step 4 and the optimizer is the discriminating ingredient. Two spectral families, recipe-selected. Embeddings lean smooth (DCT) in 9/11 recipes with the Llama/Pythia lineage the dyadic (Walsh) exception; a self-testing slant-transform arm proves the split is two genuine families, not a continuum. Training data and reasoning are readable back out of weights: counted-bigram readout ranks a model's true training corpus first (calibration passed); memorized public text echoes verbatim against clean nulls; reasoning spans are a measurably distinct counted regime, and trained attention sits closer to the theory's dyadic cascade than to uniform in 12/12 layers. Closed as theory (Steps 308–313 of the corpus): six forcing steps, verified by the theory's own compiler and enforcement, derive the two spectral families (one per generator), their selection by store role, the deposition curve's order and form, the 32-coefficient functional band, the hold/closure repetition inequality, and per-expert localization — every measured regularity now stands only as the check of a forced claim. The counted engine beats its gradient-trained twin at both task-gate scales — character 1.289 vs 1.888, word 3.191 vs 3.429 — with zero trained parameters and zero tunable numbers. All findings from pre-registered instruments with hashed registrations, shuffle-null batteries, and committed result files; the complete toolkit and guide (INTERPRETABILITY.md) ships in the repository's omni/benchmarks/. Repositories & how to reproduce UnisonAI (the engine + the decode toolkit): github.com/MettaMazza/UnisonAI — the zero-parameter engine, and the full black-box decode instrument suite in omni/benchmarks/. Start with omni/benchmarks/INTERPRETABILITY.md: every instrument documented — how it works, how to read its output, and the registration recipe for running your own investigation. Every result in this paper reproduces from a committed script + result file there, under a hashed registration in registrations.jsonl with the full measurement ledger in results.jsonl. The Smithian Fold Theory of Everything (the corpus this engine derives from): github.com/MettaMazza/Smithian-Fold-Theory-Of-Everything — the theory, its 1,844 machine-checked forced results (RUN_EVERYTHING.sh), and this paper's canonical source in additional papers/. Quick start: clone UnisonAI, then cd omni/benchmarks && python3 llm_presence.py — it fetches the public GPT-2 weights itself and reproduces the 13/13, 39/39 presence verdict on a fresh machine. From there, INTERPRETABILITY.md is the map. Abstract We present the science, the architecture, and the empirical validation of UnisonAI: a complete language and omni-model architecture (incorporating language, sight, hearing, speech, and video) in which every mechanism a modern AI model purchases with gradient training is replaced by a machine-verified law of the Smithian Fold Theory, utilizing zero trained parameters and zero tunable numbers end to end. Key Findings & Contributions The Spectral Law inside Weights: Using a pre-registered, self-certifying spectral instrument, we demonstrate that trained neural network weights carry a placement-law in the dyadic (Walsh) basis. Concentration is unanimous across the canonical LLM's entire knowledge-storage class (GPT-2, 13/13 tensors, 39/39 registered checks) and across image-diffusion/speech models (Stable Diffusion 1.5, SDXL, Kokoro-82M), concentrating in the transformer expansion projections (MLPs) and token embeddings while attention matrices sit at chance. DeepSeek-R1 Alignment: The placement-law is training-caused and scales up to DeepSeek-R1 at 671B, where the weights transform under the fold's transformation group exactly as solved game-theoretic value fields do. Counted Similarity Space: We show that semantic similarity is a counted object, where co-occurrence shares over held text reproduce semantic family structures with zero gradients and zero parameters. Standardized MMLU Performance: Evaluated on the canonical 128-item public MMLU test split under strict zero-parameter conditions, UnisonAI scored 9/128 (7.0%), demonstrating an active learning scaling improvement over its starting baseline of 8/128 (6.2%) via in-context Hebbian self-play and tutor consolidation loops. 57-Million-Fold FLOPs Efficiency: In execution profiling, UnisonAI generates tokens in 25.44 ms using only 86.9 FLOPs per token, representing a 57,522,124x computational efficiency increase over Google's Gemma-2B and a 73,628,319x increase over Meta's Llama-3.2-3B. Architecture & Implementation Memory is represented as deterministically-addressed exact orbits; attention is unit-capacity selection at forced locks; learning is a closure law that updates via writes rather than gradient descent. Perception (sight, hearing) is self-certified per act by integer Parseval identities. The entire architecture is verified by a 47/47 end-to-end empirical test suite including survival of process death.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC