Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration

Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whether models can still solve code tasks after this transformation. We examine a different systems question: how commercial APIs count the resulting requests. We present a reproducible measurement case study of provider-reported input tokens for raw source text and a compact rendered-image representation. The benchmark pairs requests across five programming languages, nine source lengths from 20 to 2,000 lines, and 15 available model aliases exposed by Anthropic, OpenAI, and Google Vertex AI. These aliases collapse to approximately five distinct accounting signatures and are not independent model replications. Across 675 complete text/image pairs, aggregate image-to-text ratios are 0.135, 0.194, and 0.242, corresponding to reported input-token reductions of 86.5\%, 80.6\%, and 75.8\%, respectively. These totals conceal materially different break-even behavior: Anthropic and OpenAI images receive lower counts at every tested size, while Gemini images require 6.95 times as many tokens at 20 lines and cross below text only at 200 lines in the aggregate. A targeted audit also reproduces non-monotonic Gemini image accounting across a page boundary. This study measures black-box request accounting for one compact rendering pipeline. It does not measure semantic fidelity, task accuracy, latency, monetary cost, or coding-agent efficiency. We release the scripts, revision-pinned corpus specification, raw usage records, validators, and deterministic analysis needed to reproduce and extend the study.

Paper

References (15)

05Empirical Standards for Software Engineering Research2020
062022. Language Modelling with Pixels. (2022)arXiv
072025. Glyph:ScalingContextWindowsviaVisual-Text Compression.
082026. CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding. (2026)arXiv
092023. Code Llama: Open Foundation Models for Code. (2023)arXiv
102023. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. (2023arXiv
112025. VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression? (2025)arXiv
12a targeted modality audit demonstrating that reported image usage can change non-monotonically at a page boundary

Scroll for more · 3 remaining

Similar papers

© 2026 NYSGPT2525 LLC