Data accompanying Yakura, Lopez-Lopez, Brinkmann et al., "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication" (arXiv:2409.01754). The reproduction and analysis code is archived separately (Zenodo 10.5281/zenodo.21296093). This deposit contains three files: manuscript-data-figures.tar.gz (~0.4 GB) — precomputed figure-source outputs: synthetic-control results, in-space placebo time series, and Bayesian change-point (Stan) results for each manuscript figure and its robustness variants (Fig. 1, Fig. 2, Fig. 3, the YouTube appendix, and the Supplementary figures). Extracts into observational_pipeline/runs/; sufficient to render the paper figures without re-running the pipeline. manuscript-data-substrate.tar.gz (~2.0 GB) — the analysis substrate: per-episode word-count matrices (counts_50k.parquet, raw; counts_50k_audited_merged.parquet, sense-audited) and the trained spontaneity-classifier weights (model_8_128_4.pt). Input for re-running the sense audit, synthetic-control, and change-point analyses from counts. manuscript-data-ngrams.parquet (~5.0 GB) — monthly podcast n-gram frequency tables, one row per (analysis group, month, n-gram): surface-form counts for n = 1–5 with a ≥5-occurrence cutoff (~445 million rows). Derived aggregate lexical statistics. See REPRODUCE.md in the code repository for how each file is used. Verbatim transcripts are not included and are available through an institutional data-use agreement.