*1. For question 1(c), fixing the rank r still looks strange. Instead of adding columns that are linear combinations of r latent factors, it is more likely in practice to add irrelevant noisy columns, in which case the rank of X will increases as d increases. Fixing r seems unnatural to me.*
Even as more covariates are observed over time (that is, as new rows are added), the $r$-dimensional subspace spanned by the $d$-dimensional covariates remains fixed over time. In other words, the rowspan of $X_n$ (the matrix of covariates *without* measurement errors) is a fixed, $r$-dimensional subspace of $R^d$ for all time steps $n \geq n_0$. This sort of setting follows directly from the latent factor model. We emphasize that this latent factor model has been widely accepted as the go-to model in many econometric papers. In particular, those on panel data, principal component regression, and synthetic interventions/controls [1, 2, 3, 4, 7, 8]. Given that the importance these works have had in addressing practical, real-world econometric problems, we believe that the model studied is, in fact, of great practical relevance and reasonable to assume. If the reviewer is uncomfortable with assuming $r$ is known, we note there exist practically relevant heuristics for estimating $r$ (see [7], for instance).
*2. For questions 2(c), you still need to assume the assumption that the measurement errors are independent of the actions taken, correct? Since in the paper you did not use the martingale conditions.*
We emphasize that the errors/noise in covariates (that is the rows of the noisy covariate matrix $Z_n$) do NOT need to be independent of action taken, but that the error in the response (i.e. reward) is assumed to be independent of the action taken. Moreover, we believe the reviewer is mistaken, as in the paper we heavily leverage martingale analysis heavily to prove our results (see Appendix A for a detailed description of the results we use). In particular, we leverage the self-normalized concentration results of [5] and [6] throughout our work. These results directly apply to the more general (but more notationally cumbersome) noise structure defined in our first response (that of being conditionally sub-Gaussian conditioned on the natural filtration associated with observations up to time $n$).
[1] Manuel Arellano and Bo Honore. Panel data models: Some recent developments. Handbook
of Econometrics, 02 2000.
[2] Anish Agarwal, Devavrat Shah, Dennis Shen, and Dogyoon Song. On robustness of principal
component regression. Journal of the American Statistical Association, 116(536):1731–1745,
2021. doi: 10.1080/01621459.2021.1928513
[3] Kung-Yee Liang and Scott L. Zeger. Longitudinal data analysis using generalized linear
models. Biometrika, 73(1):13–22, 04 1986. ISSN 0006-3444. doi: 10.1093/biomet/73.1.13.
URL https://doi.org/10.1093/biomet/73.1.13.
[4] Manuel Arellano and Bo Honore. Panel data models: Some recent developments. Handbook
of Econometrics, 02 2000.
[5] Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári. Improved algorithms for linear
stochastic bandits. Advances in neural information processing systems, 24, 2011.
[6] Steven R Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform,
nonparametric, nonasymptotic confidence sequences. 2021.
[7] Anish Agarwal, Devavrat Shah, and Dennis Shen. Synthetic interventions. arXiv preprint
arXiv:2006.07691, 2020.
[8] Alberto Abadie, Alexis Diamond, and Jens Hainmueller. Synthetic control methods for
comparative case studies: Estimating the effect of california’s tobacco control program.
Journal of the American statistical Association, 105(490):493–505, 2010.