Summary
This work studies the random forests stablity for regression problem, and the authors presents theoretical analysis on the upper and lower boudns for the coverage probability of prediction intervals constructed from the out-of-bag error of random forests. The theoretical guarantee is based on a light-tail assumption of the marginal distribution of the squared response.
------------------after response---------------------------------
After reading the authors' response, I do not think the authors answer my concerns, in particularly for the novelties and significances.
As a theoreical work, it is very important to evaluate from theoretical novelties and techniques, while I find some inremental results based on well-known techniques.
I do not find any experiments, and I do not know why the authors could claim that "Our work applies to many variants of random forests ... which makes it particularly relevant in theory and practice".
Strengths
1) It is an interesting problem on the theoretical understanding of random forests.
2) Some theoretical results on the convergence probability of prediction intervals constructed from the out-of-bag error of random forests.
3) Limited theoretical techinical contributions
Weaknesses
1) The problem is not very clear. The authors should first present the studied problem, i.e. the original ranfom forests for regression, or randomf forest interval, or prediction intervals constructed from the out-of-bag error of random forests. It is very confused to understand the main contributions in the current submission and relevant work. For completeness, it would be better to present the detailed algorithm, rather than finding some other research work for other readers.
2) The main conclusions are not clear. The main contribution of this work is the stability of random forests in Theorem 1. Generally, a theoretical work concerns seriously the convergence rate of stability, and how tight of this rate. It would be better to present the specific expression for \nu_{n,B}, and make necessary discussions. What's factors affect the stability rate?
3) Some important definitions and notions are missing. For example, where is the definition of "light tail", whichi is the basic assumption in main theoretical results. How to characaterize light tail and its relevant factors.
4) The authros should clearify the novelty and significance of the main results, for example, how about the theoretical new insghts on the technical proof in this work. As an pure theoretical problem, it is importnat to present some new technical proof, rather than simple extension of the current techiniques. What's the siginificance of the main results, is it possible to present some pratical guidance and suggest some new algorithms.
5) The authors should have a good background on the theoretical analysis on random forests, for example,
G. Biau, L. Devroye, G. Lugosi, Consistency of random forests and other averaging classifiers, JMLR 2008.
M. Denil, D. Matheson, N. De Freitas, Narrowing the gap: random forests in theory and in practice, ICLM2014.
W. Gao, F. Xu and Z.-H. Zhou. Towards convergence rate analysis of random forests for classification. AIJ, 2022.
Questions
1) The problem is not very clear. The authors should first present the studied problem, i.e. the original ranfom forests for regression, or randomf forest interval, or prediction intervals constructed from the out-of-bag error of random forests. It is very confused to understand the main contributions in the current submission and relevant work. For completeness, it would be better to present the detailed algorithm, rather than finding some other research work for other readers.
2) The main conclusions are not clear. The main contribution of this work is the stability of random forests in Theorem 1. Generally, a theoretical work concerns seriously the convergence rate of stability, and how tight of this rate. It would be better to present the specific expression for \nu_{n,B}, and make necessary discussions. What's factors affect the stability rate?
3) Some important definitions and notions are missing. For example, where is the definition of "light tail", whichi is the basic assumption in main theoretical results. How to characaterize light tail and its relevant factors.
4) The authros should clearify the novelty and significance of the main results, for example, how about the theoretical new insghts on the technical proof in this work. As an pure theoretical problem, it is importnat to present some new technical proof, rather than simple extension of the current techiniques. What's the siginificance of the main results, is it possible to present some pratical guidance and suggest some new algorithms.
Rating
2: Strong Reject: For instance, a paper with major technical flaws, and/or poor evaluation, limited impact, poor reproducibility and mostly unaddressed ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Limitations
This is a pure theoretical work