Optimal testing using combined test statistics across independent studies

Combining test statistics from independent trials or experiments is a popular method of meta-analysis. However, there is very limited theoretical understanding of the power of the combined test, especially in high-dimensional models considering composite hypotheses tests. We derive a mathematical framework to study standard {meta-analysis} testing approaches in the context of the many normal means model, which serves as the platform to investigate more complex models. We introduce a natural and mild restriction on the meta-level combination functions of the local trials. This allows us to mathematically quantify the cost of compressing $m$ trials into real-valued test statistics and combining these. We then derive minimax lower and matching upper bounds for the separation rates of standard combination methods for e.g. p-values and e-values, quantifying the loss relative to using the full, pooled data. We observe an elbow effect, revealing that in certain cases combining the locally optimal tests in each trial results in a sub-optimal {meta-analysis} method and develop approaches to achieve the global optima. We also explore the possible gains of allowing limited coordination between the trial designs. Our results connect meta-analysis with bandwidth constraint distributed inference and build on recent information theoretic developments in the latter field.

Paper

References (54)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer ThPF6/10 · confidence 3/52023-06-29

Summary

[Update1: During the rebuttal, I updated my score from 5 to 6. The reason is that I want to stronger weigh in that the paper is technically solid. My concern whether NeurIPS is a perfect fit for this paper remains, and I ask the AC to judge that part.] The paper studies aggregation strategies for combining test statistics obtained from independent experiments. It is intuitively clear that the optimal strategy would be to aggregate the whole *data* and then compute one test statistic on the entirety of the data. However, the authors make the assumption that from each experiment one can only use a single real number (aka test statistic) for the final analysis. The setting they consider is that they have observations of random variables $$ X^{(j)} = f + \frac{1}{\sqrt{n}} Z^{(j)}, $$ where $Z^{(j)}$ are independent $d$ dimensional normals and $n\in \mathbb{N}$ and additional parameter. The null hypothesis is $f=0$ which is tested against $\|f\|_2 \geq \rho$ for some $\rho >0$. If pooling the data was allowed, the chi-square test is optimal for this setting. They derive theoretical results that give two different regimes: - When $m \lesssim d^2$ the optimal rate can be achieved by combining the individually optimal test statistics $\|\sqrt(n) X^{(j)}\|_2^2$, which is undirected. - When $m \gtrsim d^2$ then taking the directionality into account leads to better rates. Two examples they discuss are (phrasing in my own words that the authors might adapt): - Split the observations into $d$ groups of size $\sim m/d$ and test with the $i$-th group whether the null hypothesis holds for the $i$-th dimension. Then combine those. - Choose a 1-dimensional projection of the $d$-dimensional data uniformly at random. Aggregate all observations along this projections (now these are 1-dimensional values, as required in the setting) and combine those directly to test the null hypothesis along the projected direction. (Note that this requires shared randomness between all experiments). Finally the authors demonstrate their theoretical findings with toy experiments and numerically confirm their findings. Overall I enjoyed reading the paper and was able to learn something new. Thank you to the authors.

Strengths

- The paper is extremely well written and has "textbook" quality. It focuses on a simple toy problem and provides an exhaustive analysis thereof. - Combining outcomes from different experiments is certainly an important problem in statistics. - The paper seems original, however, admittedly I am unable to judge to what extent similar analyses exist in the stats literature. - The theory and experiments are in perfect accordance and complement each other. - The paper provides a good overview of existing methods to combine test statistics and discusses both $p$- and $e$-values as examples. They relate existing methods to their contribution.

Weaknesses

1. My first point concerns the significance of the approach. In the introductions the paper states > Given multiple data sets relating to the same hypothesis, one would like to combine the evidence. Sometimes, the full data sets are not available (e.g. due to privacy or proprietary reasons) or difficult to combine directly (e.g. due to the different experimental or observational setups). In such cases, the analysis must be carried out on the basis of the published results for each of the studies. - This motivation reads that each study publishes a test statistic without knowing of the others. I thus think that the second type of approaches the authors provide does not fall in this category. - The second type of approaches require that the "meta-analyser" can make some choices of what statistic the individual studies report. Hence *somewhere* all data must still be available. It is unclear why we cannot use that then. I think this needs more motivation. - I presume that the first approach of combining the undirected tests has been studied exhaustively. 2. While the paper is relevant for machine learning in general, it is in itself a very statistical paper. No learning involved etc. Hence, I am posing the question whether NeurIPS is the appropriate venue. I would see a stats journal as a better fit. 3. I think the paper could be a bit more accessible if a bit more intuition about the approaches is provided. I wrote my understanding in the summary, maybe the authors want to correct/modify this and include something like that.

Questions

Please comment my first two concerns above in the rebuttal. Note that my reservation against acceptance is based on those comments. And I am happy to increase my score after the rebuttal if I am convinced otherwise. Minor: - l 124 "where" --> "were" - l. 339 "form" --> "from"

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

2 fair

Limitations

This paper only studies a simple toy problem. Limitation of my review: I did not read the appendix.

Reviewer GjrH5/10 · confidence 3/52023-07-02

Summary

This theory paper provides a minimax lower and upper bounds for the testing risk (sum of Type-I and Type-II errors) for different combination methods in the specific setting of many normal means model. With the testing goal is to detect the presence or absence of the signal component in this normal means model, their results show several combination methods of test statistics (e.g. aggregate p-values and e-values) cannot consitently detect signals below a certain threshold that depend on the number of trials, samples and dimensions of the problem.

Strengths

* The theoretical results are sound and based on several established techniques in distributed testing, assuming several assumptions hold true. These are proof to be indeed the case in the Appendix for the test statistics combination methods proposed in Section 3 of the main text. * The paper is mostly well-written.

Weaknesses

1. In general I think the phrasing of on the paper's contributions could be make clearer in the last part of the introduction section. 2. The authors should discuss more on the relevance of the setting -- the many normal means model -- in some more concrete applications. It is arguable that although this is a theoretical paper, the theories inside it is an attempt to quantify a very practical problem of meta learning. I see a lack of evidence for the popularity of many normal means in practice. 3. Slightly related but not as equally importance, but the authors should have acknowledged that a limitation of their work is that the theoretical results only hold with many normal means model assumption. 4. Experimental results could include more settings with a variety number of of sizes/dimensions (perhaps in the appendix) to support the theory.

Questions

* I have only skimmed through the proof in the appendix, so this might be my mistake, but I do not see the appearance of the $\epsilon$ term for binary approximation of the statistics in the main results. Could the authors clarify on this point?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

* See weaknesses.

Reviewer vLRc7/10 · confidence 2/52023-07-03

Summary

Authors study a problem of optimal combination of p-values in a meta-analysis context. Specifically, they focused on characterizing the minimax separation rate for a family of "smooth-ish" combination methods that aggregate p-values (or e-values). They show that: * The family contains a lot of methods that are used in practice. * The separation rate for methods in the family is nearly optimal. * Optimal combination method depends on the problem setting (sample size, number of p-vals, dimension of the problem); they describe practical consequences. Furthermore, they explored how the rate can be improved by allowing coordination between the experiments that generate the p-values.

Strengths

The paper is written in a clear and without excessive statistical jargon. For instance, the concept of separation rate is introduced and explained in a simple and intuitive way, allowing readers to understand the results presented in the paper (in contrast to the mat description found in "Non-asymptotic minimax rates of testing in signal detection"). The significance of the results is solidly established on two grounds: derivation of a separation rate and practical guidance for method selection based on n, m, and d (Section 2.2 and 2.3 add extra value). From my (limited) understanding of the literature, these results are both novel and original.

Weaknesses

I am genuinely surprised that this problem has not been previously studied. While preparing the review, I came across similar/related results concerning the Family-Wise Error Rate (FWER) in the paper "Family-Wise Separation Rates for multiple testing." However, the combination methods studied are different (Holm–Bonferroni procedure type). I'd suggest a more comprehensive and thorough discussion of the existing literature on minmax testing for multiple hypothesis testing is added to the paper.

Questions

See above

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

yes

Reviewer MHU25/10 · confidence 2/52023-07-24

Summary

The paper addresses the problem of combining test statistics from multiple independent studies in the context of null-hypothesis significance testing. The authors derive a mathematical framework to quantify the cost of compressing multiple independent trials of a study into one real-valued test statistics, and they derive minimax lower and matching upper bounds for the testing errors. The many normal means models is used as toy example.

Strengths

The paper addresses the problem of combining the results of multiple empirical studies towards one common hypothesis, an important problem in meta-analysis.

Weaknesses

This submission seems to be out of the scope of NeurIPS. While the authors draw a connection between meta-analysis as used in statistics and meta-learning as used in the context of machine learning, this connection is not clarified further. For the rest of the paper, the authors seem to focus on the problem of meta-analysis. It was challenging to read to paper as it lacks clarity in the introduction. To improve clarity, I suggest to start the introduction by clearly stating the problem that will be addressed in the paper and clearly introducing the terms used in the text. For example, in line 23, the authors introduce “meta-analysis”, (probably) referring to the technique of combining the results of multiple scientific empirical studies and then set it equal to “meta-learning”, referring to the machine learning technique of improving a learning algorithm to perform well multiple tasks. I did not go through the technical parts of the paper in all detail. However, there are several statements in the paper that appear to be misleading, e.g., - line 68: “… includes many standard meta-learning techniques, for instance the standard p-value combination methods […]”. I am not aware of p-value combination being a standard technique in meta-learning. - line 313: “Common examples of e-values are Bayes factors and likelihood ratios.” e-values and Bayes factors are closely related, but Bayes factors or likelihood ratios are not e-values, see https://arxiv.org/abs/1912.06116 Appendix A for clarifications.

Questions

Can you clarify to connection between "meta-analysis" and "meta-learning" (meaning the approach reviewed in your reference number [14]?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

The authors do not discuss limitations or potential negative impact of their work. Although being a more technical paper, I think it would have been appropriate to discuss the limitations of the meta-analysis approach as a whole.

Reviewer ThPF2023-08-15

Is Meta-Learning vs. Meta-Analysis clarified?

Dear reviewer colleague MHU2, in order to get a final judgment of the paper, I would like to understand whether your concerns have been addressed by the rebuttal or not? I share your definition of "Meta-Learning" and also think that at least within the NeurIPS community using "meta-analysis" is much more appropriate. But the authors promised to change that and I do not see any further problem with it. Do you? If it was clarified, then frankly I feel that a low score of 3 with a (quite high) confidence of 3 is not adequately reflected in your review. I think there is about a day left to ask further questions to the authors. I would like to be able to read their answer, if there are any questions still. Thank you!

Reviewer qxf16/10 · confidence 3/52023-07-25

Summary

The paper lies in the context of meta-analysis of multidimensional models. Usually, meta-analyses are performed by combining p-values or e-values. In both cases, the statistical power is not well-known. The authors provide a constrained framework of many means normal model. Based on this framework, they derive lower (theorem 1) and upper (theorem 2, for "rate optimal methods") bounds for testing methods. Theorem 3 takes advantage of possible shared randomness between trials, especially when the dimension is small relative to the number of trials considered. The authors then compare the performances of rate optimal methods (described in section 2.1), chi-square test on pooled data and single trial approach on simulated datasets.

Strengths

The paper is well-written and organised. It provides a good overview of state-of-the-art combination techniques for meta-analyses and identifies the lack of knowledge about their relative power. The paper derives bounds for the testing errors by introducing a principled mathematical framework based on multidimensional models, where a loss in power is expected. It also gives insights into rate optimal combination methods and the effects of sharing randomness between trials. The simulation study provides some results on comparing combinations methods for meta-analysis, which is not properly addressed in the current literature. About the supplementary material: Proofs as well as an R script to reproduce the simulation study are provided.

Weaknesses

The main weakness of this paper is the clarity of the mathematical developments. The framework implies several assumptions that could be explained more. The 13-page-long supplementary material provides proof of the theorems but is sometimes hard to follow. In the theorems and their proofs, sometimes arbitrary values are chosen (for $\alpha$ and $\beta$ notably) but it seems to make the reasoning more confusing. Also, when running the R script, the following error is returned: Error: object 'dat_long' not found

Questions

- $\mathbb{E}_0$ is first used in line 141 but not introduced beforehand. Is it the expectation under the null hypothesis? - I understand that Assumption 1 aims at restricting the values of S. Is it possible to give a small interpretation of the assumption? - Theorem 3 indicates that "there exists a constant $C_\beta$" but the formula indicates "$C_{\alpha,\beta}$". Is it a typo error? - In theorems 1 and 3, arbitrary values are used for $\alpha$ and $\beta$, not in theorem 2. Why make the choice of using these values and not giving general results for the corresponding intervals of validity? - A similar remark on the proofs, for example in proof A.1. Why use $\kappa_{1/10}$ and $\kappa_{1/8}$? - It might also be more comprehensible to explicitly add the results taken from the literature and used in the proofs. - The provided R script needs to be reviewed. When running it, I get the following error is returned: "Error: object 'dat_long' not found" I understand that the chosen mathematical framework provides lower and upper bounds for testing errors. The paper describes some meta-learning techniques that attain these bounds. I am not sure how the simulation study demonstrates the theoretical results. Is it by comparing these meta-learning techniques to the "chi-squared pooled" approach and the single trial approach? The indicator of performance for "optimal testing" is the ROC curve. Would it be interesting to consider other criteria, such as sensitivity, specificity, precision, or F-score? Overall, I encourage the authors to add more explanations and interpretations, especially in the mathematical development part. Note that the 9-page limit does not include references. It might also be worth submitting the paper to another journal where the format might be more adequate.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors have delimited the framework of their contribution by providing constraints inherent to their model.

Reviewer aj9v6/10 · confidence 4/52023-07-31

Summary

The paper considers methods to aggregate test statistics from different, independant, sources, in order to construct an aggregated test with hopefully more power. The key contribution of the paper is the study of the minimal treatment effect which can be detected in a standard gaussian noise setting, for which they obtain minimax rates. These rates exhibit an elbow effect when the number of aggregated statistics (m) is close to the square of the dimension of the signal (d^2), which the authors relate to the use of the signal direction in the test (when $m < d^2$, the standard chi2 test would give near optimal power, while for $m > d^2$, the test statistics must encode directional information if optimal power is to be obtained). These rates bring two insights: First, aggregating one dimensional statistics in a multi dimensional setting comes at a price. Second, there is no single optimal aggregating method.

Strengths

Overall, the paper is well presented and obtains conclusive results in the scope considered, in the form of minimax rates. These rates justify previous empirical insights on aggregated testing strategies, notably the need for different aggregating strategies depending on the number of tests and the dimension of the problem. Methods achieving the rates (up to a log factor) are specified. As far as I could assess, the mathematical proofs are, up to small typos (see weakness), correct. The presentation of the main results in section 2 can be easily followed (minimax rates in the general case, optimal combination methods then improved minimax rate using coordination between tests).

Weaknesses

The proof in the appendix suffers from some small typos. Notably, I believe that in equation (S.1), the $2\epsilon$ term should be $\epsilon$ (or $\epsilon< \frac{1}{2}\left(\kappa_{1/10} - \kappa_{1/8}\right)$ in the definition of $\epsilon$), while in line 538, the conclusion of Markov's inequality is that $D^c$, not $D$, has mass less than $1/64$. The methodology used to obtain Figure 1. could be improved. Notably, the Roc Curves for Chi-square combined and Chi-square pooled should not exhibit any randomness, since these two curves can be computed in closed form using the cumulative distribution functions of the chi2 square and non central chi square. If numerical approximations are to be used, it could be possible to obtain curves exhibiting much less noise by increasing the number of repeats and recycling them for all FPR (I could obtain curves exhibiting little to no noise robustly using 10 000 repeats and 100 FDR in less than <2s on my personal computer, so computation time is not an issue). Moreover, the way the $f_i$ are drawn, using Rademacher random variables, might have an impact on the directional methods. While this might or might not be the case, I would suggest recomputing the curve, drawing a random f uniformly on the sphere.

Questions

Could the methodology be extended beyond the current setting? Notably, is there any natural generalisation in the case where the noise level $\sigma = 1/\sqrt{n}$ can no longer be assumed to be identical ? Is there any explanation of the results in terms of the distribution of the p-values under the alternative hypothesis? If so, is there any insight on the best way to aggregate a given set of statistics (instead of considering the best way to aggregate the best statistics for a given m, d)?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The paper derives optimal test aggregation strategies in the context of gaussian noised signals. A first limitation of the paper is that it is assumed that the sample size $n$ considered in each collected test is identical. This can barely be expected in practice, and as such, insight about the impact of uneven tests would be welcome (i.e., does $mn$ translate into $\sum_{i=1}^m n_i$?, or rather $m \min(n_i)$?). This issue is not mentionned in the paper. Another limitation not mentionned is the fact that, in practice, the test statistics obtained from independent trials are not chosen, but set, and as such, Stouffer's method, which attain the minimax rates when $m>d^2$, is not implementable. In most settings where it could be implemented, the whole $X_i$ information would be known, and therefore the dimension reduction issue would not occur. For this reason, the best rate achievable with realistic one dimensional statistics is of particular interest. Unfortunatly, this is left to the appendix, in Theorem 4.

Reviewer ThPF2023-08-10

Thanks for the rebuttal

I thank the authors for their reply. I do not have any further questions. I encourage the authors to careful write out their distinction between a pure meta-analysis of, say, p-values and an analysis that allows for coordination between the trials. While the first one is clear, I believe the latter needs more motivation. Nevertheless after reading the other reviews and the rebuttals, **I will increase my score from 5 to 6.** I think the paper is a technically very solid contribution. [Note: ~It seems that currently I am unable to adjust my score. But I will do so later.~ Review is now updated]

Authorsrebuttal2023-08-10

reply

We thank the Referee for the reply and the suggestion, we will follow it in the revision. We also appreciate that the Reviewer went through all the comments and rebuttals and we thank her/him for the additional point.

Reviewer GjrH2023-08-11

Thank you for your rebuttal

I think the author have answered all my questions thoroughly. However, I agree with one of the reviewer point that in general this work leans on more of a statistical methodology paper, therefore I still maintain my score as Weak Accept, as I do not see a major problem with it.

Authorsrebuttal2023-08-12

reply

We are grateful to the Reviewer for their response and their thorough examination of all the comments and rebuttals. For the purpose of clarification, we would like to ask whether the final score of the reviewer is "Weak Accept" (6) or a "Borderline Accept" (5)? Thank you in advance for your reply.

Reviewer GjrH2023-08-14

Sorry for the confusion

What I meant is Borderline Accept (5) score, but of course this reflects my opinion that I would not be upset if the work is accepted to NeuRIPS, as I saw some clear contributions to the conference.

Reviewer qxf12023-08-14

I have read the response of the authors and thank them for considering my remarks and those of my fellow reviewers. As for the R script, when I run the code from the "#SIMULATION" comment downwards as suggested by the authors, I still get the same error. Overall, I wish to maintain my score of 6.

Reviewer vLRc2023-08-15

I'm confused by this comment. Bonferroni would do just fine achieving a goal of combining p-values, Bonferroni combination function is not smooth and does not fall into your framework. Your answer does not give me great confidence that you actually reviews multiple hypothesis testing thoroughly. Please formalize the difference between multiple testing and meta analysis problem. Temporarily moving down to 5.

Authorsrebuttal2023-08-16

We are sorry to hear that our response caused confusion. We provide below a detailed response on the difference between multiple testing and meta-analysis, in addition to explaining that our framework actually includes Bonferonni's method. **On Bonferonni's method:** Bonferonni's method would entail combining p-values as $m \cdot \min \{p^{(1)},\dots,p^{(m)} \}$ (see e.g. display (1) in [VOVK, V., AND WANG, R. Combining p-values via averaging - Biometrika 107]). This is in fact covered by our framework, on page 7 we discuss generalized averages as defined for example in [VOVK, V., AND WANG, R. Combining p-values via averaging - Biometrika 107]. This framework also encompasses the Bonferonni correction, which is a generalized average (with $a_{r,m} = m$ in the notation of page 7 of our paper). We would also like to highlight that our framework is more general than just smooth combination functions, as Assumption 3 concerns only Hölder continuity (which is satisfied for Tippett's method or a Bonferonni correction as well). We did not highlight this method as it is more conservative than e.g. Tippett's method ($ \min \{p^{(1)},\dots,p^{(m)} \}$), and is often considered when the p-values combined might have dependencies because they e.g. concern the same data, which is the case in multiple testing, see the definition below. **Comparing the mulitple testing to meta-analysis:** *Multiple testing*: Let $X$ be some data drawn from an unknown distribution $P$. Based on the true distribution $P$, a hypothesis $H$ is either true or false; which we denote by $H$ is being true if it belongs to a set $\mathcal{T}_0$ and false if it belongs to $\mathcal{T}_1$ otherwise. Given a collection of such hypotheses $\mathcal{H}$, using the data $X$ one tries do discern $$H \in \mathcal{T}_0 \text{ versus } H \in \mathcal{T}_1$$ for all $H \in \mathcal{H}$ simultaneously. This very general definition of multiple testing, see for example [FROMONT ET AL. Family-Wise Separation Rates for multiple testing - Ann. Statist. 44(6)] Prototypical examples are testing for each gene in a sequence separately whether the gene plays a role in a given disease, or testing the returns of different portfolios, for finding which portfolios have higher than market returns. For example, a multiple testing problem in the context of the many-normal-means model considered in our paper is $$H_{0k}: f_k = 0 \text{ versus } H_{1k}: |f_k| \geq \rho_k$$ for $k=1,\dots,k$, given data $X = f + \frac{1}{\sqrt{n}}Z$, $f \in \mathbb{R}^d$. For such a collection of hypotheses, one tries to discerns multiple, different hypotheses on the basis of the same data. Standard approaches in multiple testing are: Bonferonni correction and Holmes method (to control the family-wise error rate) and Benjamini-Hochberg (controlling the false discorvery rate). *Meta-analysis* can be performed when there are multiple scientific studies addressing the *same question* (see e.g. [Hedges et. al - Introduction to meta-analysis] or Wikipedia). In our analysis, we consider testing, where $m$ studies address the *same hypothesis* and the goal is to combine the study outcomes (e.g. their reported p-values). Prototypical examples would be multiple experiments conducted to establish whether a given drug has *any* effect (e.g. whether a given blood pressure medication indeed lowers the bloodpressure), or multiple studies concerning the question whether smoking causes cancer. In our setting, we only consider the same hypothesis (i.e. $\mathcal{H}$ is a singleton) in each study: $$H_{0}: f = 0 \text{ versus } H_{1}: \|f\| > \rho. $$ *In conclusion:* Although a Bonferonni correction falls within our framework, it is unnecessarily conservative as we do not use the same data for testing more than one hypothesis. Therefore we did not explicitly mention it, as e.g. Tippett's method is more appropriate for our setting. Nevertheless, in our updated version we explicitly refer to this method as well as an example of a generalized average. We hope that our definitions above highlighting the conceptual differences of meta-analysis and multiple testing are satisfactory.

Reviewer vLRc2023-08-21

Thank you for your response. It is satisfactory, and I am increasing the score. I appreciate the effort put into the rebuttal and the paper.

Authorsrebuttal2023-08-21

We thank the Reviewer for the reconsideration and for the increased points.

Authorsrebuttal2023-08-17

We thank the Referee for the reply and appreciate that he/she has read the other reviews and rebuttals as well. We are grateful for the reconsideration and for the additional points.

Reviewer aj9v2023-08-18

Answer to rebuttal

I have read the authors' response and thank them for taking the remarks into consideration. For the simulation study: Thank you for rerunning the simulations, the new plots are cleaner and easier to interpret. The authors have thoroughly answered my concern about the potential bias due to the Rademacher prior. Concerning the heterogeneity and different sample sizes: Including a discussion in the revised paper is sufficient, and I agree with the authors that a rigorous analysis of the setting where $m \min_j n_j \ll \sum_j n_j$ is beyond the scope of the present paper. All in all, the authors' answers were satisfactory, as they cleared up a potential weakness and added discussions on potential extensions of their work. The paper is technically solid and well presented. Overall, I wish to maintain my score to 6.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC