Response to Reviewer cFix
Table 2: Main results of federated OOD detection and generalization on Cifar100. We report the ACC of brightness as IN-C ACC, the FPR95 and AUROC of LSUN-C as OUT performance.
| Non-IID \ Method | $\alpha=0.1$ | | | | $\alpha=0.5$ | | | | $\alpha=5$ | | | |
| :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
| | ACC-IN $\uparrow$ | ACC-IN-C $\uparrow$ | FPR95 $\downarrow$ | AUROC $\uparrow$ | ACC-IN $\uparrow$ | ACC-IN-C $\uparrow$ | FPR95 $\downarrow$ | AUROC $\uparrow$ | ACC-IN $\uparrow$ | ACC-IN-C $\uparrow$ | FPR95 $\downarrow$ | AUROC $\uparrow$ |
| FedAvg | 51.67 $\pm$ 1.37 | 47.54 $\pm$ 0.48 | 78.35 $\pm$ 1.64 | 67.16 $\pm$ 1.17 | 58.28 $\pm$ 0.48 | 54.62 $\pm$ 0.67 | 72.84 $\pm$ 0.81 | 70.86 $\pm$ 1.52 | 61.40 $\pm$ 0.12 | 56.72 $\pm$ 0.17 | 72.68 $\pm$ 0.34 | 70.59 $\pm$ 0.19 |
| FedLN | 52.48 $\pm$ 1.41 | 48.15 $\pm$ 1.57 | 66.94 $\pm$ 1.61 | 74.82 $\pm$ 0.50 | 59.39 $\pm$ 0.72 | 53.86 $\pm$ 1.23 | 68.31 $\pm$ 1.24 | 73.41 $\pm$ 0.33 | 61.00 $\pm$ 0.40 | 56.33 $\pm$ 0.82 | 69.18 $\pm$ 0.46 | 75.87 $\pm$ 0.74 |
| FedATOL | 43.65 $\pm$ 0.54 | 41.08 $\pm$ 0.60 | 65.26 $\pm$ 0.96 | 81.64 $\pm$ 0.33 | 60.62 $\pm$ 0.61 | 56.63 $\pm$ 0.91 | 70.10 $\pm$ 0.81 | 79.27 $\pm$ 0.61 | 64.16 $\pm$ 0.81 | 63.61 $\pm$ 0.42 | 80.27 $\pm$ 1.61 | 60.51 $\pm$ 1.75 |
| FedT3A | 51.67 $\pm$ 1.37 | 51.50 $\pm$ 0.29 | 78.35 $\pm$ 1.64 | 67.16 $\pm$ 1.17 | 58.28 $\pm$ 0.48 | 55.42 $\pm$ 1.63 | 72.84 $\pm$ 1.56 | 70.86 $\pm$ 1.52 | 61.40 $\pm$ 0.12 | 55.51 $\pm$ 0.96 | 72.68 $\pm$ 0.34 | 70.59 $\pm$ 0.19 |
| FedIIR | 51.63 $\pm$ 0.61 | 47.88 $\pm$ 1.19 | 81.91 $\pm$ 0.47 | 63.99 $\pm$ 0.53 | 58.66 $\pm$ 0.41 | 55.72 $\pm$ 0.29 | 77.62 $\pm$ 1.10 | 65.87 $\pm$ 0.46 | 61.70 $\pm$ 0.76 | 57.65 $\pm$ 0.80 | 72.57 $\pm$ 0.37 | 69.07 $\pm$ 0.52 |
| FedAvg+FOOGD | 53.84 $\pm$ 0.83 | 51.69 $\pm$ 0.32 | 36.40 $\pm$ 1.11 | 91.41 $\pm$ 0.36 | 61.82 $\pm$ 0.20 | 59.91 $\pm$ 0.31 | 55.70 $\pm$ 0.78 | 86.42 $\pm$ 0.24 | 64.96 $\pm$ 0.51 | 64.18 $\pm$ 0.31 | 57.70 $\pm$ 0.87 | 84.03 $\pm$ 0.15 |
| FedRoD | 73.13 $\pm$ 0.85|69.26 $\pm$ 0.41| 66.34 $\pm$ 1.53 | 73.02 $\pm$ 1.82 | 66.88 $\pm$ 0.61 | 61.28 $\pm$ 0.98 | 70.13 $\pm$ 0.86 | 69.48 $\pm$ 0.65 | 61.34 $\pm$ 0.78 | 55.80 $\pm$ 1.21 | 74.86 $\pm$ 0.98 | 67.76 $\pm$ 1.31 |
| FOSTER | 72.54 $\pm$ 1.51 | 67.50 $\pm$ 0.57| 61.25 $\pm$ 1.05 | 75.44 $\pm$ 0.89 | 62.45 $\pm$ 0.55 | 57.62 $\pm$ 0.87 | 73.26 $\pm$ 1.13 | 68.71 $\pm$ 0.85 | 53.80 $\pm$ 0.31 | 49.28 $\pm$ 0.74 | 76.94 $\pm$ 1.62 | 65.47 $\pm$ 1.72 |
| FedTHE | 73.83 $\pm$ 0.48 | 69.09 $\pm$ 0.56 | 64.73 $\pm$ 0.79 | 75.16 $\pm$ 0.34 | 66.22 $\pm$ 0.68 |61.19 $\pm$ 0.92 | 72.95 $\pm$ 1.84 | 69.38 $\pm$ 1.64 | 61.03 $\pm$ 0.22 | 57.03 $\pm$ 0.16 | 71.43 $\pm$ 0.64 | 69.01 $\pm$ 0.87 |
| FedICON | 72.22 $\pm$ 0.72 |67.79 $\pm$ 0.31 | 61.36 $\pm$ 0.39 | 77.12 $\pm$ 0.55 | 65.86 $\pm$ 0.81 |61.83 $\pm$ 0.55 | 69.99 $\pm$ 1.02 | 71.03 $\pm$ 0.39 | 62.11 $\pm$ 0.74|57.62 $\pm$ 0.28 | 70.91 $\pm$ 0.97 | 70.84 $\pm$ 0.73 |
| FedRoD+FOOGD | 77.88 $\pm$ 0.28 |75.70 $\pm$ 0.26|58.81 $\pm$ 0.48 | 86.07 $\pm$ 0.39 | 70.30 $\pm$ 0.46 | 68.23 $\pm$ 0.25| 45.19 $\pm$ 0.67 | 89.59 $\pm$ 0.28 | 64.94 $\pm$ 0.79 | 62.56 $\pm$ 0.72 | 65.18 $\pm$ 1.19 |80.47 $\pm$ 0.32|
**As we can see, the model performance variance of FOOGD is very small, indicating the performance stability of FOOGD in OOD data. This also validates that our model is effective and practical for federated learning with wild data.**
**For weakness 2, we reply that ACC-IN and ACC-IN-C are commonly used accuracies for the two test sets, i.e., in-distribution data (Non-IID data) and covariate-shift data (OOD Generalization), respectively**. We have introduced them in lines 278-2799 of the main paper. This is used in federated learning considering semantic shifts, e.g., FedTHE, where ACC-IN corresponds to original local test, and ACC-IN-C corresponds to corrupted local test.
**For weakness 3, the related work, e.g., FedAvg as well as their implementations are listed in lines 664-693, please kindly refer to this part for more details.** In this work, we mainly focus on tackle the OOD shifts in federated learning, which is orthogonal to the existing work that tackles heterogeneity. We will add a section for introducing work related on FL with non-IID data.
**The reference you indicated, i.e., ref.46, is the SVHN dataset that is commonly used in detecting OUT samples.** SVHN is a real-world image dataset from Google Street View house numbers, comprising 73,257 training samples and 26,032 testing samples across 10 classes. We introduce it as an OOD detection task in line 267 and line 646 of paper.