Dear Reviewer,
We have conducted additional experiments on the Oxford-102 Flower dataset [R1] which consists of images and captions. As the images and the captions are formulated from different input sources, it reflects a real-world scenario where input modalities are less redundant. We leveraged Inception v3 [50] and doc2vec [R2] to extract visual and textual features and used samples belonging to the first ten classes (3,485 samples in total) for this experiment. We split the dataset into a training set and a test set with 0.7:0.3 ratio and set the number of context samples/inducing points for ETP, MGP, and MNP to 200. We will include more detailed experimental settings in the revised manuscript.
The same experimental procedure of the main experiment in section 5.1 was conducted, and the results are as follows:
\\begin{array}{cccc} \\hline {} & \\text{Test accuracy} \\uparrow & \\text{Test ECE} \\downarrow & \\text{Average test accuracy across 10 noise levels} \\uparrow \\\ \\hline \\text{MCD} & 94.86±1.44 & 0.195±0.012 & 79.65±0.33 \\\ \\text{DE(EF)} & 98.29±0.80 & 0.087±0.007 & \\underline{88.65±0.3} \\\ \text{SNGP} & 93.14±1.04 & 0.454±0.022 & 69.00±1.17 \\\ \\text{ETP} & 98.10±0.95 & 0.068±0.014 & 82.09±0.85 \\\ \\text{DE(LF)} & 96.76±1.09 & 0.394±0.018 & 83.40±1.05 \\\ \\text{TMC} & 94.67±1.09 & 0.123±0.008 & 81.40±0.84 \\\ \\text{MGP} & \\mathbf{98.67±0.52} & \\underline{0.037±0.008} & 76.55±0.55 \\\ \\text{MNP (Ours)} & \\underline{98.48±1.09} & \\mathbf{0.017±0.005} & \\mathbf{94.19±0.42} \\\ \\hline \\end{array}
While there is a marginal difference in test accuracy of DE(EF), ETP, MGP, and MNP, a large gap in test ECE and average test accuracy with noisy samples was observed. This illustrates MNP’s robustness and calibration performance that outperform other baselines. More importantly, MNP is able to maintain test accuracy under noisy conditions with a slight decrease in accuracy (98.48->94.19). Other baselines have shown significant decrease in accuracy (e.g., MGP: 98.67->76.55 or ETP: 98.10->82.09). This experiment, along with CUB, shows the effectiveness of MNP's uncertainty estimation performance across diverse input modalities.
Additional references:
[R1] S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee. Generative adversarial text to image synthesis. In M. F. Balcan and K. Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1060–1069, New York, New York, USA, 20–22 Jun 2016. PMLR.
[R2] Q. Le and T. Mikolov. Distributed representations of sentences and documents. In E. P. Xing and T. Jebara, editors, Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pages 1188–1196, Bejing, China, 22–24 Jun 2014. PMLR.
Kind regards,
The authors