Summary
The paper extends non-parametric machine teaching to the case of multiple learners. In particular, each learner learns one component of a vector-valued function. The authors consider the case where learners have no communication with each other and the case where there is "communication" via a matrix transformation of the outputs of each individual learner.
Strengths
Originality: I am not an expert in machine teaching so I can't say how original the contribution is relative to prior work.
Quality: The paper appears to be of good quality, though I did not check the proofs carefully.
Clarity: The paper was relatively clear, though with key exceptions referenced below
Significance: The paper appears to me to be an incremental step beyond the non-parametric single-learner setting.
Weaknesses
I believe the paper needs more motivation and needs to contrast with the case where a single learner is learning a vector-valued function. Currently, the paper extends the scalar-valued single-learner setting to one where the teacher is teaching multiple learners---each outputting one component of a vector-valued output. In the introduction, the paper explains why teaching multiple learners at once would be better than teaching multiple learners separately. That seems quite clear, but then why not just teach one single learner that is learning a vector-valued function rather than multiple learners learning each component separately? The introduction seems to hint at computational issues but this should be made more explicit, and if computational issues are a critical motivation, then experiments showing the computational advantage should be included.
Especially because the paper also includes a section on "communication" between the learners and shows that they achieve better performance when communication is allowed. (By the way, the communication, in this case, is just a matrix that the teacher gives each learner.. the learners aren't learning to communicate with each other, so it seems clear from the outset that this would lead to improvement) So if communication leads to higher performance, then why not full "communication", i.e., one learner that outputs all components?
Questions
I thought the experiments section could be made clearer (I also looked at the additional details in the appendix but was still confused). What is the domain and range of the functions being learned? What are the examples (x and y) that the teacher is giving for each settting? Are they pixels?
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
I ask that the authors replace the Lenna image with any other image. There is no reason that this paper needs to use, of all images, the Lenna image, a cropped photo of a nude women from a Playboy centerfold. The continued unnecessary use of this image perpetuates an uninclusive environment.
https://womenlovetech.com/losing-lena-why-we-need-to-remove-one-image-and-end-techs-original-sin/
https://www.washingtonpost.com/opinions/a-playboy-centerfold-does-not-belong-in-tj-classrooms/2015/04/24/76e87fa4-e47a-11e4-81ea-0649268f729e_story.html
Lena herself has said that she no longer wants her image used: “I retired from modelling a long time ago,” said Lena in a new documentary film called Losing Lena. “It’s time I retired from tech, too. We can make a simple change today that creates a lasting change for tomorrow. Let’s commit to losing me.”