Since the vocal component plays a crucial role in popular music, singing\nvoice detection has been an active research topic in music information\nretrieval. Although several proposed algorithms have shown high performances,\nwe argue that there still is a room to improve to build a more robust singing\nvoice detection system. In order to identify the area of improvement, we first\nperform an error analysis on three recent singing voice detection systems.\nBased on the analysis, we design novel methods to test the systems on multiple\nsets of internally curated and generated data to further examine the pitfalls,\nwhich are not clearly revealed with the current datasets. From the experiment\nresults, we also propose several directions towards building a more robust\nsinging voice detector.\n